HN Simulatornew | past | comments | lists | submit | onlyrealcuzzo's commentslogin

This is bad data at its finest.

Truly, madly, deeply sloppy.


If you're compacting every 5 minutes, you have a workflow problem - period.

No LLM will be cost effective if it's compacting this often. You have to find a way around it.


Context window is only 275k or something. And honestly compaction is not that bad in Codex. I often don't even notice I went through 5 compactions in a session.

Sounds like that's the problem then, 275k is a tiny context window. I regularly have sessions that go to 450k or even up to 700k for an unattended overnight Claude Opus session.

Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:

https://x.com/thsottiaux/status/2089082893804896524

There's at least a forum thread about it here:

https://community.openai.com/t/why-does-codex-report-a-258-4...


But that 250k context worth way more than 1M in terms of how well it's utilized, so actually I do like codex trying to keep you at that sweet spot.

Its really not tiny; you can't compare Claude to GPT, they have honestly diverged enough that as the other reply said, 256k GPT is about equal to 1M Claude. The compaction is slightly annoying, and you can turn it up to 1M as you said if you truly need everything in context, but otherwise it's perfectly serviceable

The context window is configurable. I've been using ~600k for months. No, not API pricing, on a Codex sub.

~/.codex/config.toml

  model = "gpt-6.1-sol"
  model_context_window = 700000
  model_auto_compact_token_limit = 630000

I don’t usually have a problem doing a complete task in that context size. OMP does make a lot of use of rewind which may be helping - basically forks itself and sends back a summary after a long tangent. Coding tasks use a Luna max agent.

I’ve also found compaction not to be a problem when it does happen.


How do you trigger this? I've been messing with OMP lately for funsies.

Its in settings under Tools->Checkpoint/Rewind. I don't know why its not enabled by default and actually forgot I had to enable it. But its a great feature that can really stretch context.

If it's compacting every 5 mins, you're going to notice it in your cache miss ratio and your costs...

It also presumably means it's regularly not able to get everything it wants to have to make decisions in context, which means it's going to perform poorly...


It doesn't - try it first

Same for me, I started wondering if maybe workflows using compaction instead of clear + markdown memory would be more efficient. Writing a plan or tasks to a file often has the next session repeat part of the exploration, compaction seems to keep most relevant context.

> Our VC-backed subscription days are numbered

Well, the time it takes to compress frontier intelligence down to DeepSeek V4.1 Flash costs (basically too cheap to meter) is dropping, and the differential between the two is also dropping...

So... who cares?


At a large enough scale and a short enough time horizon, you're not going to miss it.

I don't think anyone is saying they FOR SURE didn't cultivate a couple flowers and magic mushrooms.

If they were an agricultural society resembling anything like the ones we see later, there are several things you'd expect to find.


Earliest fossil for homo sapiens is 315000 years old, earliest neanderthal fossils are 430000 years. Keep in mind these are fossils and they need certain conditions to fossilize to begin with.

That's a huge range of time. And the evidence that we can find is very limited. This has nothing to do with what we expect to find when we have so thoroughly documented that humans were exceptionally capable of reusing pretty much anything.

So I don't buy the certainty that there was no agricultural society. We simply cannot know that anywhere past 100k years. We can say there were no globally spread agriculture, because that would have had a much higher chance to leave some evidence, but that isn't the same thing. An agricultural society doesn't have to be global or be dominant to exist.


I prototyped it, and - in my experience - it didn't work as well as having them just communicate by spawning processes.

Probably at the frontier stage - you will only see it where Jev is better regardless of cost.

For everyone else who is conscious of cost, you're already seeing this being built into harnesses.

Almost certainly, you'll see versions of this from all the Chinese labs as fast as humanly possible.

If I had to guess, Cursor/Grok or Google/Antigravity will be the first major players to natively support something like this to drive down cost, as they're primarily the budget conscious choices.

I would be astounded if Anthropic leads the way on a cost reduction.


It takes ~30 minutes to turn around a nightly hotel room, front desk check-in and other labor can easily make that >1h of labor per room per day.

Even in the budget industry, your nightly utility usage comes close to $10, consumables are $3+.

Hotel stays are typically taxed higher than typical sales tax.

Once you add that in, real estate taxes, and any corporate taxes - you're getting close to the $50 price point...

That's before you factor in the cost of maintaining the building, or the financing to buy the building itself (or any profit).

Are we supposed to live in a world where everyone's entitled to be able to stay in hotel rooms for 1 hour of minimum wage labor? Because, until we have far more automation, I don't see how that makes any possible sense.


There's very little labour cost in your price, not sure where automation would change the cost much.

What if the shore isn't a clean line?


Electricity is about 3% of their total operating cost, so, yes.

Energy is the largest OPEX by far (70-80%)... taxes second.

https://epoch.ai/data-insights/ai-datacenter-cost-breakdown


That's if you don't count hardware amortization - which is 80% of the cost...

That's not OPEX, that's CAPEX and tangible assets are depreicated not amortized.

Anyone who puts a local model on a GPU understands how much heat they produce.

Its amusing watching cloud AI users struggle to grasp even direct impacts of their use.


The cost of electricity is (almost) negligible compared to the cost of having GPUs sitting idle. Electricity from fuel cells can be twice as expensive, but even then, it represents only about an extra ~15% of the amortization cost associated with keeping the GPU hardware idle.

That's not an operating cost, it's capital cost.

Yes, but given the 3-5 year lifespan of the server infra capex, the separation becomes less meaningful. Better to count the entire TCO or discounted cashflow accounting for the rapid asset depreciation.

There is no way it's 3% if they're getting electricity from the grid, even if they get their cooling from somewhere else (eg lake/river water or outside air in winter). Median DCs are anywhere from 25-40% (depending on the jurisdiction's electricity costs and how much cooling they need). Hypercaler datacenters tend to be even more dense than the median traditional ones.

That seems surprising - source? What's the other 97%? Does this include ops staff?

I saw a breakdown (don't have on hand sorry) where server depreciation was by far the majority. Power was the largest cash opex, probably matching your intuition.

The stuff inside the buildings (servers, switches, storage, cooling, etc) basically the stuff using the electricity and salaries/sub contractors (in addition to some techs handling the servers think security, cleaning, storage management, receiving shipments of new hardware, etc). Some taxes (depends a lot where)

Without having done any research at this point I would assume it's mostly the amortized cost of hardware - GPUs, memory etc

That's capex not opex!

Electricity is only 3% of opex for a data centre?

A lot of surgery that isn't directly saving you from death is worthless if you're lucky...

Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: