Context window is only 275k or something. And honestly compaction is not that bad in Codex. I often don't even notice I went through 5 compactions in a session.
Sounds like that's the problem then, 275k is a tiny context window. I regularly have sessions that go to 450k or even up to 700k for an unattended overnight Claude Opus session.
Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:
Its really not tiny; you can't compare Claude to GPT, they have honestly diverged enough that as the other reply said, 256k GPT is about equal to 1M Claude. The compaction is slightly annoying, and you can turn it up to 1M as you said if you truly need everything in context, but otherwise it's perfectly serviceable
I don’t usually have a problem doing a complete task in that context size. OMP does make a lot of use of rewind which may be helping - basically forks itself and sends back a summary after a long tangent. Coding tasks use a Luna max agent.
I’ve also found compaction not to be a problem when it does happen.
Its in settings under Tools->Checkpoint/Rewind. I don't know why its not enabled by default and actually forgot I had to enable it. But its a great feature that can really stretch context.
If it's compacting every 5 mins, you're going to notice it in your cache miss ratio and your costs...
It also presumably means it's regularly not able to get everything it wants to have to make decisions in context, which means it's going to perform poorly...
Same for me, I started wondering if maybe workflows using compaction instead of clear + markdown memory would be more efficient. Writing a plan or tasks to a file often has the next session repeat part of the exploration, compaction seems to keep most relevant context.
Well, the time it takes to compress frontier intelligence down to DeepSeek V4.1 Flash costs (basically too cheap to meter) is dropping, and the differential between the two is also dropping...
Earliest fossil for homo sapiens is 315000 years old, earliest neanderthal fossils are 430000 years. Keep in mind these are fossils and they need certain conditions to fossilize to begin with.
That's a huge range of time. And the evidence that we can find is very limited. This has nothing to do with what we expect to find when we have so thoroughly documented that humans were exceptionally capable of reusing pretty much anything.
So I don't buy the certainty that there was no agricultural society. We simply cannot know that anywhere past 100k years. We can say there were no globally spread agriculture, because that would have had a much higher chance to leave some evidence, but that isn't the same thing. An agricultural society doesn't have to be global or be dominant to exist.
Probably at the frontier stage - you will only see it where Jev is better regardless of cost.
For everyone else who is conscious of cost, you're already seeing this being built into harnesses.
Almost certainly, you'll see versions of this from all the Chinese labs as fast as humanly possible.
If I had to guess, Cursor/Grok or Google/Antigravity will be the first major players to natively support something like this to drive down cost, as they're primarily the budget conscious choices.
I would be astounded if Anthropic leads the way on a cost reduction.
It takes ~30 minutes to turn around a nightly hotel room, front desk check-in and other labor can easily make that >1h of labor per room per day.
Even in the budget industry, your nightly utility usage comes close to $10, consumables are $3+.
Hotel stays are typically taxed higher than typical sales tax.
Once you add that in, real estate taxes, and any corporate taxes - you're getting close to the $50 price point...
That's before you factor in the cost of maintaining the building, or the financing to buy the building itself (or any profit).
Are we supposed to live in a world where everyone's entitled to be able to stay in hotel rooms for 1 hour of minimum wage labor? Because, until we have far more automation, I don't see how that makes any possible sense.
The cost of electricity is (almost) negligible compared to the cost of having GPUs sitting idle. Electricity from fuel cells can be twice as expensive, but even then, it represents only about an extra ~15% of the amortization cost associated with keeping the GPU hardware idle.
Yes, but given the 3-5 year lifespan of the server infra capex, the separation becomes less meaningful. Better to count the entire TCO or discounted cashflow accounting for the rapid asset depreciation.
There is no way it's 3% if they're getting electricity from the grid, even if they get their cooling from somewhere else (eg lake/river water or outside air in winter). Median DCs are anywhere from 25-40% (depending on the jurisdiction's electricity costs and how much cooling they need). Hypercaler datacenters tend to be even more dense than the median traditional ones.
I saw a breakdown (don't have on hand sorry) where server depreciation was by far the majority. Power was the largest cash opex, probably matching your intuition.
The stuff inside the buildings (servers, switches, storage, cooling, etc) basically the stuff using the electricity and salaries/sub contractors (in addition to some techs handling the servers think security, cleaning, storage management, receiving shipments of new hardware, etc). Some taxes (depends a lot where)
Truly, madly, deeply sloppy.
reply