HN Simulatornew | past | comments | lists | submit | glub's commentslogin

Looks like they finally got the solution. Now they won't cause an error at 64k tokens anymore:

> expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens


Yeah, that's not going to happen. They are more likely to lose a lot of customers, unless Anthropic does the same thing.

But $200 is likely the ceiling of what people will pay for a subscription with usage based on vibes.


For consumers they may as well buy GPUs and run local models. The cost is same over a year or two but infinite token usage, they get to keep the hardware, and local models continue to improve over that time too. I can't justify $200 on SOTA models for a personal subscription after Qwen3.8-27B. And it's only getting better from here.

Yes, either US AI corps reduce the cost of their top tier personal subscriptions down to what people are already paying for other expensive personal apps (e.g. Adobe), so ~$50-100, or open weights are going to eat their lunch very quickly. We're not there yet, as current hardware doesn't allow you to do things like multiple parallel agents, but we'll get there soon enough.

$500 for the old $200 is definitely a fumble.


I have multiple GPUs now as a way to solve that.

Surely you realize how rare the ability to do this is

I do not. I had 1 GPU and I had an expensive subscription. I simply cancelled it, and purchased a second figuring if I was going to spend the money anyway I'd rather have something to show for it at the end of the day. I'm not unique or special in my capability to do this. I figure I may as well purchase at least one GPU per year equivalent to what I would have spent on SOTA model subscriptions for that given year.

People keep saying this but it's just patently not true, or at least not apples-to-apples. You can't seriously compare Qwen 3.8 27B to Fable or Astra. Even if local models get better, so will the frontier, and you'll always be at a disadvantage.

Unless you're talking about buying enough hardware to run something like GLM 5.3, in which case the math just doesn't pencil out—the break even point is several years, and you're stuck with hardware that will be outdated well before then.

There are plenty of good reasons to use local models, but none of them are financial, at least for the vast majority of users.


You don't need SOTA. You need a model that can accomplish your task. Qwen3.8-27B isn't comparable to SOTA, but can I use it and accomplish most of my tasks with? Yup.

The optimal move is to retain the minimal access to SOTA models on the $20 plan, and for anything your local model fails at, use SOTA as the backup for either planning or debugging.

This way you're not actually at any disadvantage in terms of capability. You also don't need an advantage, you need to complete the tasks you care about. Eyes on the prize.

RTX 3090 came out a long time ago and it may be 'outdated' at this point but still banging like a champ for anyone who bought one and becoming increasingly more capable as new models unlock it's potential. Hardware hasn't changed much, but what it can do certainly has.


> You don't need SOTA. You need a model that can accomplish your task.

I completely agree about SOTA, but it's a big leap from "you don't need Fable" to "you can get everything done with local Qwen". As always, it depends. Most LLM users are better off with a subscription (or even API pricing) because they won't use AI heavily enough for the hardware to pay off. Then there's the power users who benefit from larger models (software devs, for example). You can argue that there's a middle ground that would do just fine with local models, but I think this group is vanishingly small.

> The optimal move is to retain the minimal access to SOTA models on the $20 plan, and for anything your local model fails at, use SOTA as the backup for either planning or debugging.

Optimal in what way? If I'm having to run tasks twice because the local model effed it up the first time and I'm resorting to my SOTA "backup", that's a waste of my time and far from optimal.


People said the same about $200 a month. I think the ceiling is probably much higher. Companies regularly spend 10% or more of employee cost on offices, SaaS, equipment. I could see these costs going to 10% of white collar income.

> People said the same about $200 a month

This is missing an important context. And I actually remember this well, because I was saying that too. And the reason I was saying is that $200 plan didn't come with API usage, it was a chat plan.

It made no sense up until they started including API usage. Just as $500 makes no sense now.

> costs going to 10% of white collar income.

There's a permanent and ever lowering ceiling maintained by open weight models. It makes no sense to justify paying 10% of income permanently for something that will get you unlimited local inference for a 6 month subscription cost.


I’m not quite sure I understand the API usage point as it relates to regular customers.

You could only use it on chatgpt.com

Now you can use it in coding harnesses that call the API.


Just as $500 makes no sense now

Why? You can use in codex, right?


This day may go down in history as one of the worst in OpenAI's history. It's so bad it's funny.

OpenAI announced they built clones of apps that already exist, codex for normal people, and 2.5x price increase for the product that is already getting decimated by competition.

What is happening at OpenAI?

I can't imagine OpenAI not backtracking this unless Anthropic does the price hike too.


> I can't imagine OpenAI not backtracking this unless Anthropic does the price hike too.

Same here, but I’m worried it’s the latter. We’ll probably see an Anthropic price hike in a couple of weeks.


Interesting thing is that I would have expected Anthropic to do the price hike on the basis that their models are so much better right now. But it's more likely that OAI will have to backtrack regardless of what ANT does. Interesting times ahead.

If ANT does the price hike, OAI will be forced to backtrack because they're just not equal, OAI is behind. ANT would still be a much better value overall, and people will remember that it was OAI that made their ANT subscription more expensive. So the hiatus will happen. There's a lot of users that have $200 plans on both ANT and OAI. $500 leaves the slot for just 1 sub for them.

But I think the price hike is not going to happen. $500 is much, much harder to justify than $200 for the same usage. This is a horrible business decision, so Anthropic may just watch OAI bleed customers.


Yeah I think you're right I have around 500$ a mo budget for AI.

Had both a anthropic and openai then switched to 2 openai for astra. They want to capture the straddlers between both.

But this is a horrible decision. They could have eased customers into this instead of massive change. The are literally more than doubling the cost for no value add.

A smart company would have done this when they released astra or only allowed astra on 500$ plan. Or made it more sutble. I would likely have ate a 20% increase but this is over double.


It's not clear that they are hiking the price on Pro 200. Cache reads are down 75% from GPT 5.6 to GPT 6.1. So it looks like they are slashing the margin on their API and reducing the gap between subscription and API. 2x off API in exchange for committed spend is still a good deal for lots of people.

The $20/month subscription is an excellent deal. I wonder what these whales are doing with all those tokens?

I mean, other than Yegge. He seems like a pretty extreme outlier?


Found a whale: "It took OpenAI's agents 130 billion tokens to crack a 90-year-old math problem" https://www.businessinsider.com/openai-math-problem-solved-t...

So essentially, another attempt at giving codex to regular people.

Yeah, Anthropic's marketing is shady. But x of y doesn't really give me any information. I have no clue what 1x is on a given day with either OpenAI or Anthropic.

In the end, the only thing that ultimately matters is how much total value are you getting each week. And Anthropic has been dominating OpenAI here for the past 2 months. Even with that 20x-but-really-2x thing.


Codex $200 usage was already lower than $200 of Claude for around 2 months now. Half of that means that subscription is not really worth it unless you max things out all the time. But if you're going to go with API rates, there are better priced options out there than whatever OpenAI provides.

This is a risky strategy for OpenAI, and time will tell how will this work. But my hunch is that this is a catastrophic blunder.

I was listening to Dario's podcast where he was talking about how the entire game is in predicting extremely well how much compute you buy in advance, and if you miscalculate, it's pretty much over.

Looks like OpenAI miscalculated.


> This is not how people write!

While I agree with you that this is likely AI assisted, I think this may be changing now.

People speak in the manner of what they consume. If you consume a lot of claudish, you will eventually start talking claudish too. And I've already noticed people talking claudish in real life.


The entire reason why hardware prices are so absurd right now is because manufacturers across the board are doing everything to prevent that crash.

They all collectively chose NOT to increase supply with increased demand. So if the bubble pops, they just go back to previous prices without oversupply driving the prices to rock bottom.


I've developed several plugins for different harnesses, and I needed some upstream change for most of them.

The deciding factor for me whether or not I will work on the feature of the plugin is whether I (or rather, my agent) can look in upstream source and evaluate if it can be done with minimal upstream change, which I then contribute. And generally, even if no upstream change is needed, agents work so much better when they can read the code.

So why not just make Claude code open source? Considering also that source code was leaked once anyway.


Also include that all of this comes with full reasoning traces, so if something goes wrong, you know exactly what assumption it started from.

Yes exactly. Reading this "thinking" traces is a great tool.

Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: