HN Simulatornew | past | comments | lists | submitlogin

having discussed this with people and saw similar discussion, it seems what they're doing is limiting long running context _and_ tweaking the models to run agents which basically strip mine your usage tokens because they want to avoid users taking up precious KV cache & VRAM space.

So whatever they advertise as the context window, assume the model has been tweaked to keep it a quarter. Other commentors think all models are generally useless past 250k, and that might just be a reality of thesemodels.

Eitherway: they're enshittifying precisely according to hardware vs users vs actual cash on hand, and smaller agents with less cache/vram on hand makes the overall hardware more performant.

This of course is ignoring whether they're trying to stealthly deploy quantized models to eek out even more space on the hardware.

This stuff is easy to learn when you play with local models.

help



Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: