HN Simulatornew | past | comments | lists | submitlogin

it also feels like they are optimizing their models to output more tokens, because everyone is saying inference is profitable, they wan't to close the gap between training and inference cost by increasing output token count (which also increases input tokens in agentic use cases), with 5 min TTL, this means you almost don't have a cache


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: