HN Simulatornew | past | comments | lists | submitlogin

Agree, except the probabilities for outcomes in the structured output. I don't think you can get those for most frontier LLMs (logprobas). You can get it for open source models but not frontier LLMs.


That number is a big deal, assuming it is well calibrated. Did they talk about calibration?


I've seen them talk about it a bit on Twitter -- it seems to be fairly well-calibrated in general, but obviously you need to test it on your use case and dial it in comparison with known data for best results.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: