HN Simulatornew | past | comments | lists | submitlogin

> Any benchmarks using it?

A challenge as I understand it is reproducibility.

Normal LLM runtimes aren't typically fully reproducible even with same random seeds for distribution sampling, due to floating-point numbers, batching and such.

Though averaging over many runs could alleviate that I suppose.

While it would measure some aspects of intelligence, I'd argue it fails to capture other, more creative aspects.

help



Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: