Personally, a combination of low tech and semi scientific tests based on what you normally do (redo the same task you used an older model for with a newer one).
So far, Simon Willison's pelican bike "benchmark" is the only one I've found that shows Fable 5.1 beating Opus 5.5. My personal experience has been Opus is unusable on design work it's so terrible. Evaluating whether we should consolidate AWS DMS tasks (Postgres full load and change data capture) into fewer tasks with more tables, Opus 5.5 was factually wrong and needed correction roughly every other turn.
On a "help me find a sandbox solution for agents embedded in a web app to run untrusted code" research project it kept misrepresenting security boundaries and ended up recommending DuckDB which ironically specifically says it does not provide a strong security boundary in its own documentation. GPT 6 (can't remember if it was Sol or Astra) and Fable 5.1 both recommended FaaS like Cloudflare Workers and AWS Lambda which fit fairly well with the requirements.
I switched from Opus to Fable in the session going badly sideways and told it to "Review the previous conversation and come up with a correct comparison table and corrected recommendations grounded in objectivity supported by citations. Do research as necessary to understand the current ecosystem" and that was a full 180 back to coherency...
So far, Simon Willison's pelican bike "benchmark" is the only one I've found that shows Fable 5.1 beating Opus 5.5. My personal experience has been Opus is unusable on design work it's so terrible. Evaluating whether we should consolidate AWS DMS tasks (Postgres full load and change data capture) into fewer tasks with more tables, Opus 5.5 was factually wrong and needed correction roughly every other turn.
On a "help me find a sandbox solution for agents embedded in a web app to run untrusted code" research project it kept misrepresenting security boundaries and ended up recommending DuckDB which ironically specifically says it does not provide a strong security boundary in its own documentation. GPT 6 (can't remember if it was Sol or Astra) and Fable 5.1 both recommended FaaS like Cloudflare Workers and AWS Lambda which fit fairly well with the requirements.
I switched from Opus to Fable in the session going badly sideways and told it to "Review the previous conversation and come up with a correct comparison table and corrected recommendations grounded in objectivity supported by citations. Do research as necessary to understand the current ecosystem" and that was a full 180 back to coherency...