I've found Droid-ify more reliable at doing background updates, though it still only manages it sometimes (possibly due in some way to how I installed the apps). Hopefully Fdroid's app will be better at this.
Personally, a combination of low tech and semi scientific tests based on what you normally do (redo the same task you used an older model for with a newer one).
So far, Simon Willison's pelican bike "benchmark" is the only one I've found that shows Fable 5.1 beating Opus 5.5. My personal experience has been Opus is unusable on design work it's so terrible. Evaluating whether we should consolidate AWS DMS tasks (Postgres full load and change data capture) into fewer tasks with more tables, Opus 5.5 was factually wrong and needed correction roughly every other turn.
On a "help me find a sandbox solution for agents embedded in a web app to run untrusted code" research project it kept misrepresenting security boundaries and ended up recommending DuckDB which ironically specifically says it does not provide a strong security boundary in its own documentation. GPT 6 (can't remember if it was Sol or Astra) and Fable 5.1 both recommended FaaS like Cloudflare Workers and AWS Lambda which fit fairly well with the requirements.
I switched from Opus to Fable in the session going badly sideways and told it to "Review the previous conversation and come up with a correct comparison table and corrected recommendations grounded in objectivity supported by citations. Do research as necessary to understand the current ecosystem" and that was a full 180 back to coherency...
I did that but I recognize that even though your submission was a few minutes later than that one, you posted the better link, and you're also an established account (the other post was from a new/throwaway account), so I've restored this submission and moved the comments back to it to reward you.
reply