I use models.dev's CLI tool, which I think gets data from OpenRouter, and ArtificialAnalysis so your coding agent can help you narrow it down.
Similarly, HuggingFace has a CLI + a few skills, and they are very useful.
I had a production image processing using Gemini 2.5 Flash Lite (which is getting discontinued in October), and in 20 minutes Claude Code + HF Cli recommended the best replacement small model (Qwen VL 3B something) and proceeded to fine tune it on my datataset. All this while I was in a rush to get dressed and go to the store.
It cost ~$3 I think, and results were excellent. Not perfect, but not far from perfect either.
We didn't replace Gemini in prod at the time, because we didn't have time to do all the math on how to end up with a smaller bill/mo.
Similarly, HuggingFace has a CLI + a few skills, and they are very useful.
I had a production image processing using Gemini 2.5 Flash Lite (which is getting discontinued in October), and in 20 minutes Claude Code + HF Cli recommended the best replacement small model (Qwen VL 3B something) and proceeded to fine tune it on my datataset. All this while I was in a rush to get dressed and go to the store.
It cost ~$3 I think, and results were excellent. Not perfect, but not far from perfect either.
We didn't replace Gemini in prod at the time, because we didn't have time to do all the math on how to end up with a smaller bill/mo.