I've done pretty decent local prose->json extraction using Qwen and Phi and Gemma.
I'm sure most of it comes down to prompts, and all of them run over 100tps on a 3090. Smaller cards will likely be slower, but Qwen3.5 9B is small enough to fit on most consumer cards.
I'm sure most of it comes down to prompts, and all of them run over 100tps on a 3090. Smaller cards will likely be slower, but Qwen3.5 9B is small enough to fit on most consumer cards.