Also if you have less than 24GB VRAM, then ollama defaults to 4K context. If that "Qwen 3.8" uses thinking, it might be running out of context and forgetting what it was even answering mid-generation. If that's the case, then try increasing context size: https://docs.ollama.com/context-length , but also: https://sleepingrobots.com/dreams/stop-using-ollama/
Probably this:
https://huggingface.co/empero-ai/Qwen3.8-4B-Distill
> Qwen3.8-4B is a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-4B architecture.
https://huggingface.co/Qwen/Qwen3-4B/blob/main/README.md
SO, if that isn't official it explains the results I got. They were dreadful.
https://huggingface.co/Qwen/Qwen3.5-4B
> commited on May 21, 2025, over 1 year ago
Fairly old update to the README.md of (instead of Qwen3.8), should have raised some flags?
Also if you have less than 24GB VRAM, then ollama defaults to 4K context. If that "Qwen 3.8" uses thinking, it might be running out of context and forgetting what it was even answering mid-generation. If that's the case, then try increasing context size: https://docs.ollama.com/context-length , but also: https://sleepingrobots.com/dreams/stop-using-ollama/