> though I found it better than Deepseek V4 Flash previously
Same experience here.
But man, switch to V4.1 now! It is much better.
I don't event need to test it for long run and I believe it's crazy good. I call it "AI era model taste" when I judge the model by it's output without reading the bench scores.
Using AI for design will always be derivative. For many cases that is good enough but for good product design, a human element will most likely always be necessary. Until the machine gains the ability to do just that
Sometimes I try to add comments in a new session and the agent just don't have enough context for it to give a comprehensive sentence with full context on the why, then the agent will just describe what it does.
Human comment is in another level to answer the questions mainly like "why do it like this" for the later collaborators or the forget-ed self, so the important blocks live when it is needed and can be eliminated when it does not.
Everyone does this because there's no other options out in the market to create a "new" model easily enough. I just hope this one is good, and what's better is if they could open weight it.
The flash model will always use an outdated Treafik version that is not compatible with the newer docker engine, I tried to deploy some personal services with Traefik and everytime it uses this wrong version, and then fixes the version issue in the thinking chain.
I was thinking to switch to Caddy but with your experience I'm gonna stay with Traefik and bare with the version issue...
I've tried or sometimes be stupid to work on bugs/features and ask with almost identical prompts with same modal and harness set, and yes, they generate totally different results.
Sometimes the output is unusable and even with extended guidance it will still drift away from what I was expecting.
Sometimes the output is just one shot and follows almost whatever I want.
I then be used to work like this, if the model and harness set does not work for one time, I just start a new session and do it again. And currently there is one of my task working like this.
Same experience here.
But man, switch to V4.1 now! It is much better.
I don't event need to test it for long run and I believe it's crazy good. I call it "AI era model taste" when I judge the model by it's output without reading the bench scores.