I don't work with large language models, but out of curiosity: mimo's score on DeepSWE keeps going up, so why are they stopping training early? Is it due to budget constraints?
It looks like the curve is flattering, and the current state actually looks slightly cherry picked (it matches a previous spike that looks a bit of an outlier before the result went down). The longer they train, the more they risk getting scooped by another release by someone else. Etc etc it's a judgement call based on all of these factors (and more, including cost/occupying a big cluster as you mention)
+1 at some point, you need to expect to train a much better base model using everything you've learnt. At the least, you probably want to bring on line the next 10 clever RL environments and ideas your team has been cooking up (which will pipeline into v2.7 etc.)
Language is meant for sharing information. I don't think there's anything wrong with using AI to write - the real issue is that AI-generated writing can be hard to understand. Once AI solves the readability problem, its writing will actually be better suited for communicating with humans.
Not only hard to understand, but if one uses it without first understanding the topic, it's a disaster.
Couple weeks ago I had someone linking to a doc page in my ticket "Made this summary, hope it helps". It was 100% AI generated, the person clearly had no understanding of the problem, it only mentioned UI changes although the thing required changes in many systems. Complete waste of time, I don't think I will humor anyone who makes AI summaries anymore.
reply