They are not too difficult to train if you already have infra to train regular LLMs. You can typically replace a few layers train them alone and you're off to the races.
Getting training data that works well for calibrated classification objectives is difficult.
I hear conflicting opinions (including my own) about how well calibrated each of these are. Jev seems to be the best.
But the jev release made obvious the PMF for these models, and the underlying reality is that calibration really doesn't matter much when you're replacing usecases where people were using damn LM head softmax probabilities before, which are nowhere near calibrated.
So now everyone simply finetunes qwen and makes a compared-to-regular-LLM vastly cheaper decision model. And it works for majority of usecases. People mostly only care about accuracy, not confidence.
That was exactly my estimation when I played "raw" against Opus. But Opus with stockfish as a tool and much time available can, from what I experienced, generate good comments.
Synthetic data is used in specific contexts. Slopped up hacker news comments along side natural ones is where watermarking will be used to delimitate them. Not all synthetic data is good.
The thing is, exercise becomes a thing time is even a concern for only if it has been neglected for long enough. It is as important as bathing or brushing and the primary issue is that its not framed as equally important as those things. If it was, people would do it every single day. If they did it every single day, 15 minutes a day is more than enough.
However, given that people "graduate" their 20s without having done this, then everything you say becomes true, people end up so weak that they have to rebuild even basic core strength and posture. So practically what I suggested is useless except for maybe the next generation.
Thank you for sharing, that provided some good context for how to interpret this
Post content:
_____
I wish we didn’t need these again, but here is the honest version of Anthropic’s biology announcement
(Caveat: I haven’t worked in bioinformatics for many years.)
The good: Anthropic ran ~950 Claude agents over a large biological sequence database. Claude searched, wrote code, compared sequences and genomic neighborhoods, and found an interesting pattern that apparently had not been noticed before: a known reverse transcriptase associated with another gene and a repetitive DNA array.
That is cool. Automating this kind of open-ended bioinformatics search at scale is useful, and Claude may have found a lead a human would have missed.
But: Claude did not do a biological experiment. It searched databases and analyzed data.
Humans then took the candidate into the wet lab. And the wet-lab result so far is modest: they showed that the repeat array produces short RNAs.
We still don’t know what the system does. No function, mechanism, phenotype, targeting, defense activity, or programmability has been demonstrated.
This is also where the CRISPR framing gets ahead of the result. Right now, “it has some features reminiscent of known programmable systems” is a hypothesis for what to investigate next, not a discovery that it behaves like CRISPR.
And there is a missing baseline: bioinformatics has had tools for finding unusual gene neighborhoods and candidate systems for years. The interesting comparison is 950 Claude agents vs. an expert using the best existing computational pipelines - not Claude vs. someone manually looking through 200,000 sequences.
So my honest announcement would be:
Claude autonomously found an interesting candidate for a previously uncharacterized biological system. A small human wet-lab experiment confirmed that part of the candidate is expressed. We don’t yet know what it does.
That is a good result.
But in a regular biology lab, this isn’t the finished paper. It is the result you show at lab meeting and say: “This looks interesting. Now we need to figure out what the hell it does.”
Maybe that next step leads to a major discovery. But that discovery hasn’t happened yet.
I am irked by the CRISPR framing. That seems to be IPO positioning.
Good hypotheses are a dime a dozen in life sciences. Biology is very unforgiving and most hypotheses lead to nothing when thoroughly tested. This is true for something as "simple" as enzymes as in this case, but even more true for curing diseases. Otherwise, there would not be any failures of phase III clinical trials, after billions USD spent on preclinical research and prior clinical trials.
When overinterpreting these (interesting) results, you are entering Andy Grove Fallacy [0] territory very fast.
He was a PhD student. He knows the significance level of this result. He knows that if he had walked into Bill’s office (his advisor) with “we found an interesting system, but we still don’t know what it does” and said he was ready to graduate, Bill would have kicked him out of the room.
But somehow, when the IPO is around the corner, this becomes “AI is starting to drive biological discovery.”
100% AI per Pangram. I caught it at "This is also where the CRISPR framing gets ahead of the result." -- somehow this is not a sentence anybody non-obnoxious would write. It's a weird structure where the AI talks about something specific as if it were an example of a common theme. This paragraph is an even clearer ekample:
"But in a regular biology lab, this isn’t the finished paper. It is the result you show at lab meeting and say: “This looks interesting. Now we need to figure out what the hell it does.”"
While I agree with you that this is likely AI assisted, I think this may be changing now.
People speak in the manner of what they consume. If you consume a lot of claudish, you will eventually start talking claudish too. And I've already noticed people talking claudish in real life.
I had to check and he does not seem to have the real qualifications to make his comments. In particular, he did computational neuro, not bioinformatics, and I can't find publications to support his claim.
The confidence score is not trivial to compute. That is the whole point of the model. Even if you are using a proper scoring function such as NLL, it is not enough to ensure calibration in deep nets. So you have to do good post training to ensure it. These are all known techniques, but they are far from trivial, especially on large scale datasets.
Their docs at https://docs.typesafe.ai/confidence state "confidence is a statistic computed from the probability distribution the answer already gives you. TypeSafe computes it for you"
And further down "TypeSafe computes confidence from how the probability is spread across the options. All of it on one option gives 1.0; the more evenly it spreads, the lower the confidence. This demo uses (3 × largest probability − 1) / 2 to approximate confidence for three options."
So while we don't know the exact formula they use, it is just a function over the probabilities
I am open to the argument that this does not work well if you just plug in a qwen model instead of a model that is trained to output more statistically useful token distributions
we agree then, that is the entirety of my argument. Getting a deep net especially one that is anywhere near even SLM size to be calibrated is tough, especially across domains. They claim calibration across a variety of datasets which is interesting.
Well, running an LLM X amount of times does give you better results provided you are willing to select the best one out of the X yourself.
But I agree with your general point. One of the reasons subscription plans are cheaper because they modulate usage in this way based on demand. They can also recover compute more coarsely via usage resets (which give positive PR).
Well yeah. If for some task you find it easier to just do it yourself then you should. But you can improve the situation even if you can't entirely automate it by automating parts of the verification thereby making it easier to human-do larger verifications when X>1. But in many cases even that is not possible.
The progress however is such that the number of tasks that you can do with >p% automated and X=1 keeps increasing. So many times just waiting works. Of course, here also it changes from field to field. There are some tasks at which AI hasn't even gotten started, others where it has already peaked, others where it's increasing slowly, and others where it's increasing fast.
Getting training data that works well for calibrated classification objectives is difficult.
I hear conflicting opinions (including my own) about how well calibrated each of these are. Jev seems to be the best.
But the jev release made obvious the PMF for these models, and the underlying reality is that calibration really doesn't matter much when you're replacing usecases where people were using damn LM head softmax probabilities before, which are nowhere near calibrated.
So now everyone simply finetunes qwen and makes a compared-to-regular-LLM vastly cheaper decision model. And it works for majority of usecases. People mostly only care about accuracy, not confidence.
reply