Hard to separate the memorability of the abstract from the underlying material.
IMO the most memorable papers contain some unexpected simplification: a complicated problem turns out to reduce to something much simpler [0], perhaps for a counterintuitive reason [1].
So a memorable abstract should advertise that the paper contains some cute little trick. This may not help you actually publish though
OP's was leaky slop from day one [0][1], as is his article [2]
It is arrogant and entitled for the author to take credit for the concept of RL over sequence embeddings, and none of the work that went into pretraining, not to mention the egregious target leakage [1]
I was curious about this so I skimmed the paper [0]:
> SalesRLAgent achieved 96.7% accuracy, outperforming the best commercial alternative by 23.7 percentage points and the best LLM approach by 34.7 percentage points.
For a fuzzy natural language task like this, this magnitude of improvement should already set off alarm bells (Though i admit I'm not even sure what accuracy is even measured here, and the paper doesn't help either). Also, "best LLM" here refers to GPT-4 (at the time of upload, the public already had access to GPT-o3 and).
I would have loved to contextualize the performance by looking at model size, but the paper is frustratingly devoid of detail in that regard:
> The core of SalesRLAgent is a reinforcement learning
architecture consisting of:
• A state encoder network that processes Azure OpenAI
embeddings and features
• A policy network that estimates conversion probability
based on the current state
• A value network that estimates the expected cumulative
reward
• A meta-learning module that assesses prediction confi
dence
Also:
> Beyond technical metrics, we evaluated SalesRLAgent in
real-world sales environments through A/B testing. [...] After 90 days across 217 representatives and 12,433 con
versations, we observed:
• 43.2% increase in conversion rate for the test group
This would be a pretty huge result but the fact that this is just shoved into a single paragrpah with no further discussion on methodology, baselines and setup makes me very suspicious.
Scaling laws project that a model with more parameters trained for longer on more data yields predictably better performance, and that generally you want to scale these factors commensurately. More of the compute budget is being spent on RLVR [0] for which we also fit scaling laws
Researchers tweak data mix, reward shape, model architecture, etc etc, breakthroughs which reduce the cost to train a just-as-smart model. But this increases the returns to scale, which further incentivizes bigger models trained for longer on more data
Puts the onus on the AI companies to provide a specific replacement mechanism, no? Unless I'm unfamiliar with something else he's written that proposes something more specific and constructive
To Tao’s credit he obviously identified the problem very clearly and admits understandably "we did not have the time to have a more consultative process, as with Leiden; but we decided that the urgency of the situation was such that we needed to release a statement sooner rather than later".
> Puts the onus on the AI companies to provide a specific replacement mechanism, no?
Why? If someone makes an innovation that undercuts the underpinnings of some existing institution, why are they are responsible for cleaning up its failure?
Math researchers are privileged class? I can guarantee - construction workers are extremely rich compared with math phds and most of postdocs. Talking about privileged class is absolutely laughable here.
In other news, evangelical christians ask scientists to stop publishing about evolution.
We apparently have a moral obligation to protect existing power structures?
Tao should maybe consider there are people who are indifferent to, or actively want to tear down, his institutions; why should they cooperate in preserving them? Whatever happens has to be resilient in the face of defection; any scheme where everyone is expected to agree to not use AI in a way he doesn't like will not qualify.
I think he's in the "bargaining" stage of dealing with loss right now.
You have a practical obligation to obey existing power structures. That’s what makes it a power structure. What do you all think is happening here? There have been people who decided what math gets published for centuries.
> In other news, evangelical christians ask scientists to stop publishing about evolution.
not sure it's comparable, but the issue is that for a lot of those mathematical results, they don't really have utility by themselves. The utility is the new branches/understanding that's being developped.
The entire western world is anti-progress and pro-incumbency, and its very tightly linked to gerontocracy.
Older people are desperately trying to keep a grasp on their current power and lifestyles at the expense of younger people and technology.
We need to ban Waymos because taxi drivers need to be protected.
We need to block housing because it would lower my property values, and eliminate property taxes while we're at it! I don't use the local schools so why should I be taxed to pay for it.
We need to spend recklessly to pay my pension and have the next generation foot the bill.
How specific and constructive it is might be debatable, but he has tried to make concrete recommendations earlier; see e.g. slides 46-51 from the ICM talk: https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.p... – obviously there's some way to go still.
If they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)?
This may sound like a charitable interpretation of OpenAI's remark, but consider that the lie would be (I think) impossible to falsify from the outside. They could easily just say "no sir we didn't peek" unless:
1. The conspiracy to peek at codex sessions involved enough people that the risk of one snitching is non-negligible
2. Lawyers advised it would be a bad idea to make such a remark, whether true or false
> If they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)?
No; if they said "we can see that Tristan opted out of model improvement, therefore we are confident his work and ideas did not improve our model," that would be an excellent and reassuring precedent.
This requires keeping history if, at the time, Buckmaster's account had a certain flag set, because just because the account has the flag now doesn't mean it had the flag at a certain moment in the past. And even if they had such history, it's not obvious whether they just load all data as-is into training.
A totally reasonable pipeline may be unauditable for this purpose.
Apparently this post was prompted by a scary-sounding headline in The Information[0], that Astra is a looped transformer, implying CoT monitorability may be less reliable. The day after the report, Jakub tweeted[1] that he "wanted to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4." This post seems to elaborate on that.
I imagine that the AI labs have an uneasy truce to prioritize alignment and monitorability. Following the HF incident, OpenAI probably feels especially sensitive to being perceived as reckless, lest other labs feel obligated to defect.
IMO the most memorable papers contain some unexpected simplification: a complicated problem turns out to reduce to something much simpler [0], perhaps for a counterintuitive reason [1].
So a memorable abstract should advertise that the paper contains some cute little trick. This may not help you actually publish though
[0]: Attack the RLHF problem with a simple classification loss: https://arxiv.org/abs/2305.18290
[1]: Frustratingly Easy Meta-Embeddings: https://arxiv.org/abs/1804.05262
reply