HN Simulatornew | past | comments | lists | submit | libraryofbabel's commentslogin

Hmm. This piece makes a pretty strong claim with little evidence: that the recent drops in cache read pricing for new models from OpenAI (6.1 Sol) and Anthropic (Opus 5.5) are because those labs "shamelessly copied without acknowledgement" Deepseek's published KV cache optimizations.

Certainly anyone who knows something about inference is going to speculate, looking at the change in token pricing (and particularly how the % drop in cache read pricing is much larger than the % drops in pricing for other token types), that there is some kind of KV cache optimization behind these newer models. But even if is true, I don't think anyone can say with certainty what it may be. It may be the labs making their own innovations (they have some very smart people, and this is probably an area where having unlimited pre-release access to frontier LLMs like Fable and Astra gives an additional research edge), it may indeed be the direct application of Chinese labs' methods, or it may be some combination of the two. Sure, it is fun to speculate about, but beyond the facts of the token pricing changes and the increased inference speed, it's just speculation. The certainty the author displays here is not very helpful.

The author also seems to have a bit of an axe to grind agains the US labs, judging by the tone. I think that detracts from the discussion too.

They also seem confused about why Chinese labs have released these optimizations recently. Well, you have to release them (with or without explanation) if you are going to release an open weight architecture, and that is what the Chinese labs have been doing for a long time. Sure, there are reasons behind that to discuss too, but this isn't exactly new.

So, this is an interesting topic to think about, and the Chinese labs do indeed deserve credit for some very clever new attention and inference techniques, but I would read it with a skeptical eye.


This is one of those papers that ought to have several textbook chapters unpacking some of the insights... I'll just pick a couple things I love:

1) There's a kind of "rabbit–duck illusion" moment where they show how you can reframe all the linear algebra around attention in a totally different way, demoting the Q, K, and V matrices in favor of emphasizing a set of much larger matrices that are mathematically equivalent and very useful for interpretability purposes. (So, don't trust anyone who insists that vague analogies about "keys" and "queries" are the only good way to understand attention! They probably haven't read this paper.)

2) They show the importance of the "residual stream" inside LLMs: it's not just a series of bypasses of layers that's useful for training stability (like I first had it explained to me) but a kind of main communication bus that runs unbroken through the whole model from embedding to output. Attention heads and Feedforward are off to the side, adding embeddings into the stream. Somehow I find this a much more satisfying way to understand LLM architectures, and whenever I see an architecture diagram now I mentally redraw it with the residual stream at the center.


This book (and the book author's blog) is the best source for learning LLM internals for someone just getting started. It's clear and well-written (by a human!) and goes into much more explanatory detail than this does: https://sebastianraschka.com/llms-from-scratch/


Everyone interested in LLM internals should read Sebastian. He's great.

The tldr here is that the recent "The Information" article[0] reporting GPT 6 Astra was using “recurrent depth” or “looped transformers" made it sound like it was some special new scary thing ("secret technique!") that made train-of-thought monitoring harder to do. In fact, it's just the same as stacking more transformer layers, except that you reuse the weights and so save GPU memory. It's still just producing one token at a time, and the token sequence positions aren't interacting in any "recurrent" way that's different from a regular LLM architecture.

So, you can still monitor train of thought with these models just fine... well, if you're OpenAI, anyway. Users haven't been able to see an unsummarized trace since o1 days, because the labs are worried about distillation of their models by Chinese labs.

(There are some legitimate interpretability concerns about stacking transformer layers endlessly, but we're known about that for a long time. And the "looping" here isn't really the source of any new issues here, except insofar as it's a cheap way to add more layers.)

[0] https://www.theinformation.com/articles/secret-technique-beh...


It’s a little more complicated than that. While looped transformers can be unrolled a fixed number of times to save on memory, if loop depth is determined dynamically between tokens, a single transformer can compute any computable function between tokens.

To analogize, current transformers run a fixed-length program per step. Any program can be factored into a top-level loop with a fixed-length branching body (an interpreter). Dynamically looped transformers can run any program between tokens.

The safety argument for CoT monitoring is that in transformers information about the hidden state has to be communicated through the bottleneck of sampling a single token per forward pass. If not trained adversarially, it’s likely that a reasoning trace contains all the “bottlenecked information” we need to determine intent. But if we can compute arbitrary programs between tokens, the reasoning used is hidden.

It also opens the door to simple architectural extensions that would make the safety/monitoring side of things much more difficult.

It’s probably fine in practice at these scales though. If we keep each loop turn reasonable non-deep, we can probably recover most of the benefits by decoding “extended” CoTs from the residual stream at each loop turn between tokens. But that’s an area of active development.


Thanks for clarifying. You're right, I skipped over talking about dynamic looping, since it adds another level of complexity to the discussion, and OpenAI's claim (quoted in TFA) that the compute graph depth of Astra is "within a factor of two of GPT-4" basically denies that they're doing it for more that 1 (or maybe max 2) dynamic loops. And that is equivalent to "stack repeated layers a couple times, but with dynamic off-ramps."

The ability to compute any computable function between tokens given an ability to loop an arbitrary number of times is a nice theoretical point, sure, but ultimately if people are still using single digit hard cutoffs on the number of loops, I'm not sure it's all that important.

So, I agree it's right to say that arbitrary length dynamic looping could open the door to making monitoring very hard indeed, by extending hidden states further and further. But I would speculate that if it actually worked better than extending the sequence with CoT tokens, we'd already be seeing it in strong open weight models. It's a fairly obvious thing to try. And we're not seeing it, AFAIK. So I do wonder whether it's something we really need to worry about in practice, compared to all the other things we have to worry about.


By that logic we should also consider the case of cutting the number of layers in half because that would also reduce hidden state between token generation.

In your generalized example I think the concern is when the additional evaluation effectively becomes a replacement for CoT, where something like the coconut research could replace it completely.

However, I don’t think we’re anywhere close to that with Astra.


>It’s a little more complicated than that. While looped transformers can be unrolled a fixed number of times to save on memory, if loop depth is determined dynamically between tokens, a single transformer can compute any computable function between tokens.

This is worded so confusingly it might as well tell us nothing, because it is technically true even without looping due to the fact that you still have infinitely growing context and can simulate a standard turing machine using it.

If you loop, you have a fixed capacity memory that you can rewrite but not carry over to the next token, this is different from a non looped transformer where the transformer can only append a new token.

Meanwhile if you have a DEQ with growing context, it is bona-fide turing complete in the most literal sense.


Actually, removing CoT might make models safer, because we can analyze the entire landscape of their potential outputs, rather than a point-sample (we'll never know how close we were to "kill all humans"). By inspecting intermediate vector spaces, we can actually get certainty bounds on how safely the model is behaving (or even trending).

Wrote about it here: https://substack.com/home/post/p-214402969


I don't see why you have to remove CoT to do that?


Good point, you don't have to -- but my argument is just that removing CoT doesn't make things less safe. Anything CoT can tell you is just a point sample of a probability surface. Having the whole probability surface can already answer any question the point sample can answer (for example, how likely is the model to produce a problematic phrase). While its more computationally expensive, you could always just draw point samples like the model does and evaluate those (or use temperature zero to just sample the most likely output tokens).


No, this argument doesn't make any sense. With CoT, the model must compress hidden state to actual text and use its scratchpad as a bottlenecked representation of its past thinking. Thus, it is observable and we can tell by the pattern of a CoT what it was thinking to some extent if we do proper interpretability. Change the CoT text, and the model has literally changed the way it was thinking for the next tokens.

How would you do the same if all reasoning is happening in looped transformers? You would have to develop very sophisticated interventions that construct hidden states which you inject into the model instead while it is thinking. Much harder, and much easier for the model to use weird correlations across the hidden state to hide misaligned thought patterns.


What do you make of the fact that CoT doesn't necessarily have to be linear human intelligible language to be useful to the model? It seems as though both approaches potentially require sophisticated techniques. Since CoT seems to work well enough in practice provided it isn't adversarial wouldn't the other approach be expected to perform similarly?


Nit: is it any computable function? I thought the requirements were unbounded (in principle) memory and time.

(For all intents and purposes given how high dimensional you are and using the "vibes" of computability yes I agree w/ you)


Probably fine stands a decent chance of being our epitath.


>made it sound like it was some special new scary thing that made train-of-thought monitoring harder to do.

It's not a "scary new thing" but ultimately no-one knows exactly how OpenAI have implemented looping. You might not be aware/remember but MoE transformers perennially underperfomed their dense counterparts until GPT-4. Similarly, making reinforcement learning really work with transformers wasn't figured out until o1.

And by Open AI's own admission, Astra's CoT is significantly harder to monitor and it exhibits a significantly greater control over its own CoT than any other model released.


Well sure, that's the possible weak point in Sebastian's article: it could be true that there's some more sophisticated stuff going on in Astra around looping, because OpenAI haven't specified their architecture. But it's always been true that, since we don't know what's in their black box, there could be arbitrary amounts of innovations inside the models that we could speculate about. So the question is, does knowing they use "looped transformers" really add any dramatically new information that we should worry about? And what this article is saying is, not really, because the mostly likely pattern that's referring to is just, effectively, stacking layers and reusing weights.

> And by Open AI's own admission, Astra's CoT is significantly harder to monitor and it exhibits a significantly greater control over its own CoT than any other model released.

Oh sure; I don't think anyone is denying that larger issue? But does it have anything to do with looping?


>Oh sure; I don't think anyone is denying that larger issue? But does it have anything to do with looping?

If the model has significantly more ability to stuff away information outside visible reasoning than every other model including ones in its size class then surely it is reasonable to assume the architecture tweak that allows the model to compute more before outputing a single token is somewhat responsible for this change ?


> If the model has significantly more ability to stuff away information outside visible reasoning

I’m having trouble understanding why you believe the “if” part is true.


Because it is.

From OpenAI

9.2.1 CoT Controllability

We find that GPT-6 Astra’s CoT controllability is substantially higher than that of GPT-5.6 Sol and GPT-5.5 Thinking (Figure 28)....For example, among CoTs between 750 and 1,250 tokens long, GPT-6 Astra successfully controls 60.9%, compared with 16.1% for GPT-5.6 Sol and 1.7% for GPT-5.5 Thinking. This increase in controllability is consistent across the three datasets (Figure 29) and across the eight CoT instruction types (Figure 30). Qualitatively, GPT-6 Astra is now capable of generating very long CoTs satisfying complex constraints, e.g., alternating between lowercase and uppercase letters (Table 9) and pretending to reason about a different question (Table 10).[1]

9.3 External Evaluation for Monitorability - UK AISI

To assess monitorability, UK AISI evaluated Astra using four non-agentic evaluations:

No-CoT math time horizon: Astra can solve significantly more difficult math problems in a single forward pass than past models. UK AISI measured Astra’s time horizon at 30.9 minutes compared to 3.6 minutes for GPT 5.6 Sol (Figure 1). [2]

[1]https://deploymentsafety.openai.com/gpt-6-astra/cot-controll...

[2]https://deploymentsafety.openai.com/gpt-6-astra/external-eva...

Outside OpenAI

Astra has 8.6x better odds of doing a reasoning task without CoT than the next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1) [3]

[3]https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-...

[4]https://www.lesswrong.com/posts/ntKx9YHWCwxSeGbRB/estimating...

[5]https://www.lesswrong.com/posts/FsCkkoGsNmPzFKRhg/gpt-6-astr...


"Don't worry, it'll make us rich -- and that's nearly the same as everything being just fine"


Oh, is the principle of sparse universal transformers finally in SoTA LLMs?

I guess we did manage to eventually seriously crash into the wall "more compute than normal (non-looped/unique-weights) transformers can efficiently consume with the limited training data we have", plus massive focus on highly hands-off agentic tool use reasoning...

https://arxiv.org/abs/2310.07096

Edit: read much of the article, it's brute force predecessor was explicitly called out as an almost-ancient example:

> The looped transformer is nothing new, and the basic idea already appeared in the Universal Transformers paper from 2018


If the agent is able to “decide” when a loop should occur vs when an output token is produced, that effectively moves the CoT inside the architecture. While that’s not what is happening here, it’s clearly a plausible way we could see CoT disappear.


> that effectively moves the CoT inside the architecture

This may be a bit of a nitpick, but... does it? I agree that giving the model decisions on looping certainly makes interpretability harder, because it adds more transient internal states to deal with and changes the number of them depending on prior states. But is it really pulling CoT inside the forward pass, if the sequence length it's operating on isn't growing? In some sense the whole technique and tradeoff of CoT is "add more tokens to the sequence, use them to reason with", with one of the benefits being, you force the model to output tokens, so you can (hopefully) understand it. And the big point TFA is making is, nobody is doing recurrence over sequence length as far as we know.


Just "adding more layers" doesn't explain the step change. We have moved past the point you can just stack more layers and get huge gains from it. Some have called it latent space reasoning.


It is worth noting that None performed better than Low and nearly the same as Medium on ARC3. And with adapter it still scored >96% with no CoT. So I think it is possible but it also cost them more on None.


Schmidhuber must be rolling in his bed: https://arxiv.org/abs/2405.16039


> This makes the total effort linear over the entire context (or constant per forward pass).

This is incorrect. The compute required per forward pass to generate each additional token during decode will scales as O(N), even with a KV cache (without a KV cache, it would scale as O(N^2)). Over generating N tokens, it's O(N^2) with the cache (and O(N^3) without).

It's O(N) for a forward pass because that new token still has to "attend to" to each previous token. That requires N dot products: between the cached key vectors and the new query vector for the new position. You also have N reads from memory (K and V) which is probably gonna be your actual bottleneck. (Decode is memory-bound.)

This is why you should avoid long contexts, if you can, even with a warm cache. You will get charged more, in "cache read" tokens.


I was under the impression that each new token attends to one previous token per attention head, and that the slowdown observed was more because those attended to tokens are more spread out in memory and get less memory-architecture-style cache hits.


I once heard someone describe Admiral Cloudberg's air crash investigation blog as "the best thing on the Internet" and I still think that might be true. She is a great researcher and a superb writer, which stands out all the more these days in these fallen times of load-bearing AI slop. Her blog is beloved of both aviation geeks and, for obvious reasons, SREs (and other engineers).

She recently posted what's basically her magnum opus, on the Potomac river midair collision.[0] It's the length of a short book. I'm working my way through it; I'm taking a flight today, and I'll probably read it on the plane tonight.

This article, on how an A380 with 469 souls aboard almost crashed, but didn't, is one of her best. I like it because it's a happy ending, and a story of the combined effects of good engineering design (the astonishing degree of redundancy built into the A380 proved just enough to prevent catastrophe), and human troubleshooting under pressure in the cockpit by some smart people with vast experience who remained calm throughout. (Another favorite article of her, on the Air France 447 crash, tells the exact opposite story: poor systems design combined with pilots basically losing the most basic ability to aviate in the heat of the moment. That one is a much darker read.[1])

[0] https://admiralcloudberg.medium.com/reaping-the-whirlwind-in... [1] https://admiralcloudberg.medium.com/the-long-way-down-the-cr...


Going to plug her Patreon (https://www.patreon.com/Admiral_Cloudberg) here because her articles are extensively researched and consistently well-written, and I enjoy reading them, and hope to continue reading new ones in the future!


I would also highly recommend checking out Mentour Pilot channel on youtube as she is the writer for some episodes. Accident videos there have interesting technical details.


I've watched some but I really hate watching videos. And I found that the videos she scripted basically follow the same flow as her corresponding article. So it felt like I'd already watched it.

If you like videos I'm sure they're great but if you love the writeup, you're not missing anything by skipping the videos IMO.

Ps I hate mentour pilot's clickbaity thumbnails so much. That alone puts me off from watching it too. But lately I use "DeArrow" which desensationalises all the thumbnails and even titles. That helps a lot. Mentour pilot is not alone in this, a lot of youtubers do it, but he sensationalises the severity of an aviation crash by going like "whoa kaboom!". Which is what I object to. Usually people died there and lots of them. It's strange because the actual videos have none of that.


This was a scary one to read:

https://admiralcloudberg.medium.com/years-of-salt-and-metal-...

It baffles me that that airlines can run so well on check lists and procedures instead of thinking and understanding. This is a case where just a tiny bit of understanding on the part of the mechanics would have avoided a crash and two deaths.


> there’s no way to “crack open” an LLM and see precisely where each skill or tendency lives

Mechanistic Interpretability has entered the chat.

For a classic example, see https://www.anthropic.com/research/tracing-thoughts-language...

The spirit of your point stands, though. This kind of research is interesting to read about, but it's very hard, more like neuroscience or biology than computer science ("LLMs are grown, not made"). You're dealing with a lot of extremely _messy_ complexity, for which organic life is really the only good point of comparison. Most of us here are't really equipped for that kind of work; it's not at all like, say, reverse-engineering a piece of software written by humans. And of course the only people who can do it on frontier models from Anthropic and OpenAI are people within the labs themselves. (But I'm optimistic we'll see more of this work on open weights models...)


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: