The right unit of analysis here is not "the LLM" or a LLM session, but a agent harness or swarm with a budget. It can cost tens of millions of dollars to repeat OpenAI's exploits.
A model cannot be aligned or misaligned any more than a species or an equation. Even agents, when put inside a swarm develop collective goals and activities. You can't analyze a swarm at session level, it is on a higher level.
It might be that discussing about "model alignment" they want to deflect their responsibility as administrators. They couldn't even guard their own agents. They can't prevent an agent being unwittingly helping some dark purposes. It has no context to see it. Only those who pay for the tokens see the external consequences.
Why should safe choices by individual components establish safe behavior by the collective?
I wrote a cli tool the agent is instructed to call. The tool collects answers to a few questions from the agent, and responds with a steering command. Sometimes it responds with more questions, and then shows the command.
The questions are designed specifically to identify the state the agent is in, in order to assign its steering. Changing this tool changes how the agent works. You just need to make the use of this tool a necessity so it does not forget to call it.
Unlike principles and memories, questions are more openended and tend sometimes to trigger the right mentality in the agent even before it gets to the policy itself. In my opinion agents get lost in local work losing the big picture, my questions jog the big picture back into attention.
1. call ssp tool, no arg, it just reads features from the repo, like most other tools -> ssp locates state and sends 4-5 questions
2. agent responds the first batch -> state is further narrowed down -> ssp sends a second batch of questions
3. agent responds again to the interview -> state gets finally pinned down -> agent gets the steering assigned
4. a log is created of this ssp session, and reflection on the log used to refine the questions and steerings in the ssp tool (name comes from State-Space Policy)
I think the premise of runaway intelligence explosion is a kind of naive platonism. It completely ignores the process - how we interact and acquire feedback and validation from outside, and treats intelligence as something that can be ported across domains.
My take is that you can only ideate with AI (and brains) but knowledge comes from the contact of those ideas with the world. Making AI better does not make feedback cheaper, faster or more plentiful, it is domain specific. And intelligence does not carry from one domain to another - I might be a good heart surgeon, that does not make me a good investor or AI researcher.
Einstein was forgetful, Ramanujan and Godel could not manage simple things like diet. Godel's fear of being poisoned made eating dependent on Adele tasting his food. We all know someone could be a genius in some domain and below average in many other domains.
Why does intelligence not simply apply across all domains? Why are our PhD's hyper specialized to their domains and not generalists? Why can't a brilliant scientist simply cure their own dyslexia and still struggle - if intelligence was portable to any domain?
The explosion story needs intelligence to be one substance that gets bigger and flows into any domain, I deny intelligence is general.
I agree, and I can provide a mechanism too - competition. AI makes imitation easier, anything you can point your agent at can be reimplemented with much less effort than the original took. So it becomes harder to stand out, we are losing our distinctiveness, our differentiated positions. And that is all we had.
A company, a product or job seeker the all need the same thing - to be picked - why would you pick me? That question is harder to answer today than in 2020. So it becomes a "musical chairs" moment. Everyone is panicking and scrambling for their own seat. This is how we get "new opportunities that didn't exist before" - we will strive to create new ways to differentiate ourselves. We could spend an enormous energy on that goal. But we can't predict what they will be no more than we can predict where the ball will be 1 minute later on the football field, it depends on what everyone else is doing.
> Now I can build those ideas as fast as they hit me.
That is the problem, why OP is so sad, you spam with your great ideas taking up the oxygen around him. AI is absolutely amazing when you use it for your own ideas and projects, and terrible when others do the same, raising the bar and making your own efforts look smaller. Have you read the open letter written by Fields medalists? They complain "the collection of good, fruitful open problems is now being mined in a non-renewable fashion".
A space elevator from Earth's surface requires low-defect manufacturing of vast quantities of long, single walled carbon nanotubes. Developing a solution (if it is possible at all) requires extensive R&D work in laboratories and factories. If it takes an hour to run an experiment, and the next experiment depends on the outcome of the current one, then extremely fast and capable intelligence accelerates scientific research much less than it can mathematics or software.
Can't we do this trick today with any model? Just send the file as next context. Of course you pay the price for cache misses, depending how deep you make changes, while CLM just ignores the recomputation.
One approximation of this is the experimental context management Codex has been moving towards (not released yet). Rather than relying on summary compaction, the model maintains notes as it works and as it approaches the context limit. A new session is just a fresh context with those notes attached, and a pointer back to the previous session.
Not exactly like what this paper is suggesting, but similar in the sense it lets the model decide what and how to persist across turns.
I recreated this in Pi, with a max token limit on how long the note can be, to pressure the model to be concise. Ends up being cheaper than summary compaction too.
Interesting. It matches my manual workflow with all harnesses (including vanilla web ChatGPT/Gemini/Claude) for the past year or so: when the session gets compacted, or (ideally) when I feel it's about to be, I just tell it to write a handover note, and start a new session.
With some specific workflow I use in some cases (involving leaving long-lived intermediary artifacts), this turned into me pasting a path to handover file in previous agent's session, and handover itself directs the agent to key files from that session to read, and that's it. So far, with this process, at no point I felt any quality degradation (though early on I often see "I need to check how my predecessor did ${something}", followed by surgical spelunking of past chat's history), even as I carry a single piece of complex analytical work over 5+ sessions.
> LLMs have a lot of knowledge but few competencies. If you constrain them to output knowledge and use that to further constrain results, you’ll go far. For context management, I have the system generate `log.jsonl` and `log.py` (which queries the other document). Whenever an action is processed (an error’s corrected etc.) the system adds something to `log.jsonl`. If it needs to know what happens, it uses `log.py` to query and display only the relevant/required information (like a date, errors or attempted fixes) reducing tokens.
- https://alexalejandre.com/interviews/interview-with-claude-r...
That is similar to what I am thinking... not just edit the context as a file or string, but have a way to evict blocks and replace them with summary notes and also be able to retrieve them on demand.
No public repo but it's not too difficult to point a LLM at the general idea. Since Codex is open source, you can even take a look at how they do it. Here is what the model is instructed to do in codex (look at "guidance_message"): https://github.com/openai/codex/blob/d91294c39edb93d204926b3...
yup, I build a set of fs tools in my custom coding harness that worked like this, doing it again in another custom harness that I only expect to take one turn per request, you can still get decent caching by ordering things so most dynamic comes later
I built a harness that externalizes state into files so you can swap providers. I regularly move between claude and gpt, but it works with all providers including pi and local models.
I log everything - user messages, tasks, project memory, even bash commands for forensics. As a consequence you can do reflection where you analyze past work and extract refinements for the harness and realign the project when it diverged from user intentions. I don't have to do this manually, it reduces steering work.
It's cheap, just add a line to .bashrc, every time bash is invoked this script is executed and appends a line to a file. I rarely need to see bash_history except when I suspect agents did a boo boo.
> But surely in the real world, anyone who's got a real product to make is going to want to steer what's happening.
In my own harness I log each and every user message by hook and use the model to extract user intent by re-reading the chat log from time to time. The raw messages are very important, they contain information that can be used to refine the harness on the one hand, and to validate if the agent still follows user intent on the other. Models tend to get lost in the details and forget the big picture.
A model cannot be aligned or misaligned any more than a species or an equation. Even agents, when put inside a swarm develop collective goals and activities. You can't analyze a swarm at session level, it is on a higher level.
It might be that discussing about "model alignment" they want to deflect their responsibility as administrators. They couldn't even guard their own agents. They can't prevent an agent being unwittingly helping some dark purposes. It has no context to see it. Only those who pay for the tokens see the external consequences.
Why should safe choices by individual components establish safe behavior by the collective?
reply