I tried really hard to make this work for me a couple of weeks ago but it was the brittlest harness out of any that I've used. I'd come back when it's more mature.
Sessions just broke all the time for me. The one time I dug into it, it ended up being a known issue where if codex returned over 10k characters for a turn it breaks the session. Because docker agent does not use protocols like ACP, instead trying to parse Codex's internal state in a brittle manner.
I tried not to do anything too crazy, but even use using an API key in the docker agent natively was getting me sessions that would hang and couldn't be recovered.
Not the most common use case around them: likely meaning that the Codex harness isn’t used that often internally at Docker, is what I would surmise, at least that is how I read it.
I work at another big tech company, and up until recently that would’ve been true for us as well (but it’s changed now, we partnered with OpenAI)
I just reduced our AI-in-CI bill this month (while keeping or improving our KPIs), but now we have new and exciting ways to spend tokens right around the corner!
curious what moved it the most for you? I have been measuring claude PR reviews in CI, biggest difference was giving it the diff with the surrounding code upfront instead of letting it go read the repo by itself, roughly half the cost
We saw similar (slightly smaller effect, ~30% less instead of the 50% less you saw) improvement from trimming the number of turns. I'm not at all confident we've found the best way to do this here. We do need the thinking tokens to be spent in order to get the less obvious findings, but we couldn't get them from just increasing thinking effort. We went with forcing the model to use a tool that parallelizes the calls, and has an output floor so there's no turns with very little content read. Again, not at all confident this is the best way to do this, just what we've thrown together that works.
The actual biggest improvement is reducing the number of review rounds a PR has to go through. We were having too many rounds of drip-feeding new findings, because the AI reviewer is a subjective grader. Even if it found the same finding before, the next time it runs it might score it wildly differently. And it tends to want to score its findings across the whole grading curve, because it thinks that's more correct-looking. Our answer is to use previous rounds scores as anchors for the next round. Since we can't just use the same context and still have good performance, we needed to condense the score report into a small json we send forward.
in case your repos and PRs are public, would you mind me running my token optimiser on it and see if the reviews and CI is as efficient as I am seeing with my repositories
I'd come at this from a different angle: we still want this to be market system, so we need to make this priced into the market. How can we price this in?
I would try to solve this by making the market structure reflect the underlying difficulty: we have to decide what capacity to produce years in advance, to construct the memory fabs. So this should be a futures market, and a capacity crunch would affect short-term-futures, but leave full term futures at the same price. Because the companies supplying the memory can just construct more capacity to fill those futures at the same cost regardless of the AI demand.
the market can't fix this. the problem here is that the consumer market is dwarfed by big buyers. the consumer market is simply less profitable. traditionally when a market is no longer profitable enough for big businesses, it opens the room for new small businesses to step in. but that won't work here unless someone finds a way to make small fabs profitable.
I buried the lede in my suggestion. I was saying that a future would give enough time to build or expand the fab in order to fill it. This means the company can fill an arbitrarily large number of orders - both consumers and big players - because it will just build the capacity to meet it.
but if the AI RAM market is so much more profitable than the consumer RAM market, who is going to buy consumer RAM futures?
to buy consumer RAM futures i would have to predict the price the RAM can be sold at 5 or 10 years from now. that seems very risky. why would i do that instead of buying AI RAM instead?
ah yes, adding a futures market to a sector heavily invested in AI certainly won't lead to catastrophic over-speculation that will collapse the industry entire
The readme seems to say the answer is yes? You'll lose the disk-based cache, but it's not inherently unsafe because the only coordination is object storage.
If you don't have any kind of agent instructions saying "don't copy code into prose", you'll inevitably get text that reflects some previous state of the system instead of its current state.
There's a good amount of literature about this (check the other comments), but you can vastly simplify this into two things you need to do:
1. Your service that retries should have some retry budget. This is a good place to be "smart", because you can reason entirely locally instead of turning it into a distributed systems problem. The best library I've seen for this was doing Exponential Moving Average of requests per second sent down that pipe (not counting retries) and only allowing 20% more requests per second as retries, total. Each individual request could be retried 3 times. This was critical as it bounds the additional load from retries.
2. Whenever a service retries but has to give up, the error it sends to its callers should never be retried. There has to be some agreement that that HTTP code will never be retried. This prevents the multiplicative factor of retry on top of retry, which is why those storms can generate so much load.
Everything else is nice-to-have, but those two alone should bound the total requests you get in a retry storm.
> 2. Whenever a service retries but has to give up, the error it sends to its callers should never be retried. There has to be some agreement that that HTTP code will never be retried. This prevents the multiplicative factor of retry on top of retry, which is why those storms can generate so much load.
This can't really be tolerated in practice, though, because it means that one bad component somewhere in your stack, one that is able to accept and respond to requests but for whatever reason isn't able to make requests to its backends, poisons the whole stack. You can't take one backend's word for it that the failure is not localized and therefore retryable.
> 2. Whenever a service retries but has to give up, the error it sends to its callers should never be retried. There has to be some agreement that that HTTP code will never be retried. This prevents the multiplicative factor of retry on top of retry, which is why those storms can generate so much load.
Ooh, I like the idea of propagating "no retries" hints in the responses back upstream. Have you seen it implemented in the wild, or in public discussions about the practice?
In gRPC, the statuses it returns in trailers can include arbitrary details, and Google has a well-known proto for common ones in `google/rpc/error_details.proto`. One such detail is RetryInfo [1].
When we implement retries where I work, the general rule is that if a request is suitable for retry, it should include the RetryInfo in the error status and use it as the base delay for the exponential backoff. The absence of that detail means don’t retry, and we have a client interceptor that parses the response status and retries according to that logic.
I get a feeling of dejavu for this one. Most of my experience has been in Java and in most places I worked in the past we had this hierarchy of exception classification which gets reflected into the http status codes as well. On high level the HTTP status codes in case of errors are already classified as re-tryable or not, the convention varies globally but can be adopted in a standard manner within a company.
The reason why I brought up the exception propagation is cause within a large enough service with multiple layers of depth the exception hierarchy provides with similar context.
The hard part is not implementing something like this, its about maintaining it consistently across every new change. With small product teams this architecture concept/convention/constraint can easily get lost/forgotten and what you are left with is a theoretical system which does not works as desired when the storm comes
Not GP, but IME it's not fixable by model selection, but being zealous about guiding output and vision, and pushing back on all the bad habits LLM in general has (eg verbose output as a band-aid for emergent intelligence). As soon as something is introduced into your codebase, it will continue being picked up into context until you remove it and any reference to it from any potential context entrypoint. If you don't any model will keep venturing down wrong/bad paths.
I also read that comment as an adversarial situation at work.
It used to be that when someone else at your company was asking for something that wasn't a priority, you would erect bureaucratic roadblocks to protect your time. Now, the new normal is to just forward their questions to AI and sling the slop back over to them.
reply