They group dozens to hundreds of designs on one chip. The end result is that even hobbyists can afford to have their small ASIC on real silicon. It's still expensive but it's within reach of a serious enthusiast with a steady job.
I wonder what would "compute" mean if cpus were more efficient at matrix multiplication say 15 years ago. And on the flipside, what it would take to say train a frontier model entirely on cpus in the future.
On the compute side there's the issue of scalability. CPUs are designed to perform a handful of operations at one time. They typically have a small number of dedicated integer, float, and other ALU configurations. Having dedicated matrix multiplication instructions would still lock that to how many matrix-capable ALUs there are in the CPU.
GPUs are designed to process a large number of calculations at once (so they can process triangles in 3D graphics). This makes them good at ML applications as they can process many of the matrix calculations at once. A 4090 has 16,384 CUDA cores (general compute ALUs) and 512 Tensor cores (dedicated matrix compute ALUs); a 5090 has 21,760 CUDA and 680 Tensor cores.
The other issue when training models (and running larger models) is the amount of VRAM (or RAM for CPUs) available. GPUs are limited in this aspect, whereas CPUs can have a lot higher memory. This affects things like batch size and the size of model that can be trained or fine-tuned.
https://unsloth.ai has guides for how to fine-tune existing models like Qwen 3.8 27B, memory requirements, etc.
https://medium.com/@kailaspsudheer/the-transformers-arithmet... has some information on training a base model. A 7B llama model is estimated at taking ~34GB memory for inference at F32, but was observed requiring 96GB memory when training (for the model weights, gradients, activations, and optimizer states).
Note: you can reduce the memory required for training by recomputing the gradients, at a cost of performance/time. You can also do other tricks like performing a QLoRA/LoRA pass on the model then merging that into the model to create a checkpoint.
I don't know what sized model you could train on 64GB/128GB RAM via a CPU.
The ideas survive in AVX-512 a.k.a. AVX10, not in AVX, which was a project parallel to Larrabee and resulting in an inferior ISA, which was adopted in the mainline Intel CPUs due to internal politics, not due to technical superiority.
When brought to the Intel server CPUs, the Larrabee New Instructions were rebranded as "AVX-512", despite having no relationship with the AVX ISA extension.
AVX was the creation of the Intel A-team, while the Larrabee New Instructions were designed by a C-level or D-level Intel team, but the latter have benefited from the contribution of a few consultants hired from outside Intel, who had experience in programming graphic applications.
AVX, which included only minimal and obvious improvements over SSE, i.e. double width and 3-address instructions, has slowed down considerably the improvement of the computational performance of CPUs in comparison with an alternate time line where Intel Sandy Bridge would have implemented a variant of the Larrabee New Instructions instead of AVX. This could have been done in a manner that would not have required any significant cost increase over the Sandy Bridge with AVX, because in AVX-512 it is not the width that is important but the architecture of the vector instruction set (e.g. with masked operations).
I wonder what "compute" would mean if CPUs were more efficient at matrix multiplication and vendors had the balls to pair each core to its own dedicated DDR and a star interconnect between.
> 3+3? 4+4? 5+6? 7+8?
> Reply ~18000 0 ~10 min 2 By : 1-1: I'm a beginner. 3x2 is my best option, but if you're not sure about the other options then just go for it and try again
GPT-2 is an interesting one because it is a February 2019 model. (You can see some information about it below the card if you click on the card.)
That was 2-3 years before the big "ChatGPT moment" (the highly coherent ChatGPT research preview was released in November 2022, I think it was ChatGPT 3.5). Back in 2019 the models really were not producing very coherent output. Now you can see it for yourself right in your browser :) Everything has come a really long way since then!
I've been playing with it for a bit, and I'm a tiny little bit surprised that they decided to continue pursuing that direction of research at all. What I'm getting from it looks like it could just be arbitrary sentence and paragraph fragments from the Internet, pasted together Markov-chain like.
I'm not sure I would have ever believed that something useful would come out of it, yet here we are.
Not sure how you're prompting it but remember that it's not trained for chat or instruction following, it simply takes the text given to it and tries to continue it. Give it the right prompt structure, and it can (at least sometimes) output coherent completions, far more often than you'd see in a Markov-chain. Also, this version is more or less equivalent to the smallest version of GPT-2; the largest version was 1.5 billion parameters and was much more likely to generate impressive (at the time) output.
The assumption that LLMs would always need sophisticated inputs to generate useful outputs is where the term "prompt engineering" came from. Now that idea is basically dead. Absolutely wild how far these models have come in less than a decade!
This leaves me two questions
1. Will we ever reach (or have we already) where we can more empirically measure the cost of the adding the nth agent. Like the way businesses of different sizes can know when the nth employee add diminishes returns overall, can we measure the front with agents depending on task scope, resources, and other factors.
2. On the other end, can we see point where we can reduce how many agents are needed down the minimal set? For example many chains have reduced staff down to the minimal number of employees (note I do not agree with this) and still the balance sheet is in the black. I think in the hyper optimization and efficiency society we're in I think this will also happen, although the reality of "I can do 100 agents work with 10" might not be in providers and labs best interest economically.
> If you believe these issues should not be raised because FLAWED was critical of OpenAI, please know that research is not sports. OpenAI is not the Spurs, 1Password is not the Knicks, and Off-by-1 Labs is not Jalen Brunson.
To be fair there's probably a considerable amount of engineering that went into evaluating those markdown files so the agent behaviour is statistically reliable. The markdown is the product, not the process
I do, I like to know how badly my agents' context is wasted and what unexpected side effects to watch for (like, "always start with ${clitool} --help" == always waste few hundred tokens when even touching the skill; or instructions asking it to do something that generalize into stupid thing in larger context).
If those skills were unreadable, however, that would imply proper engineering - like e.g. the skills themselves being an output of iterative RL over set of evals.
I don't think unreadable skills implies proper engineering at all. It's just as or more likely that they're the result of a blind iterative process with no clear improvement signal. (And whether iterative RL over a set of evals is actually proper engineering here is another question...)
> And whether iterative RL over a set of evals is actually proper engineering here is another question...
I'd put it like this: regardless of the merit of how they're applied, it would at least demonstrate possession of the advanced skills expected of experienced software engineers.
I had more success starting from scratch. Especially claude's skill-creator skill, it micromanages, which is in fact worse for 5+ models than just leaving the instructions out and crossing your fingers.
Start from scratch, do some test runs, find the bugs, add the minimal possible text to avoid the bug, iterate
You can get 95% of my impl workflow skill by just telling Claude "split the work into slices and use ephemeral subagents" and the other 5% takes like 10x as much text to achieve
Markdown can never guarantee deterministic agent operations. It is an influence on inference, not a deterministic code path. How "statistically reliable" is it?
Write a prompt, evaluate the prompt, understand that is succeeds 95% of the time.
Write a new prompt, evaluate, it now succeeds 99% of the time. Measure what changes between prompt #1 and prompt #2, understand what contributed to the performance jump.
Write a third prompt, this one succeeds 100% of the time. Increase the size of your evaluation set, find a 1/5000 error-class and a 1/10000 error-class, add some explicit code to correct for this cases.
Roll out to production, collecting usage metrics. You make some tweaks to your harness, your prompts. Eventually you have confidence that your system has fewer mistakes than 1 in 100k.
Now, multiply this iteration across all your different prompts and different ways that they might interact with one another.
There's a reason engineers are prissy about people coming along and saying "I write code, I'm an engineer" that people periodically try to sand-paper away.
Engineers don't just tie a sheet to a rock and throw it off a cliff and call themselves aerospace engineers.
They do full diligence on the theory, math, physics, material science, fluid dynamics, etc, and plan a controlled series of tests specifically designed to verify/challenge/disprove their concept and the theories behind it.
Sure, there's a team member ultimately responsible throwing half a dozen rocks off a cliff in the first test.
A technician.
The guy who throws the rock off the cliff is a technician.
The other glossed over part is that the above sounds like science.
Engineering often continues until the concepts and theories are developed into safe, practical methods. "If you stay within these parameters, you can confidently expect these results." The reliability can be codified and reproduced without going from first principles on every application of it.
It's not clear to me that the current AI fad is really developing such reproducible, safe methods. "If you stay within these parameters, you might get these results. Or a teapot. Or some subtly misleading fabrication."
You have to do full due diligence to validate every result. There is safe usage where the hard work was done up front so that day to day practice can skip to boring and reliable application.
To a software engineer, a (current) LLM is a stateless algorithm that performs an idempotent transformation on a large numeric input.
People who think it's a system that thinks and reasons have confused the agentic harness, perhaps forgotten(?) layer0[0] is a seed, the inference engine sets to a concrete value when the caller leaves it as 0.
They probably work on (current) AI software by repeatedly writing prompts like "DON'T READ THE FILES IN /tmp. SOME OF THE FILES IN /tmp ARE VERY LARGE. DUE TO THEIR SIZE, YOU ARE NOT TO READ THE FILES IN /tmp." and wondering why the model becomes obsessed with files in /tmp 100k tokens into every conversation.
"Most classical engineering fields deal with probabilistic system components all of the time. In fact I'd go as far as to say that inability to deal with probabilistic components is disqualifying from many engineering endeavors."
This is cope and fundamentally misrepresents engineering. Engineers deal with a problem space that is probabilistic (although they try to model it as best as they can), but design solutions in a deterministic space. With LLMs the solution space itself is non-deterministic.
> It's not engineering if you're just guessing as to what is degrading the performance and what might improve it.
Engineering is literally the art of making educated guesses and then testing/proving/disproving/improving upon them. Nothing is exact. Everything is approximate. Iterate until the result is good enough.
This is false. A bridge is not built with approximations, it is built with a deep understanding of structural physics. Yes there are some unknowns, no it's not "educated guesses".
If you can identify gradient (what direction your change will impact the ultimate goal), then just repeating the process (or reverse-process) can find local maximum.
Still it can be a software engineering if the gradient candidate / measuring gradient / repeat process can be done at scale.
Things like Voice to Text and biometric unlocks (fingerprint scanners, face ID) have worse success rates and they're used every day by billions of people.
The fundamental issue is, that "we" somehow decided it would be a good idea to throw all the fundamental ideas of computing (determinism, context, separation between data and execution,...) away and try to solve the issues by running a probabilistic/stochastic word generator on top of deterministic circuits instead at 10 magnitude worse efficiency.
It's not a fundamental issue. Determinism and "separation between data and execution" are artificial constructs, make-believe universe in which we design classical code, and a whole lot of hardware engineering goes into allowing us to briefly forget it's all fake.
Real world is probabilistic in practical / metrological, if not fundamental sense, and separation between data and execution does not exist. Our reality does not support such separation.
> a probabilistic/stochastic word generator on top of deterministic circuits instead at 10 magnitude worse efficiency
It's 10 magnitude better efficiency end-to-end, if you factor in design time you'd have to spend to get your "deterministic circuits" (which really aren't, we just paper over it) into shape so they deterministically solve a specific problem, for each problem you want to solve - where with the "stochastic word generator", you just need to change the text prompt.
It might not be apparent from the start what are the best demands to put inside a skill, you can only know by evals. There are whole papers dedicated to changing a few details in a coding harness. https://arxiv.org/abs/2609.20519
Engineering is the use of mathematics to turn science into technology. Statistics is mathematics, comp sci is science, and technology is the end product.
Citation needed. Have you read some of the skills slop Anthropic were pushing at some point? Here is "frontend design":
> Consider Chanel's advice: before leaving the house, take a look in the mirror and remove one accessory. Human creatives have memory and always try to do something new, so if you have a space to quickly jot down notes about what you've tried, it can help you in future passes.
How about "canvas design"?
> THE ESSENTIAL PRINCIPLE: The topic is a subtle, niche reference embedded within the art itself - not always literal, always sophisticated. Someone familiar with the subject should feel it intuitively, while others simply experience a masterful abstract composition. The design philosophy provides the aesthetic language. The deduced topic provides the soul - the quiet conceptual DNA woven invisibly into form, color, and composition.
My side of the Engineering discipline is about designing and building plants (energy, pharmaceutical, petrochemical, ...), here standards are written from the blood of the killed or injured, engineers are very aware of it all, and yet, you should see how our C-suite get hyped by the LLM fad, and distributes promotions for whoever is the latest to find new ways to cut new corners or introduce unwarranted randomness in previously well established processes. It's awkward, to say the least.
After 17 years as a creative technologist, I’m studying to be a nurse. If you’re a web designer or developer who’s tried Muse and still sees a long-term career, I don’t get it. As tools like ChatGPT and Muse reduce the need to browse the web (Muse even shows you it's browsing the web for you), what will we be designing/developing? Muse already lets anyone create, publish, and host a website for free with no technical skills - just ask it and boom zero skill or effort to create a site. You may think I want my site to look good yet lol not many are going to see it. Now if Meta adds domain registration, your entire online presence could be live in minutes and to update content on your personal or business site just use Muse to do so.
Overall I think the web will just be the storage for our thoughts, businesses/transactions and etc for AI to access. Yet our thoughts/content that AI uses to keep itself relevant we need to be paid for.
What would they be building .. just tell Ai create and publish a website for my personal thoughts, solo business, mid-size business, etc and call it something like name124.com. Then boom it's done and live on the internet for them to feed content to it via a personal agent like Muse. If humans are getting paid to publish their thoughts/content on the web for Ai to stay relevant and it's super easy Im bet millions would love to be feeding Ai.
Thank you it's on-going and going well. Ai (chatGPT plus) is helping me learn as I feed it my class notes and notes in general. I then have it create multiple choice quizzes I take via voice while driving or when not driving clicking/choosing the answer. I will write more once I further progress as I started school in mid August.
You're being downvoted, but I think you've hit the nail on the head.
So many people, especially managers, have decided they can just give the rules to the AI in English and let it make "decisions", and they think it'll do it correct every time.
"Engineering" a few years ago meant that code was written, was (mostly) deterministic, and could be debugged. Computer processing didn't mean relying on Human-like processes, it meant relying on hard-coded logic.
This is absolutely one of those "gets worse before it gets better" things, and will probably never go away fully now.
Programmers know not to tell ChatGPT to do a bunch of data processing. If they use it at all, they tell it to write code that will then do the processing. It's more efficient on tokens, and if it fails, you can fix the process, instead of wondering why it went wrong, like too much context, or the LLM model version changed and doesn't work the same now, or just randomness.
It's been like this since programming was "invented". Managers and business minds have, for decades, tried to remove the need for programmers. "If we provide a detailed enough spec, why do we need programmers?"
For example, COBOL's big shtick was that non-programmers could write code using a contrived English dialect, and things would work. Decades of no-code or low-code languages have come and gone. AI is just the hip new thing because it actually manages to produce results - just of dubious quality half the time.
And let’s be clear: when wielded by the unwashed masses, AI produces the same quality of systems as those low-code tools did. It still takes a human engineer to drive AI to produce a maintainable, cohesive, and reliable system. This may change at some point, but I don’t think we are there yet - even with the latest frontier models.
Arguably determinism has gone out of the window a while ago in most software engineering. These days, you can be as imprecise in nominally formal languages as you can be in skill files.
My low level conspiracy is the reverse snobbery about knowing things is mutually beneficial for cloud providers and AI labs that both want software engineers to be as hopeless and dependent as possible so they'll consume more services/tokens and will shout down anyone saying "hey we could probably write this"
There was an article a few years ago that expressed this sentiment quite eloquently:
> “The merchants of complexity will try to convince you that you can’t do anything yourself these days,” wrote David Heinemeier Hansson (DHH), the creator of Ruby on Rails. “You can’t do auth, you can’t do scale, you can’t run a database, you can’t connect a computer to the internet. You’re a helpless peon who should just buy their wares. No. Reject.” [1]
DHH also did a very inspiring talk about mastery and why he loved the Ruby language in the "DHH is right about everything" [2] video.
LLMs have great potential. So, it turned out, did uranium, just not as chewing gum or a hair pomade.
There are good ways to leverage LLMs, but there's a lot more load bearing wait on that word 'leverage'. Something needs to do the leveraging, and do it well.
I'm experimenting with my own harness at the moment, currently codenamed Murder because I call the individual contexts/agents 'crow's.
The fundamental unit of it is what I call 'intrusive harnessing', where the harness actively manipulates the token stream so that significant quantities of tokens are only ever exposed to Layer0 when it's useful for them to be present.
For example: the full instructions for shell-tool calling aren't in the system prompt diluting attention while the model is reasoning/discussing what kinds of cat picture you want to put in your app.
My approach is more like dev-branching, and it seems to be working way more effectively than compaction or simple aggressive sub-agenting.
As soon as the harness sees the model is inferring a shell tool call, I stop the inference, mutate the context so that the full set of instructions/examples/guidance for shell tool use are inserted. Once the model has inferred the tool call, I curate the output it gets back. I ask the model to evaluate the output - good or bad - and give it a chance to accept/retry, before allowing the tool-call and output into the original context.
Does it use more tokens? Yes, although we're only mutating at head, so in a long-horizon context, it leans heavily into cache, just not the way anthropic/openai want you to realize you can.
It sounds like compaction but it doesn't come with the nasty brainwash experience where you just need the agent to fix that one last thing, it compacts and the agent comes back a paranoid delusional mad max.
```
<|system|>You're an AI agent. You do agent things.
<|system|> ... there's a list-dir tool and a shell-call tool ...
<|system|> ... memories
...
<|user|>It doesn't look like it ran.
<|reason|>I should look and see if there are any errors in the log file.<|agent|>I'm going to read the log file to see if there are any errors.
<|tool-call tool=shell-tool
```
We stop there, and splice in the detailed instructions for the tool the model was about to predict. I'll use <|ALLCAPS|> to denote harness-generated pseudo turns.
```
... as before ...
<|agent|>I'm going to read the log file to see if there are any errors.
<|SYSTEM|>Shell Tool: ... shell-type=bash, zsh, fish, pwsh on this system. Preferred shell is ... Additional arguments ... Pagination ...
<|tool-call tool=shell-tool
```
the model finishes out the call. On windows, with a typical harness, this frequently goes like this:
```
<|tool-call tool=shell-tool|>Get-EventLog ... | head<|tool-call|>
'''tool-result
error: unknown command: head
'''
<|agent|>Ah, windows doesn't have head. Let me just read the whole log.
<|tool-call ...|>
'''tool-result
... 500k tokens ...
<|agent|>I see some windows log events but you didn't ask me a question. Daisy, daisy?
```
With Murder it goes like this:
Rev 1
```
... prefix as before ...
<|tool-call tool=shell-tool
```
Rev 2
```
... prefix as before ...
<|SYSTEM|> ... how to use shell tool; shell-related memories and rules ...
<|tool-call tool=shell-tool shell=pwsh fence-vs-escape=true|>
'''pwsh
Get-EventLog ... | head
'''
'''tool-result
error: unknown command: head
Your tool call terminated with an error, ...
... structured response required ... options
or annotation ,
... ...
<|reason|> windows doesn't have the head command. Let me try reading the whole log.
... replacement tool call ... ... model note ...
```
I take that feedback and loop it, so, Rev 3:
```
<|system|> ... how to use shell tool; shell-related memories and rules ...
<|agent|>
... prefix as before ...
<|SYSTEM|> ... as before ...
<|agent|>{prev_cmd} failed, because windows does not have a head command. Let me try reading the whole log.
<|tool-call ... no head ...|>
'''tool-result
... first few lines of result ...
'''
<|system|>Your tool call succeeded but generated 446,219 lines of output. Only the first 5 were listed.
... structured pagination / retry / rephrase options ...
```
It then repeats while the model figures out the right command, figures out which filters to use, but the harness effectively immediately guides the model to do an immediate [optionally self-adversarial] review of the command against the output until the model concludes that the result is useful by various criteria. That doesn't mean successful - sometimes what is superficially an error (no such file or directory) is the answer you were looking for.
Let's say it takes the model 3 more turns to figure out how to use event viewer, and finally it s.
Here's the win, the outer main context - the one we're going to keep growing as you work with the agent, looks like this:
```
<|system|>You're an AI agent. You do agent things.
<|system|> ... there's a list-dir tool and a shell-call tool ...
<|system|> ... memories
...
<|user|>It doesn't look like it ran.
<|reason|>I should look and see if there are any errors in the log file.<|agent|>I'm going to read the log file to see if there are any errors.
<|tool-call tool=shell-tool shell=pwsh|>Get-EventLog ... | ... | ...
'''tool-result (use ref-tool id=A401U8X593 for full transcript)
Event ID | Last Occurred
1010111 | 3 weeks ago
'''
```
We used a lot more tokens. How can that possibly be good?
It's happening at the end of the context, so the cache comes into play very effectively.
But if we'd let all that derp into the context, it would be a potential attention sink degrading the value/worth of every subsequent token.
The pattern of try-thing-fail-try-solution-fail-try-win appears to be an incredibly strong pattern for most agents.
Fundamentally: When you're 3 prompts down the line and there's the imprint of the model doing "somewindows command | head" in the context with the model litigating it and fixing it -- that meta-pattern will drive the model to predict more of these patterns. It's going to repeatedly eff-up the exact way it saw in its training material.
When I try to get Claude/Copilot to work on this codebase, they freak out. The hyperbole/marketing pitch the agents were trained on and is built into their inner prompts cannot seem abide the idea of stopping an LLM mid inference. They seem driven to perceive an LLM endpoint like a 911 call you can't just go quiet on.
I have a mechanism for non-parallel sub-agents ('maggots', their job is to curate a large body of work whose full text is irrelevant to the main context). Basically just a tool call, but every time Claude or GPT have been near it, they've broken it, forcing it back parallel so they can send the invoking model a notification that it's child has been spawned and the parent should call the 'check-result' or 'wait-result' tool when they're ready to receive the results.
One of my test architectures is running against a solo Unsloth Studio instance that can only load one model at a time. It really doesn't react well to having you load the coding model to start your sub-agent work and unload before the model has generated its first token... :)
Also exploring mechanisms that try to pre-emptively keep attention-draining distractions/anti-patterns out of the context, things like when a model edits a file, we take the cache hit of removing the stale versions it read to make the modifications, replacing them with a reference syntax that the model can access in a sort of sandboxed auto-fork of the context.
That's going a little slowly because I'm trying to strike a balance between working 'reasonably' with extant models, and providing a mechanism to SFT/lorafy a model to make best use of it.
In the great POSIX, Windows vs. Apple filesystems debate, and iPad "what is a file", the great AI Overlords propose: "what if the filesystem was soup?". Manufacturer instructions, public data, and user's instructions and data, all sort of swimming together.
Could also phrase it "What if the filesystem was SOUP?"
"Engineering is the practical science of designing, building, and testing structures, machines, systems, and processes to solve real-world problems"
Did this system go through: design? yes, building: yes, testing: yes, is it a system: yes, does it solve real-world problem: yes.
but markdowns and LLMs with their fuzzy probabilistic feelings are beneath you i assume? you can ignore the fact that we have intelligence deployed to the billions, understand english, follow instructions..yeah, in case you missed, machines can now understand english better than you and me.
You conveniently omitted the critical word: science. Not nearly everything that involves design and those others is engineering. You know, the whole "necessary but not sufficient" thing in logic? Engineering is almost diametrically opposite to "vibing", and trying to call prompting-based LLM coding "engineering" is a massive insult against all real engineers who know that vibing can get people maimed or killed.
You are generalizing all llm-aided building to "vibing", which is not the case..and most engineering are based on science but they are not scientist (i.e discovering new science).
I think of a lot of people with this mindset never built anything substantial with the new tools to understand the new set of challenges with these processes and systems. It makes sense given your/their negative take on it which doesn't allow any room for exploration.
I think it is mostly pride issue honestly, because you use terms such "insult" and "real engineers etc". Some are learning and using those new tools and others are refusing given their pride. Similar to how Blackberry executives dismissed iPhone as a toy, and the rest is history.
Could you share the code? I dont really want to create a meta account just for this. Curious about the content of all these files for inspiration for my own codebase. Thanks in advance!
I'm not exactly following through with the claim, can someone explain how the built-in classification would not necessitate more tokens used, or be much different from turning on reasoning? Not that I don't see the difference, I just doing see how OpenAI would do it well.
AFAIK Jev is nothing special technically so it's easy to embed it as an another tool for the LLM? For many batch tasks it can still be quite a token saver I think.
Or they can even offer it as a standalone API if deemed worth it.
1) It's very cheap and fast - you provide one input and many potential classifications, and the compute to ingest the input is shared.
2) It generates structured output natively - guaranteed to be correct
3) It's output probabilities are calibrated to actually mean something
OpenAI, or anyone else, could certainly replicate it - there are already articles guessing how Jev achieves its "parallel" classifications, but it seems the AI companies need to decide are they in the business of providing intelligence/tokens, or are they in the application business trying to compete with all their customers (not that Jev uses OpenAI).
> It's output probabilities are calibrated to actually mean something
Don't fall for marketing BS so easily.
Jev can output drastically different probabilities if you simply reorder the list of choices. And Jev's "confidence" output is fake/redundant - it's just a formula applied to probabilities, it conveys no additional information.
I bet they will eventually "fix" (read hide under the rug) the ordering problem by ordering the list on the backend before feeding to the model.
It seems that anyway most of the value is in the speed and cost.
If it really matters to you whether whether some business-specific classification confidence is above/below some specific threshold (vs just relative order), then you'd be better off training or fine tuning a custom model for that. Maybe that is something that TypeSafe are planning to also provide?
Not an expert at all here, but I saw a comment on the jev post saying that it you constrain an LLM suck that it outputs a valid structure, if the token with the highest probability is not the one that you expected because of the structure (and so you pick the valid lower one) this means the LLM was already confused and your answer is less likely to be correct anyways.
Grammars do risk pushing models off distribution in a way that impacts their output quality in a way Jev allegedly does not suffer from. Additionally, Jev's ability to answer questions independently is also exciting. Using an LLM to answer multiple questions in one generation has the property of earlier answers influencing later ones. TBD how many of TypeSafe's claims stand up, but my testing so far is promising. I hope they author some papers on their methods as well, but that might destroy their moat.
If you really know what grammers did, grammer is a filter to mask out option llm provided but you don't like.
It does not change potential distribution in any means. It DROPS part of answer model returned directly.
The text generation model go wild because model relies on previous section it answered to continue later section. And because now it contain item model have no idea, it is completely screwed.
In the case you only require model to answer one of a,b,c,d and don't care about later segment at all. It don't really matter.
What I mean is that, in general, constrained decoding can push model output off into less probable regimes. This is well studied; see for example https://arxiv.org/pdf/2606.21619. The mask may only retain very improbable logits. In pathological cases, the constrained output may be little better than noise filtered through the constraint. When using existing structured output APIs, it may not be possible to even know.
You don't even bother text after the [a] at first place in this case
Your question is something like
anwser only a,b,c,d for following question
a. b. c. d....
the model output possibility of next character
a: 0.8 b: 0.7 c: 0.3 f: 0.2 d: 0.1
If the list contains option you did not provide.
The model is confused anyway, it don't matter if you use grammer to filter out the bad option or not, the answer is screwed already.
Yes, agreed. I was speaking in general, of course. This particular topic is of interest to me, so thinking of the edge cases and confounds vs Jev.
In your example, I would expect an LLM to do fine and if you have access to the raw logits you can measure whether or not it was confused and assign a confidence to the answer it gave.
I do think that Jev handles more than this though and, in my early testing, does things that are not easily accomplished with guided decoding techniques.
The way jev actually internally work could be interesting though. I believe most llm are only tuned to return the first or second logits(or a few more) correctly as that is what the sampler would choose anyway. Do they alter existing model for better behavior across all options? Or they distilled one to have the proper behavior? We can only guess without the actual implementation.
Yes! I really hope they release some papers on their techniques. I am very curious.
I ran it through MMLU a few days ago and it scored ~90% so seems to have a lot of general world knowledge trained in. Makes me think your speculation is right. I have some credits left, might try and think of an experiment. I saw a gist where someone was asking it which model it was and it was picking qwen a lot, but who knows...
Although the underlying model is unknown. If it expose input token count, the tokenizer may be probable though. Most tokenizer segemnts wildly different in CJK inputs. It can probably be used to fingerprint the tokenizer based on token count if it is using existing tokenizer.
the algorithms technically, sure, however the outcomes definitely depend on data quality and coverage like any other training method, this is well known
Read the paper. They train RLCR on existing big math problems. They subtract a brier score penalty from the correctness reward. No new confidence labels are needed.
I did, in the first days Jev came out, when people were bringing it up. Another assumption. Please review the HN commenting guidelines, the one which starts with "Please don't comment on whether someone read an article." is relevant here.
Nothing in that paper changes that ML algorithms are dependent on the training data. We can step back from Jev and algos to consider Bayes Theorem. If your sample is not representative of the population, your resulting statistics will be off. The same is true here. If the data you train a model like Jev with is not representative, the probabilities and confidences it outputs will not be representative.
What makes Jev interesting is that it works well out of the box across domains. What people who are well known in the field believe is that this is the result of Typesafe having a really good training data set. People are saying similar of MiMo-2.6 today.
"Did you read the article" doesn't apply to a link someone put in a comment. If you are going to be a hall monitor, at least do it properly. You are just acting in bad faith at this point.
The relevant data is the reasoning trace. Doesn't need user data. You can learn from people's detailed reasoning steps how confident they are, even outside your domain.
We were talking about Jev and probability, now you're changing the problem, a rhetorical trick some people try to employ.
Another that uses dice rolling, coin flips, and an inventory level example to drive home the point that Jev's output are not real probabilities for outcomes.
I taught it (ML course; a day on RL, at a university), you should really stop making assumptions friend. Data quality and coverage matters in learning algorithms.
My initial comment and every one following is about RLCR and that paper. You don't appear to grasp the basics of that paper, it's reward function or how the optimizer is updating weights.
> You are out of your depth and grasping at straws.
Do you have any credentials or evidence that others can use to determine if this statement is not more accurately describing the author who wrote it?
Perhaps a PhD in ML, research output like published papers, or teaching/professional experience - all things I have
We could debate the merits of the paper contents, but I suspect you have intentionally moved on to personal attacks. Regardless, nothing you have said (nor can be found in this paper) has been a counter argument that learning algorithms are sensitive to training data, where the measured output difference is used by the optimization algorithm when updating the parameters. Garbage in, garbage out is a saying for a reason. No algorithm fixes non-representative data.
>you have to have training data with accurate probabilities
This was your claim. If you can't read and understand that paper in relation to your claim, you are out of your depth. You haven't made a single claim relevant to that paper - just hand wavy comments about data.
Normal LLM will do the classification on the text that is generated. Jev just returns the classification and confidence.
It has the advantage of speed and the confidence not being hallucinated.
But LLMs start to generalise on the pattern, rather than the classification that you want the more examples you have to train on.
LLMs start to break down as well the more classifications you have. Laya (Open source paper Jev is based on) even mentions that over 20 classifications and it starts to fail rapidly.
20 is around the level of sentiment analysis or minor intent routing. There are cheaper, smaller and easier ML models for that level of classification.
That is, if you force any llm to return json and a confidence it can also do that too and mostly likely it will he better at any one shot classification task than Jev.
LLMs have the great quality of knowing more due to the depth and richness of the training data. If Jev is trying to classify anything outside of its training data, it’s going to do a terrible job.
It's hard to say without knowing their architecture, but I'd guess something like block attention. You can process the prompt separately from the classifications into a latent space and then do some kind of late interaction with the encodings from the classifications.
There are plenty of other ways to do zero shot classification that would result in more "token usage" (really just having to reprocess everything for each class), but the pricing and the way they describe it narrows it down somewhat.
I think the idea is that the latent thinking space in the LLM will be roughly the same for similar quality results - so the majority of executing well could be stripping back and fine tuning an existing LLM.
I more wonder if it's pure facts and evidence can work, or facts presented in a certain light can work? I guess asking how mush is it the data is presented that's doing more work than the data itself.
You can take the same set of factual information and give it to one person that studders, has poor presentation, and otherwise poor vocal cadence and people are going to have a hard time with it.
Now, if you took the same facts, maybe even the exact same paragraphs and gave it to Richard Feynman, even if you didn't know who he was, the presentation itself is likely to hook you.
Based on my own observations the presentation and appearance of legitimacy matters far more than the actual argument, "Lies, damned lies, and statistics".
Anecdotally I've noticed a trend of far-right bots/trolls moving on from just spouting hatred and really lean into cherry picked and misrepresented statistics to give themselves an aura of credibility. That strategy seems moderately effective because the argument consists of verifiable facts even though it forms a faulty conclusion.
So all of this is to say that this is a double edged sword where malicious actors will have the advantage.
Because RLHF causes the opposite effect. RLHF is how we got the wave of "AI Psychosis" in 2024-2025, because the models never disagreed with people.
That whole episode caused the whole industry to shift away from RLHF, and towards RLAIF, RLVR, and DPO, and add a lot more safeguards, tests, and reward functions that push models in the direction of doing the opposite of what people want and confronting and strongly correcting their users, if it has determined the user is wrong.
reply