Looks like yet another non-price-competitive Jev competitor.
Microsoft only compares the price of theirs to GPT Sol(!), not GPT Terra, or GPT Luna (which is what OpenAI's Jev wannabe is based on), and certainly not Jev (4/10 the cost of Luna).
I can't remember when a new product created So many competitors so quickly. What is very clear is that everyone is saying "Doh!", slapping themselves on the forehead, and scrambling to get a slice of this obvious-in-retrospect massive pie.
What no-one appears to have done yet is to come close to Jev on pricing!
Well let's see if the Chinese open weight models come in with even cheaper models. Deepseek came out of Algo traders they probably already have small fast decision model they use internally
Open weight model pricing isn't really up to the company that built them (unless they are also serving them, in which case their API price may differ). It's really up to the serving cost and/or pricing of whoever is actually serving the model.
The main market for Jev is going to be business, who may not want to deal with a low cost open weights server like Fireworks AI and are more likely to go with a provider like AWS where pricing is higher.
A very major part of Jev is the cost and speed. Yes, classification is/will be a commodity business, just like LLMs are, and similarly there is no moat only production cost and pricing.
Yes, anyone can wrap a decisions API around an LLM, but so what? If you want to compete then you need to compete on price, and it's not clear if OpenAI and/or Anthropic are able or willing to do that without building a custom architecture, and even then is a race to the bottom on pricing really what they want to pursue?
I'm not sure if OpenAI have announced pricing for their Decisions API, but they have said it's based on Luna which costs $0.10/M input, not even remotely competitive with Jev's $0.04/M input, which I'd expect has some headroom built into it.
Assuming that the architecture behind Jev is not just an LLM, and gives them some inherent efficiency/cost and speed advantage, then the question is whether OpenAI and Anthropic really want to duplicate this and have a race to the bottom on pricing for what may be a large part of the business automation market they are addressing. Is that what they want as their IPO pitch - we're selling potatoes, and think can grow them cheaper than Typesafe ?
If you are referring to the recent UAE datacenter attacks, AFAIK Amazon's services worked as advertised, but obviously a redundant resource is only as good as the number of copies you have left, and eventually 3 of 3 datacenters were hit.
Apparently not all AWS resources/services are redundant - some are advertised as single location, and then it's up to the customer whether they choose to pay to make backups elsewhere (e.g. in a different region) or not.
A lot of the smaller drones that Ukraine is using only do limited damage - they may start a fire that can spread, but are not going to destroy an entire larger facility (although the Flamingo can do significant damage). Many Russian energy facilities just get hit repeatedly over time - Russia will fix as fast as they can. Staggering hits seems to make sense.
Energy facilities have the destructive energy stockpiled on-site. Don't need to fly in much force yourself.
Also why I think governments should mass-produce & subsidize blow-molded 25L plastic gas cans instead of building centralized storage tanks/caverns. No way those things cost anywhere near US$25 to make.
I think the energy facilities being hit are mostly refineries - not storage, and anyways any bulk storage is just temporary for bulk transport for export or to the front, etc, so having it in 25gal cans wouldn't be very helpful!
Countries still have strategic reserves, but usually of crude, not refined product --> decentralized reserves would help with supply cuts by any cause. The benefit would be too individual vs export/frontlines, so won't happen.
There are chemical based emp bombs, NNEMPs, if the penetrate through to the inside and set that off it seems like it would be pretty effective on datacenters.
Russian air defense is multi layered and relatively strong, so one of Ukraine's main strategies so far has been just to overwhelm them with volume. You want mass volume cheap and cheerful, not expensive ones, unless for the occasional strategic target.
Ukraine only has limited number of more capable missiles/drones like Flamingo FP-5 and Neptune, although now the FP-7 is about to hit volume production, followed by the FP-9, as well as Europe-supplied missile and anti-missile systems, so it remains to be seen if this will change Ukraine's strategy ... I tend to doubt it - they will have firepower to retaliate much harder, and at further distance, when they want to, but I expect they don't want to unilaterally escalate.
I seem to recall that this is optimal game theoretic strategy - don't initiate, but always immediately retaliate.
Whatever is happening now is payback for Putin's birthday present to himself from 2 days ago, firing a missile into a Ukrainian apartment building that killed 22 civilians (mostly women & children).
Apparently Wildberries carry/carried over 1/4 million items tagged as "for SMO" (Special Military Operation), so they had become a major part of the supply chain for Putin's war, which is why they were hit.
In the case of Wildberries we're not talking about "dual use" items like a sleeping bag that your government fails to provide you, we're talking about specifically tagging items (1/4 million of them!!) as being for the war. It's Wildberries' business (literally) if they want to be a supplier for the war, and profit from it, but obviously that does mean they become a target.
Cutting off the enemy supply line is a classic and obvious military strategy. Ukraine didn't choose to be attacked, nor did they choose for Wildberries to become part of the supply line. They are playing the cards they were dealt, and I would say doing so with remarkable restraint.
of course that'd be fine, under the Geneva convention a civilian object loses its protective status when it makes effective military contributions and its destruction confers a military advantage.
What's supposed to be the argument against it, inconvenience for American shoppers during a war of aggression?
I don't think the Geneva convention is that unequivocal, I believe there's also a duty to assess proportionality of military advantage vs. civilian risk.
Apparently "only" needing the problem statement to be correct (i.e. expressing/rewriting a mathematical proof in Lean) isn't so simple, and if you don't get it correct then you haven't proved (or disproved) what you were intending.
Notable that one suggested use case for Haiku is "classification requests", i.e. Jev competitor, and the pricing matches GPT-6 Luna which is behind OpenAI's "Decisions API" Jev competitor.
For this application 100K token input is plenty.
Of course Anthropic and OpenAI, both at $0.10/M, are still 2.5x the cost of Jev's $0.04/M.
The classification performance remains to be seen, but presumably we'll soon start to see classification benchmarks.
For other tasks like summaries (another suggested usage) it's good to see Haiku and Luna now competing against each other on cost.
I'd love to know how the business automation market breaks down by volume of call type though - hard to imagine that decision making (e.g. branching, triage) isn't a very large part of it, greater than these other suggested Haiku use cases.
As you note, it's going to be based on what the LLM knows about the user, and as any good predictor knows, rich people like to buy expensive stuff.
Whether or not the article has examples, it would be surprising (an LLM prediction failure) if this expected behavior did not extend to different prices for different users, e.g. recommending the rich guy the Whole Foods bananas and the poor guy the Walmart bananas.
> The letter counting issue was due to how LLMs split text input into tokens
No - this is provably not the issue.
Take any model that fails to correctly count the letters in a word, and ask it instead to spell the word (even a made up word), and it will be successful - they have no problem predicting the letter sequence from the token sequence (and would be shocking if they did - this is what they are built for: seq -> seq prediction).
The reason LLMs can fail at the letter counting task (depending on model training, prompting) is because of the counting part, not because of any difficulty correctly mapping the input token sequence to the letter sequence.
even though the constant is tiny, this breaks the long held assumption that nlogn is the minimal possible bound, so further research seemed unfruitful. this discovery will now trigger more research in this area, and probably more optimal algorithms will be discovered.
Like matrix multiplication, the common assumption was that it cannot be improved past n^3. Then Strassen broke the barrier (with a more significant constant) which caused intensive research - and now we are around n^2.3..2.4.
Microsoft only compares the price of theirs to GPT Sol(!), not GPT Terra, or GPT Luna (which is what OpenAI's Jev wannabe is based on), and certainly not Jev (4/10 the cost of Luna).
I can't remember when a new product created So many competitors so quickly. What is very clear is that everyone is saying "Doh!", slapping themselves on the forehead, and scrambling to get a slice of this obvious-in-retrospect massive pie.
What no-one appears to have done yet is to come close to Jev on pricing!
reply