> But even the narrative that hey microsoft fired 1k citizens and hired 1k H1bs that are lower paying is false. These companies are huge organizations. If you laid off 2k people during your annual cycle in Xbox and start hiring 2k people in cloud 2 months from now (some of whom maybe H1Bs who are also impacted by layoffs), it's not the equivalent of firing an American and hiring a lower paid H1B.
I don't think this really holds up. They could offer the American engineers the opportunity to move divisions, but they don't because what they really want to do is what they did do: Fire Americans and hire foreigners.
> Will people really come to the US on a H1B if they know they can get scaled down like a cloud server at the first sign of economic trouble or the first sign that some politician wants some brownie points?
Then make H1B what it is supposed to be for: elite specialists, not infinite low cost non-American laborers. Then you won't even need to have this discussion.
I remember the controversy around the original announcement. It seems quaint now, honestly. Simpler times. Slop by those olds standards is practically the Mona Lisa now.
Is this fundamentally different from other text-based LLMs, or is it the same except with special reinforcement learning a safe guards around generating valid types?
Surely it’s still generating some kind unstructured data internally? For example, what if I told it to generate a short story, but the short story is output as a JSON string?
So Jev won’t write a story or emit arbitrary structured data. But if you ask it the right questions, it can make near-instant “decisions” against those questions, with accuracy and world knowledge on par with LLMs. The economic advantage is that it’s parallelizable and can give back up to 255 answers at once, in milliseconds.
Rare as in your grandfather's John Deere manual from 1982, not rare as in a test print run of The Great Gatsby. The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
> The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
Citations? Also what exactly are these 'normal means'? As a bibliophile who loves scouring used book stores for out-of-print titles this is a topic I'm very interested in.
Normal means is throwing it in the recycle bin. Especially for stuff like a 1982 John Deere manual. I've never donated an old appliance's manual to the library. Have you?
Amongst my friends, I'm one of the rare folks who donates books to the library. Most people just trash them. And I know the library only wants them to try to sell them in their book sales (or online) so they can get money. Almost nothing one donates to a library actually ends up on the library shelves.
I moved one of our daily workflows over to Kimi K3 on Fireworks. It replaces two sales assistant positions, and was about $50/day for 14 million tokens total (in/out). They surely are not subsidizing this price as it is just inference only and other providers are even less money.
The general model for subscriptions is that power users are subsidized by subscriptions of casual users, like a gym membership or whatever.
This is a little dicier in post-agent AI, because it's easier for casual users to automate power-user consumption, but the providers have done decently in discouraging that.
There's people here saying they're obviously subsidized, there's people here saying they're obviously profitable. I think they're probably subsidized, but I would hold back on saying it's obvious.
I don't want to host my code on a platform that has activist goals beyond open source advocacy, but I'm glad it exists because it's a useful filter, like the X/Bluesky divide.
Can someone explain to me the difference between this approach and using planning with a larger model, then just switching to a small model for implementation without clearing the context? I understand that it specifically does the first edit as well either way the larger model. Is there some other difference I am missing here?
When a frontier makes a succesfull edit based on the plan that it made, it leaves an procedural trace in turn biases the NEXT model, low cost model, straight into procedural action. The cheaper model doesn't need to reread everything again because it has enough information from the frontier model to complete the task. A simple "plan" of what needs to be done does not carry this information.
I don't think this really holds up. They could offer the American engineers the opportunity to move divisions, but they don't because what they really want to do is what they did do: Fire Americans and hire foreigners.
> Will people really come to the US on a H1B if they know they can get scaled down like a cloud server at the first sign of economic trouble or the first sign that some politician wants some brownie points?
Then make H1B what it is supposed to be for: elite specialists, not infinite low cost non-American laborers. Then you won't even need to have this discussion.
reply