I still have this naive notion that we don't need LLMs for code generation and editing. Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?
Maybe people smarter than me know better but couldn't there be a middle ground where an IDE/Editor has an embedded engine (doesn't need to be a full-on LLM) that doesn't require external tool calls and token spend?
If an organization is paying $2400/year per developer for tokens and a highly intelligent editor/IDE comes around that charges $1000/yr and gets more output at a fixed cost, its a no-brainer of a decision.
> Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?
Turns out that actually - no. Researchers have managed to prune half the Experts in a MoE model that had a low probability of getting activated during coding tasks, resulting in a more focused model:
Main benefit is that it greatly reduces the amount of RAM required to run these models. Of course you could just cache those unused experts on disk instead, but the main point here is that you know which ones matter.
But aside from that recent models, like Qwen3.8-27b are reportedly more durable under heavy quantisation, e.g. 3bits or even ternary. With additional techniques like TurboQuant, you can feasibly run these models on consumer hardware - even if at 1/4th the speed you'd get from rented infrastructure.
VS Code has extensions such as Kilo Code or llama-vscode which let you work with local models much like you would with cloud based solutions.
You know what, considering there was a recent "small" open weights LLM released recently that meets 90% of my coding needs I'm inclined to agree.
Qwen3.8-Flash-Next - relatively small, it runs on 6 6 year old GPUs on my home PC happily running 5 simultaneous 262k sessions with additional 10 cached in RAM (bought back when you didn't have to remortgage your house for Ram) and it has been the first local model that is not a toy.
But there is a class of problems where I still reach for Anthropic's fable...
However, I have a hunch bordering with certainty Anthropic is achieving such great results by doing a lot of harness tricks.
For example opus 4.8, is not much better on coding than before mentioned Qwen model, but gets amazing results on factual knowledge stuff (the knowing all works of Shakespeare thing). How hard would it be to add a general knowledge RAG to requests that contain relevant questions and beat all benchmarks like that? Not very hard.
So I think there is big innovation to be had in harnesses, routers, inference and so on.
As to money spent on AI per developer my current client (a fortune 200 software company) spends $500 per month. That is $6k a year. A lot more than your examples. And many people run out of their quota pretty quickly.
I was inclined to listen to your take… then you started talking about having 6 graphics cards in your computer and acting like that’s a normal thing that people do…
The other pitfall here is that we are tempted to compare this crazy setup to frontier models when the real comparison is running this same model in OpenRouter.
This 6 GPU setup will probably outspend OpenRouter on electricity alone.
we use that same model at our company, it powers not only our devs but also many business needs. We rent 1 gpu (B300), it costs around 10x less than the api costs
Do you rent from AWS or some other provider? I wonder if there are issues with on-demand rent for those high-end GPU instances, is capacity always there or sometimes it is unavailable?
> "Does a system need the knowledge of the full works of Shakespeare to be able to output Javascript?"
If the implementation brief says "attempting a reconnect in this handler would be a wild goose chase", the model needs to know enough Shakespeare, at least indirectly, to understand that expression...
Forgive my naive understanding of LLMs - but how do you get semantic understanding of a codebase, such that it knows what changes to make/why/where, without a wider understanding of language more broadly?
I'm using the word understanding loosely there, but I couldn't think of another word.
Depends on what exactly you want to change. Lsp can do a lot but only in very simple changes.
Intellij when I used to use it had a lot great features like refactoring, extracting part of code as a function, renaming and creating empty classes/boilerplate but that's it
In current job I can order LLM to take data sink from other endpoint and write new with given URL. It will fetch from endpoint, check what it gives, compare with other and write new sink. Then it needs polishing because it always create something as awful as possible with cloning data all around but the most boring and soul sucking part is done
This is why I am working on https://github.com/spockz/semantic-editor. To bring more powerful editing functionality to agents. In a way that attaches to their chain of thought and deals with their probabilistic framing in json so it all works out cheaper and faster.
It's a totally fair question. I'm personally wondering if there's a half way point. Some kind of structured language that isn't plain English that a "dumb" LLM is able to parse. It could be human written, or it could be written by a "smart" LLM at a greater cost.
Looking at what Jev has shown, and what is being done in the space. I think you are right that there’s probably a lot of room for performance improvement on specific tasks and workflows that will be happening in the next few years
I also imagine it could be a big shakeup if all of a sudden models could run on CPU. Imagine running an Astra-level coding agent, locally on your laptop. All of a sudden GPUs wouldn’t look as valuable, if you don’t need them as much
We are still some time away from that, but it seems like progress is being made
I have a feeling that when all the AI-hype dust settles, what you describe will be the killer app of AI. That, ChatBots and unstructured data processing. Huge productivity improvements but not the sci-fi hype of today.
We’re conditioned to interact with language models as chatbots and in that sense strong language understanding (implicit - some sort of world knowledge), is probably necessary for that?
But I’m sure we can have a small model that’s really strong at programming concepts, JavaScript syntax, and that’s about it. You’d interact with it differently, at specific seams in your code base - review a PR, merge two functions together, investigate these logs.
Or maybe I’m just not adequately absorbing the bitter lesson. Idk
I think the amount of knowledge to correctly work on code is more than you'd think, because at the end of the day writing code without an understanding of the environment it exists in/for is likely to not fit the problem correctly. Maybe it doesn't need knowledge of _Shakespeare_ per se, but if you were working on a virtual tabletop having knowledge of tabletop games can help with identifying the right implementation to use, knowing what kind of constraints to consider, etc.
I don't need AI in the same way that I don't need autocomplete. I can definitely program without autocompletion, but I'm a lot slower than others who use it
It is very difficult to predict ahead of time what knowledge a model would need to understand a prompt to generate a program. It could refer to all kinds of real world knowledge referring to the kinds of entities you want the program to model.
I think this already happens through Mixture of Experts which is now build in to ost models.
But finding out what an LLM needs to understand from a business side to write your code good, is an otpimzation which no one cares currently.
I'm pretty sure we either stay on big full frontier models for a long time, just use them for everything or we will start to see more and more people doing finetuning/project specific training like java + german + english + business contxt xy;
My gut says the economics, e.g hardware/data center/resource constraints, are going make the economics of small specialized models more attractive. Without any evidence whatsoever, I also think that the big frontier companies will have to de-emphasize chatbots as huge models in favor of chatbots as huge products with a much much more granular mixture of experts approach, but with much smaller models. I’ve been saying for a while now that AI products have to hit the gas on prioritizing product design to reliably solve real people’s problems in predictable-enough ways, because the current approach is only really appealing to enthusiasts, developers, or optimistic managers, and with the kind of money they’re throwing around, that’s not going to work.
Everything around "desired machine physics" is superfluous wank; historically a biz case stored as code when some UI could feed biz case params go a function generator
Come on we know what we use computers for; media consumption and 2D data entry/review. Locally we just need a core engine for geometric transforms of visual state. What all these languages give us ability to create such a generic VM filled with customized semantics that mean nothing to solving the problem but plenty to a clever coder.
Kind of like Unicode we need distilled geometry primitives like "teapot for text" and desktop metaphors and to let people put the superfluous wank at the presentation layer
Which text used to be so making UI out of layers of text, OOP, and such made sense for decades
But we're just engaged in bloating system state through def jargon_to_encapsulate { desired machine physics } when we already know it's going to be simulated 3D or 2D visual transforms. We don't need to capture all those states in code verbatim.
Things like Jev are the future of models. Fine tuned on transforms given a context. "So you want to replicate GTA5? Here's a data set of geometric shapes and gradients constraints from all observed xyz" pipe that into your local renderer
We're entering the phase of software engineering (and engineering generally) where we realized we been dramatically over playing the song and can strip out entire asides and digressions, circumlocutions of provenance, to tighten up pacing and improve enjoyment of the outcomes. Hopefully. Or we kill ourselves. Through social squabbles (political, economic, religious, whatever) due to laziness to learn etiquette, and environmental destruction.
1. Storing Shakespeare's work costs almost no $ in regards to disk space.
2. If the prompt doesn't include "Shakespeare" or relevant terms then no regression is performed for that topic and therefore there is no effective token cost.
Someone may correct me, but I think it's not a big $ win to exclude relevant topics from the models' overall capabilities. Instead you'd tune weights so that #2 better identifies what is or isn't among the relevant terms on which to run regressions.
I think you're probably missing the forest for the trees here... broadly speaking, these multibillion "parameter" (whatever that actually means) models store a lot more than shakespeare, and have storage costs in the 100s of GB/TB (which translates to $$$$$$ in SSD/RAM costs), nevermind the (kilos/mega/giga)watts involved, all the pollution, etc...
Meanwhile, a template (maybe a couple KB) costs less than a couple cents to store and run. Large Languages Models are not really interesting, (smaller) LLMs that only contain "what you need" are.
To me it feels like the difference between crows and humans. Yes crows are smart, but what we really want are all inclusive models capable of human "thinking" . I guess we probably need a mix of both so we can give the crows the easy jobs freeing up resources.
> Maybe people smarter than me know better but couldn't there be a middle ground where an IDE/Editor has an embedded engine (doesn't need to be a full-on LLM) that doesn't require external tool calls and token spend?
IDEA already have small LLM for one line code completion IIRC.
But the gain people want from LLM is generally "here, add this entire feature" or "here, go thru every dependency's changelog and update code to work with latest version". Those are not small LLM tasks
I have hope we will get there eventually, once all the hype/wealth extraction/boys club giving all their buddies money cycles end, and the specialized tools with real value start to emerge.
These specialized tools already deliver tremendous value. What happens on the backend financially is of no concern to me, as I have no influence over it.
There was a run on Mac Mini's for people running OpenClaw earlier this year. I was doing some research earlier today on an idea and that came up as an option because it (and Apple's hardware for Macs) have shared RAM across CPU and GPU so its apparently a viable option to run ollama (ironically from Meta) or other local models for simple tasks.
I think VPS are going to be expensive because of the RAM and GPU requirements - 32GB RAM is beefy for a VPS and might run like $30/month on the cheap end. Thats $360/year so running something at home might be more secure, capable and cheaper. Remote/mobile access might be viable with something like tailscale.
Meta's active user base doesn't care about privacy.
Most personal tasks take trivial amounts of time these days. Does a 5th grade teacher or an auto mechanic need a personal assistant? Do we need to be able to book a trip faster? It feels very gimmicky to me, but I realize I'm not their target audience.
Given that, I do think Meta will probably get the lions share of B2C AI and they have a proven track record of using personal data to serve ads and make a lot of profit. If they capture more of their user base's time, it keeps them engaged and away from a list of competitors. If Meta starts crushing it in this fashion I could see IPO's of OpenAI and Anthropic land like fart in church.
I mean I just used muse to plan out my entire weekend trip to Vegas recently. I booked everything myself but I asked it to shortlist dining options for me given dietary restrictions, events given restrictions (with elderly family), etc. it basically handled all the stuff I would have normally just Google searched for. Plus it made a nice web app dashboard to display it all. It seems pretty useful for common use.
It also handled payment splitting. I told it what person x spent for the group split 3 ways (it originally split by the total number of people, but I clarified that only me + x + y are paying) and it kept running totals. When I got an email from Venmo saying x sent me money, muse automatically updated the balance sheet.
We saw some signs written in an unfamiliar language. Muse told me what language it was.
I think some of those tasks are part of the fun of a trip - doing the research on locations, events, restaurants, etc. You stumble through a bunch of things and little nuggets of that sit in your brain and then when you're on the trip they might come up - one place is abruptly closed, something looks like a bummer in person, etc. Without doing that research yourself then on the trip you have a list of things and aren't really knowledgable about alternatives.
Like I said, I'm not the target audience for a personal assistant. That does sound like a more ideal UX than what chat gippity and the others offer though. I would expect that with Meta - they really get the UX mostly right for people.
Sure, I get that. For this particular case, I've been to Vegas (family lives there) "plenty" so I don't really care much about stumbling upon stuff like you pointed out. I just wanted to have some places on hand to suggest for dinner and things to do when the group when we'd inevitably get to the "so can someone decide what are we doing for dinner???" part of the day. For planning a long trip to say Tokyo (never been) for vacation I'd do the regular stuff searching Google and reading blogs, but I'd also add AI search to it, if anything to find the current new "viral/trendy" stuff that I absolutely do not keep up with.
I've been "forced" to use AI in my work since this year, but I only started experimenting with it with my own ChatGPT subscription recently. I'm able to afford $20/mo and I'm fine paying it. The average user meta is targeting wants a free assistant to replace Siri, which is what muse is.
Continuing my thought, talking to Muse and ChatGPT feel fundamentally different. The single thread chat Muse offers encourages you to talk to it like a person imo. But that's probably just me.
Yes we do care and many of us have set the privacy settings to certain things over the years
But this is where friends are and don't give me 'you could just call them'
Not everyone has access to email/phone. Facebook should work in theory, if they weren't so enthusiastic about intentionally eroding privacy (physically changing the privacy options). Multicast. I post some status to multiple people at once. I can find people I don't already know/have lost touch with, and connect with them
What Meta does, does not mean its users don't care and I wish you people would stop saying that
I think we're a few decades away, at least, from seeing a robot on something like a residential construction site. Sure they look good in a flat controlled environment, but humans have a hard time walking around a construction site. Then there is the dexterity issues that they have an human hands excel at. The startup cost of a robot is probably 4x the cost of an entry level tradesperson, at least.
Right now, today, modular homes are factory built, transported and site assembled (accurately lowered into place and levelled on biscuits to hydraulic jack together) that are multiple large Lego blocks of click together cast concrete foundation with utility piping inbuilt and "some part of a larger house" fully built above.
These can be trucked or flown into place.
Factory built simplifies full trade critical path, with pipeline schedule of custom variations moving through to be wired, plumbed, built out and finished.
This topic is similar to jigs in woodworking - you can get all fancy and over-engineer it but in the end, does it improve the product/output? If a fancy jig and a gross looking home-made jig both produce a 90 degree corner, what is the added benefit?
I think people put way too much emphasis on things like this with conclusions that they are faster and therefore its better. There is Pareto's law and a point of no return on invested time with this type of endeavor. I would feel better if people started off with "I did this because I can and I like doing things like this" vs. some argument that its better in some way.
Some of these stories are similar in nature to people escaping once they realize whats actually going on. Alignment to a company's mission is good but it shouldn't be followed like a religion.
Why do people working in tech consistently get disillusioned into some company's mission statement or the equivalent? Its easy to just say the simplest reason is money, but this has been going on for decades though. You don't see the same attraction to adult entertainment (gambling, video, etc.) software jobs so there is obviously a line a lot of people won't cross. Those industries are at least honest about what they do, its not hidden behind some mission statement.
By all indications the shallowest reasoning is once someone can "cash out" thats when their values matter more. Maybe there is an element of maturity that happens after working for 5+ years that kicks in? Maybe it really is achieving FU money? It would be interesting to hear honest accounts from people that went through that cycle across more industries than AI.
I saw a headline that Muse was leaking people's contact info. There was a DEFCON presentation in 2025 on how they (hackers) coerced private/personal data out of a model (openai maybe). That has been patched most likely.
Meta doesn't have a great track record of protecting data.
I don’t think it’d be related to this if the defcon presentation was in 2025, but your description reminds me of this HN post about getting Claude to leak memories: https://news.ycombinator.com/item?id=48916975
Whats the use case where this, or other personal assistant/agents, are helpful and not gimmicky for average consumers (i.e. not tech nerds like all of us)? Say a person that is a car salesperson at a dealer? Or a teacher?
What would they use this for in their personal life? What does it unlock that they can't do today? Scheduling things? Buying things? These are so easy now for nearly anything. Reminders? phones have apps for that, computers do too. People can already dictate all sorts of things via speech (siri, etc.)
I totally get the professional use cases, I just don't see B2C other than gimmicky things. So I think that means Muse/Meta wins this out, people using that have already succumbed to giving up their data for "free".
These are not my words but someone mentioned these (probably hard) use cases that I think make sense for these agents but they're not there yet:
- Regularly get insurance (car/home) quotes
- Negotiate medical bills / help understand bills in general
- Analyze credit card spending / sort of like mint.com experience
- Advice on financial products (e.g. use money market for X sum vs having all your money in 0% checking) - sort of like a financial advisor
- Taxes and rebates, deal with IRS
- Reach out to Utility or whatever companies and file complains (saw this on Reddit where some guy's personal bot reached out to Verizon to complain about loose pole wires)
So it's practical time saving use cases...not a restaurant reservation or remind me about X or help me clean my mailbox BS cases. That stuff if for techies like us is useless but you have to start somewhere
My specific power tool advice would be to skip the miter saw, circular saw and possibly table saw and get a decent track saw for $500 US. Its a circular saw, you can also use it for angles, maybe with proper planning doing compound ones with a couple cuts, I've never tried.
Also if you set aside a certain budget for hobbies, etc. don't feel guilty about spending on it, go with the philosophy that is guilt-free spending. If you have fun and then lose interest, as happens with hobbies, you can sell the tools. But also see what you can get done with a small set of tools and see what you like first. Taking classes helps narrow in on what you like doing. I learned I like the mystique of hand tools but don't have the patience to do that 100%.
Agree. I got a Makita track saw and a large sheet of that 1" thick foam insulation they sell in big-box stores.
Lay the foam down on a garage floor, lay your plywood (what-have-you) on the foam, lay your track across the wood (where you want the cut) and go.
As long as you make sure the depth of your cut doesn't pass all the way through the foam, you're good. (And a 1" margin for error should keep you from sparking the saw on the concrete floor.)
If you're clever, cut the 4' x 8' sheet of foam (with a long X-Acto knife) into 4 pieces: 4' x 2' each. Duct tape pairs of them back together and you can fold up your foam like a gym mat. Note: make sure the tape is on the floor-side, you don't want to cut through duct-tape.
Add a portable drill & driver combo and your woodworking journey has begun.
(Me, I added a jointer, planer, table saw, routers, built a workbench, etc. You need a lot of clamps. Got rid of my miter saw—kinda dangerous and I just didn't need it. To be sure the table saw is dangerous too, but the miter saw is stealthily so—seems undangerous for the longest time and then‚ bam!)
There are specialty chainsaws used for cutting mortises and so forth when timber framing, and some folks will rough up beams using just a chainsaw, to say nothing of the expedient lumber/board cutting jigs/fixtures.
Maybe people smarter than me know better but couldn't there be a middle ground where an IDE/Editor has an embedded engine (doesn't need to be a full-on LLM) that doesn't require external tool calls and token spend?
If an organization is paying $2400/year per developer for tokens and a highly intelligent editor/IDE comes around that charges $1000/yr and gets more output at a fixed cost, its a no-brainer of a decision.
reply