What is the strict definition of arbitrary you feel we haven’t reached? Right now the limit is scale and complexity, not domain or “type” of application.
> reliably able to entirely build and deploy arbitrary applications from a prompt
You're maybe thinking that we can build and deploy arbitrary "kinds" of application. Being able to "build and deploy arbitrary applications" would mean I could ask for any scale or complexity in my application.
Now that is moving the goalposts. The entire subthread has nothing to do with whether AI can match humans in any particular endeavour. It's simply about one user's past prediction about an AI capability.
> I don't think there'd be any controversy if the problem statement said "common applications"
I would claim "common applications" is also controversial because what is a "common application" depends insanely on the area in which you work. Even if you exclude some highly advanced scientific applications (because you don't consider these to be common), in many industrial sectors there exist applications that have grown over multiple decades, and which encode an insane amount of knowledge about the respective sector and its workflows; this is a central reason why these applications are so hard to replace.
Well, it depends on how you define "complexity." LLMs are totally rewriting the notions of what's hard for a computer but easy for a human. LLMs have really bad spatial reasoning. I would have to play around a bit, but I'm very certain you could come up with a prompt where an AI is incapable of properly generating a "simple" app that has some important layout constraints.
And just generally, anything that requires the LLM to understand something LLMs don't understand, it's going to fail. I'd hesitate to give an exact example without trying Astra/Fable but I'm sure they exist.
One-shot a minesweeper clone that isn't screwed up in an insane way. Bonus points if you choose a language/platform that LLMs aren't likely to have already seen a minesweeper clone done in/for already.
Right now, they can't build anything without help that isn't buggy in unintelligible ways. If you push the thing feature by feature, have a lot of tests and a lot of instrumenting, and you check that it isn't cheating or lying after every step, you can get a lot of work done.
Should people who missed out on formative life experiences all simply accept this? Decide they don’t deserve spiritual nourishment and trudge on with their mediocre lives? Self-actualize in a way you personally would find less cringe? We all get one spin, man.
Even if you’re right, Burning Man strikes me a symptom of a sick and hollow “normal” society.
I think the odds of a human accidentally creating a novel virus or bioweapon at the behest of a rouge AI are pretty small to be honest. That's a lot of manual labor to go "oops I didn't realize what this was!"
In the short term you're probably right. But after a couple of years, complacency will take over and more and more decision making will get ofloaded. I don't think it's out of the question before 2030.
Exactly. No way to tell what's behind this. Once waymo has total market dominance and exists in every city, it's hard to imagine them continuing to support local transit. The thing about Embrace, Extend, Extinguish is nobody suspects it during the Embrace phase because good motives and bad motives look exactly the same.
Anthropic in 2025: We can use our dominant market position to degrade the harness experiences of our competitors because they will never adopt CLAUDE.md
Anthropic in 2026: We are losing our market position. Users who adopted other harnesses have a degraded Claude Code experience because it doesn't recognize their AGENTS.md
Sounds like free market at work to me. I am just glad there is quite a lot of competition in a field that I would have assumed would have huge costs of entry
Building small models is a lower barrier. Think of something like to just train on a corpus of internal corporate data. I've see some small models that do this. It's like a super RAG thing. I think more of that will happen. Excited to see a SLM vendor emerge.
The walls are closing in and the president is gleefully lighting fires he has no intention of putting out. There's a reason they're rushing like mad to an IPO, but as we saw with OpenAI it's easier said than done when your business model is "Lose tons of money to eventually maybe dominate a market with the moat we don't have, but trust us AI is huge give us trillions."
You can get a cheap approximation of this with chromium browsers + Filesystem API and a PWA manifest for that native-ish feel. AI has reignited my interest in building based 100% on browser native features.
If you want to distribute apps to many users this would probably be a better solution. The main benefit of Capsule is the personal document like app. You create a capsule, move it to an USB stick, save it in your cloud storage and it's still the same document. Nothing has changed, not the UI, not the data. The document will stay the same.
If it’s web native, anyone can just use it. If it’s a capsule, they first need to download the capsule app to associate the file extension with the OS and have it launch, right?
That feels like a big barrier, almost like a Java runtime.
In defense of capsule, the Filesystem API is janky and has kind of shit UX. Less sophisticated users would probably prefer the UX of a purpose built tool.
I prefer worse UX for the sake of standards and zero install, but many people would prefer the opposite. I don't think Capsule is intended for the HN crowd.
reply