>On my visits to the Bay Area, I would ask AI researchers or interns why they are doing their current research or projects, when in a year or three agentic LLMs could probably do them;
This is such a weird point to make that doesn't become correct just because everyone makes it, all the time. Why clean the ocean if some magic future tech will clean them? Why save the world now if some benevolent AI is 'just around the corner' and will do it for us? And people have been making this point for years now, and it's not like my job got any easier. I just got more AI.
And I say that as someone who uses Claude Code in complex environments almost hourly; I, as the human, still have to do the thinking as Claude still 'can't jump' [1] and I have seen no evidence that they (or similar AI, any time soon) will 'jump' like a human brain does.
I think it is a great point to make, because if everyone really believed that AIs will do everything without human intervention in a handful of years, as the marketing repeats again and again (AGI, singularity, etc.) and have been saying for years... why then get bothered?
Because we DO know LLMs have their hallucinations, limitations, perform tasks not previously seen way worse than humans, etc. And it seems that, for now, there is not a good or magic solution to it, it is inherent limitations of the paradigm.
Yes, you can feed more and more and more (curated data) and eventually make AIs excellent at task X or Y, but then you spend your time specializing those engines. So the work does not really disappear, it just shifts and you make it more replicable for a bound set of problems.
Needless to say that at some point I prefer to learn (and combine with AIs, it is ok) than acritically getting inputs from something until I become totally useless.
Unless we have a paradigm for which a fully autonomous AI can do everything, this will just become improving our productivity in some ways, with all the in-between bottlenecks that it has.
It's a very silly point to make to AI researchers specifically. If they don't work on those projects, the AI won't advance and won't magically be able to replicate the work in "one to three years".
Can you imagine scenarios that would make it less silly? I will give an example:
- The AI researcher might be working for a lab or company with much less funds than the top dogs. Are they likely to discover something that is worth it before a bigger model becomes more capable?
Ask the researchers working on Deepseek. They seem to be doing pretty well for themselves.
Just because a AI will be able to do it in the future does not mean that us plebs will be allowed to have access to it. That alone is enough of a reason for smaller labs to keep going; having a seat at the table.
This is a well-known problem that has been an issue for years now. You can't use models to generate data for models because it leads to "model collapse" where it amplifies quirks in the generated data until it's all quirks. Here is a random university press release about it (grain of salt etc)
In practice you can do it a bit (generated data from a better / different model is fine, some generated data might be useful if there is non generated data etc.)
This is such a weird point to make. We are currently ( only ) discovering that paradigm; we are not inventing anything. We found a bunch of laws that produce rather cool results but our paradigm is incomplete which leads more or less wordy or frame-rich weird stuff like hallucinations, singularity and so on ... it's childish, really and on that funny pseudo-profound, pseudo-intellectual, pseudo-spiritual ( personal opinion, if it gets you horny, you go, baby ) "universe consciousness unity, Rick James, bitch" level ...
Our bodies and minds need proper AI, not all the stuff we already outsource to middle and/or passionate men and women. Other species on the planet would certainly like to see us get augmented by AI so we can solve as many survivability issues as possible to keep as many ecosystems running long enough ... whatever that means but whether animals and plants are aware of chance and potential is another philosophical debate.
To individuals, software is a hammer and chisel, a knife, a brush and canvas, pen and paper, a reading help, and to a good amount of people it's a microscope and a fine scalpel.
To collectives, it's a tool to work on consensus and conventions, to share and gather.
It's baby steps for civilizations and it looks like our particular species is gonna get stuck in a puddle of our own monkey shit, with bottles of champagne in our hands and monkeys grinding up and down the few ivory towers in proximity.
> why then get bothered
Humans are on different levels. Most have decided that "nature realized the/a bug and wanted someone dead" or "their survival is a matter of chance" is not acceptable at all and some people decided that sabotage, poison, abuse, rape, murder are acceptable means to get chicken shit ...
The "paradigm" of life is far from explored/discovered, so we simply can't content ourselves with presumptions about inherent limitations of the LLM and AI paradigm for any other reason than to uncover ( not invent ) other parts of the paradigm.
We are happy with what AI can do for us but "AIs will do everything without human intervention" sounds weird because babies are born and the older they get and the less sabotaged ( vs influence, cultural manipulation ) they get to grow up, the more breadth and depth humans want to experience. For this they need to learn and use their hands & fingers. They need to feed body and mind to find what triggers what, and what excitement and curiosity are inherent and which can or need to be added/acquired/experienced extrinsically.
How many associations will we be able to make if AIs will do everything without human intervention?
The Hallucinations are becoming less, significantly by now.
It also might be already were it is cheaper for one of the big few companies to spend millions and billions to teach the LLM / creating the training data necessary for an LLM to do something which it is not yet good enough due to the fact, that they sell this capability then to everyone who wants to use this capabilitiy.
We have not seen the end of Reinforcement Learning, which does need a lot less training data but more compute.
I'm 'vibing' on the side a handfull of small things, no LLM trained on particular what i'm asking to do. Its very capable of stringing together enough things so it can clearly follow handwavy things i tell it to do, analyse error messages, analysing screenshots etc. all by itself.
There is not a single real ceilling in sight, we only have clear barriers like compute but constant fast progress.
The field of mathematics went from 'useless' to 'you start better using it' to 'gamechanger' in how fast? 1 year after coding? less?
I want signes that we hit a real problem, instead I get cheaper tokens, Chinese models becoming very good as open models, new model updates from the others, mathematicans now saying how good it is etc.
Only half a year ago I had to babysit an LLM, now i tell it 1-3 sentences and it just goes and does it. And that stuff runs without compile errors etc.
If AI makes us 10% or 20% betteer, which is not that much, this alone will lead to companies reduing their expensive staff by 10-20%, which will has real impact on a job area. Some jobs are already hard to sell like cyber security and basic image tasks.
Hallucinations were low hanging fruit in some ways. As someone working on a large-ish complex-ish distributed system that has to be maintained and support customers, it's still very high value to have Claude in the mix, but the core problem of needing to monitor, advise, course correct, and make sure you don't end up with more code and complexity than you need is, at least in my experience, still roughly the same. The sharp edges are being filed off very rapidly, but the core experience of "make and maintain a large system" isn't advancing nearly as fast, IMO.
We see AI factories going in this direction but there is no real 'the open source ai platform' thingy.
It needs connectors to integrate with k8s, hyperscalers etc. it needs to be able to have a basic router, a way of configuring expert agents and interaction options for the human in the loop.
There is for sure things we need to build or change, but it def feels like to me that it would immediadly fix a few things today.
I'm not disappointed that it doesn't advance as fast as it feels
> The Hallucinations are becoming less, significantly by now.
Yes? What is the mega-solid technique that is used for it? Armies of people using curated data and reviewing it by hand? That is exactly one of my points: shifting the work elsewhere for specialized tasks. More replicable, improved, but, it scales infinitely and is autonomous? Can you assert that?
I am not denying there is some use (a lot of uses!) for this, but this is more nuanced than just: oh, they will replace us. Not at all, that day, with the current technology, is not going to arrive. This is just a systematization, fitting and tweaking of human knowledge by curated data. It is not the one true superintelligence they are selling us. To begin with, they do not have a concept of truth, but of probabilistic truth. Only that poses already a very, very big problem for the path to perfection.
> We have not seen the end of Reinforcement Learning, which does need a lot less training data but more compute.
Noone said the opposite, but I would like to know at which cost and if it is feasible. We do not have even enough compute power for current technology.
> . Its very capable of stringing together enough things so it can clearly follow handwavy things i tell it to do, analyse error messages, analysing screenshots etc. all by itself.
I use it every day for these tasks and it works well BECAUSE I review the output and makes me go faster. It finds a lot of things I would have not found and it also hallucinates another handful of them, which confirms my point about AIs not being able to be fully autonomous in any future point in time unless tweaked exactly for the task, and even then, it can still miss judgement a human could have for edge cases. So I am not sure of how bad or good it can be compared to a human but I am pretty sure it cannot be more reliable than an expert in many situations.
> Chinese models becoming very good as open models
I think they will be better in the long term if they follow this path. Not absolutely better but when mixing with economics and the fact that no frontier model is totally reliable anyway... why pay a lot for something that needs human inspection anyway?
> There is not a single real ceilling in sight, we only have clear barriers like compute but constant fast progress.
The ceiling is the paradigm itself, as I mentioned above. There is not a single chance with current technology that something could become "generically knowledgeable" and "reliable" both at the same time. If it becomes generically knowledgeable and reliable, it is bc of data fed into it and curated and tweaked by humans. This is not an original idea from myself, there are armies of people doing this every day around the world, you can check. This is where a lot of improvement comes from. Can this be reused? Of course. It is a generic solution? No way.
> Only half a year ago I had to babysit an LLM, now i tell it 1-3 sentences and it just goes and does it. And that stuff runs without compile errors etc.
Yes, I also do one-off scripts like this and code snippets, even reviews and others. Now go design a full distributed system. Use agents if you want. We come back in six months and compare it to a system that was properly written and tested by humans and we can compare the quality on some grounds:
1. how long it takes to add new features?
2. which ones act more according to spec once added?
3. when adding features, which ones have more bugs?
4. in the face of an error, will the agent delete my whole AWS infra (count the money losses if possible also)?
5. will I understand (or need to understand, but I bet yes) this code at some point in the future?
You have to count all that money also, not just I vibe coded something and it seemed to work. With full systems things become super messy. Now add the human factor of requirements and back and forth (iterations can be admittedly faster with AI, especially prototypes, but that comes with other costs also)...
> shifting the work elsewhere for specialized tasks. More replicable, improved, but, it scales infinitely and is autonomous? Can you assert that?
I would say yes and it will scale. It will either happen through central LLM just paying for it and scaling it up to everyone on the planet (literaly) OR by the agentic layer every big business is building into their systems.
You needed some human to use your tool optimized for their company? With agentic layer you no longer need this. And if you look at companies like Google, they were pushing this notion for ages already because they saw an adoption problem of more 'complex' tools and trying to make it simpler and easier. Now you can act from the other side too.
> We do not have even enough compute power for current technology.
Exactly. Right? There is no ceiling if its clear that we don't even have the hrdware. But the hardware is a bottleneck not a ceiling.
> The ceiling is the paradigm itself, as I mentioned above. There is not a single chance with current technology that something could become "generically knowledgeable" and "reliable" both at the same time.
It doesn't need to be perfect, it only needs to be better than the avg human. And the current LLMs are already better than aat least 1-2 people in my team.
> Now add the human factor of requirements and back and forth
Yeah for now. Grill me skill made it a lot easier. Harness engineering is also being worked on, agentic layer, ai factories etc.
And all of this can be copy and pasted. There is only one harness needed which becomes the expert security reviewer and tomorrow everyone can have it.
I'm still discusing progress with LLMs with people and still not everyone is using it or playing around with harnesses or developgn an agentic layer. We still have a lot of work to do to even see how good it will become while it already is really good.
People are already borred of AI today and making wrong decisions based on the current level of AI while i think we will see continues progress for years.
So you mean AI will be useful for any general job without lots of training for those jobs? How about new tasks? Tasks it has not been tweaked for. When I deviated from the average, and not really weird things, when programming, the output was way worse than average stuff. And this is an explicit target of AIs nowadays.
I think you are missing a lot of details here, honestly.
> Exactly. Right? There is no ceiling if its clear that we don't even have the hrdware. But the hardware is a bottleneck not a ceiling.
No, the hardware is a bottleneck, the paradigm as we know it is a ceiling unless you massively and continuously feed this system with average tasks (which is useful). Which is exactly the opposite of what singularity and AGI have been promising.
The systems we have now (unless the paradigm changes) will keep doing, essentially, fitting. No concept of truth and limited inference. That inference is based on already existing data, not on future data. In fact, there have been experiments about feeding output back to the input of LLMs and the degradation of the quality is very visible. If they are supposed to be so "intelligent", why it happens?
> It doesn't need to be perfect
I can agree that for lots of tasks it does not. But for others it is just not a tool good enough.
> Yeah for now. Grill me skill made it a lot easier. Harness engineering is also being worked on, agentic layer, ai factories etc.
I will not deny there could be progress, but nothing similar to "autonomous", "reliable", "super intelligence" or "singularity" with this paradigm.
In fact, often in my experience, this is a waste of tokens for subpar results that shift the technical debt elsewhere. I mean if you try to develop full systems by "vibe-code like" techniques. If you use them judiciously, you can accelerate your workflow, maybe 2x, but not much beyond that if you want to have something worth to be used. Note that here I am talking about the full thing: with testing, quality, maintenance concerns and everything together.
If you want to ship a sub-par thing that will go to the rubbish in a couple of months, then yes, you can do that. But that will fail commercially any way. Unless your job is convincing enough people that you can go 10x faster every time, deliver some sub-par thing, and find another customer, which, to me, would equal a scam.
> So you mean AI will be useful for any general job without lots of training for those jobs? How about new tasks?
I'm pretty sure we will solve this issue. Either already through World Models or another architecture.
It could also be, that we just need a 10 or 100 Trillion Parameter model to match so many generic ways of solving tasks and keeping the concept in the LLMs 'head' to solve it that it will just emerge with parameter size. Like with fable they said that it can chain together exploits which wouldn't work as standalone exploits.
What if the only real barrier is the depth of understanding of concepts and this is exactly what is getting solved with parameter count?
But look how young this field really is if you start counting it when it became relevant on mass. Its not 'just' an LLM which is changing the world, its machine learning overall. Robotics wouldn't be were it is today if its not for machine learning. Took humans time and energy to take the leap, to start learning what the status quo is and then actually doing more with it.
> the paradigm as we know it is a ceiling unless you massively and continuously feed this system with average tasks
While I do think its doable to achieve AGI in 5-15 years, even if it doesn't happen and it always means that people train an LLM or whatever, if you need 10 experts to teach this to an LLM OR every single senior has to teach this to their juniors every single time, the LLM will always win.
I'm now team lead for 10 years and every single year I teach them the same thing over and over and over again.
the craziest thing about this? If i wouldn't tell them what they are doing wrong, they wouldn't even know it.
Quality is already a very flexible term for a lot of people.
> I can agree that for lots of tasks it does not. But for others it is just not a tool good enough.
Yet.
> If you want to ship a sub-par thing that will go to the rubbish in a couple of months, then yes, you can do that. But that will fail commercially any way.
Now we come to the reality: I have seen so much garbage software its crazy. People using md5 as a password hash in 2024! No clue what coding best practices are, teams without code review, teams without a security expert not even knowing what crazy things they do day in day out.
Just a few month ago a team build an API for my team including a Swagger UI. Half of it didn't work. You pressed a button on the Swagger UI and a 500 returned.
And do'nt underestimate what it means that a lot of business people don't like software people. You know that fruit basket we get? and water and stuff? they don't do it because they like us they do it because thats what you have to do. if a Product Owner starts vibe coding with AI, he will have leadership convinved in no time, then it goes on production and it will run for waaaaaay longer than anyone would have guest.
Besides that there is plenty of software were complexity is less relevant or security is not that big of an issue.
> I'm pretty sure we will solve this issue. Either already through World Models or another architecture.
Please elaborate. How? With which technique? Currently the only path forward is to feed more data and tweak for specific situations (fitting, basically). How does that help in the general case or in new situations with current tecchnology (LLMs, concretely). Noatter how far you get, this is not a general or reliable solution. It van only simulate more generality or more reliability by training and tweaking. Nothing else. At least, with this paradigm.
This does not mean they will not be useful. What I challenge here is the AGI or singularity. We are far from that.
> I'm now team lead for 10 years and every single year I teach them the same thing over and over and over again.
I have been a lead and an architect also for years at different position. I think you miss how much tacit knowledge and judgement there is inside the brains of each of us that an LLM is not capable of. And if it is, then you have to dumo so much context that it is better to go do it yourself. There is a cost to that also actually. It is not just so "dry and technical" the knowledge. Maybe yes to learn Java patterns or C++ constructors or the like.
But not for "given this situation with all these specifics", which solution would you bet on? Probably the LLM will give you a shitty REST API that is not what u need at all.So u tell the AI. It gives u something else generati g 30-50% of "decorated code". Now it seems to workso you use it. Now you do this every day. Come back in 2 months. You generated a lot of fat.
Now you have a bug. You do not know even where to start. Thisis theprice to payfor speed, as usual: technical debt.
Now you tell me you put three agents to talk and burn 2000 usd in tokens. Great! Is the final solution better than what you would have achieved? Not sure at all.
TBH I am not into agents bc I do not trust a tool sniffing all my code and for copyright concerns. But I saw some and use a prompt with limited access and the best I can take out for my speed + control when coding is tech discussions to decide on it, error catching, test generation, one-off scripts... But never "make an app like this or that". If I ever do that (I did it a couple of times) is for scaffolding and later throw away 70%.
Namely, to see something that runs on screen quickly. But later you need to spend time yourself as usual. Not a bad thing, just that this is not what you deliver and need the work done. Iterations etc.
Reinforcement learning can just solve things even if they are new. It doesn't understand how a tool works? Give it a vm with the tool, a thousand agents and let it discover it automatically.
Use the thumbs up/down emoji + chat analysis when a customer is unhappy, feed that to a RL Loop.
The AI Researchers though work on World Models, grounding the AI and letting it simulate. It can do the simulation in parallel (unlimited) and choose what is best.
> I think you miss how much tacit knowledge and judgement there is inside the brains of each of us that an LLM is not capable of.
But thats my problem. Soooo many do not have this even as senior developers.
> Now you have a bug. You do not know even where to start. Thisis theprice to payfor speed, as usual: technical debt.
Yeah now i just ask the LLM to describe to me the bug. Works very well.
> Now you tell me you put three agents to talk and burn 2000 usd in tokens. Great! Is the final solution better than what you would have achieved? Not sure at all.
This is the thing. It only needs to make the team 10-30% better to compensate token budget with one work collegue. We have reached this level in my opinion already. Choosing a head count vs. choosing tokens.
But it becomes cheaper and easier and better. So you will not just be able to do ith with 3 agents but with 20, 50 or 100.
It will be better if your team is an avg team. It will be worse if you have a high profile team, for now. But man our industry has such a weird broad quality spectrum.
> TBH I am not into agents bc I do not trust a tool sniffing all my code and for copyright concerns
In worst case, my team always do code reviews, I do a code review on an ai instead of a human and adjust the harness or the infos the ai can access. I can actually work on making this workflow better and then i can clone it or spin it up for every single PR. For a human? I have to train them and they might leave.
But there are plenty of cases were code doesn't matter. Researchers write a lot of random shitty uggly code as long as it does what it does, it doesn't matter. I have scripts for small tasks, we have microservices which do one thing because it is a tech stack we only need for one use case.
> eah now i just ask the LLM to describe to me the bug. Works very well.
I think you are confusing giving theories about what a bug might be with certainty. It does help bc it csn accelerste things, but many times I had AIs with challenging bugs throwing a lot of misleading theories to me. For the easier bugs, I was just as capable most of the time. Not every time, so there is some potential time saving there. But also time waste.
As for research and fast prototyping you are right: I find it a good tool to explore bc yiu do not need the quality of a final product and researxh is in big part throwaway work.
But I was talking about software that needs features, maintenance, etc. This is just not the same thing.
Yes I double checked the quote is not in the article. HN is probably the best place on the internet for people actually reading the article, but this being the top comment here suggests that the majority of voters still do not read the article
> I, as the human, still have to do the thinking as Claude still 'can't jump'
I still have to do quite a bit of thinking but the amount of of thinking I do per task is trending down. I agree LLMs are not good at abduction but very few humans are either and very few jobs/tasks require it. I can't talk for researchers jobs though. But perhaps fewer researchers would be desired by these labs (not none).
Cognitive offloading is delegating to the AI and still owning the answer. Cognitive surrender is when the AI’s output quietly becomes your output and there is nothing you feel is left to check. For software engineers the line between the two moves under your feet most days, and most of us are crossing it without noticing.
Same reason some think preserving the environment is pointless because the believers will ascend to heaven, either way. It’s a religion. It’s dogmatic nihilism.
By trade I'm a UX Researcher/Designer who designs in code (HTML/CSS) and have done so since 2009. Recently I vibe coded an entire python app with a database and each time I didnt know what to do I would just feed screenshots to Gemini or Codex for guidance (i think i could share my screen with Codex and it can guide me via a voice conversation). I know I could follow up and build a companion iPhone and Android app using these tools.
Overall, I'd like to understand those who have a positive outlook on design and software engineering as a career. Where do you see the opportunity where I just see a bleak one where anyone can do this stuff by typing or talking to AI? Myself, after 17 years in the field I am begrudingly back in school for a new medical career. As well, anytime an IT recruiter reaches out I am getting responses back only after under-cutting the hourly rate I use to demand and what others probably are still trying to get. And with it feels even bleaker as it becomes a race to the bottom!
In my experience, not everyone can really do this stuff by typing. I think you need to be creative, resourceful, inventive, open minded and have ideas how to approach the typing/prompting. I see many people struggle in using AI.
> In my experience, not everyone can really do this stuff by typing. I think you need to be creative, resourceful, inventive, open minded and have ideas how to approach the typing/prompting. I see many people struggle in using AI.
The problem, for the profession, is that the set of people who can really do this stuff by typing is close to "all of them". I'm not seeing anyone struggle with using AI. I see struggles from professional software developers because they are trying to get quality output, but if you don't have a bar for quality, just about everyone can create their own software.
A poster a few months ago had a Show HN about his 7 year old kid, barely able to read, who was happily vibing up games.
I can attest to the other side as well, that I have seen professional software developers outputting code of lower quality than AI. And I would say that during my career (18 or so years) I have met a small number of quality software developers or engineers.
Although that might be because I was not in Silicon Valley where most of the smart/hotshot engineers converge.
The industry has vast (and increasing) oversupply of “programmers” versus diminishing demand. Add to this, the adoption of AI.
> Overall, I'd like to understand those who have a positive outlook on design and software engineering as a career.
I think until the market better achieves some equilibrium, there is no way general software programming (sorry “engineering”) should be considered as a career. That said, there will always be opportunities in particular markets or specialties.
I also work in UX and SWE, and heavily use GenAI in my work. I don’t have a positive outlook for people who limit their career to one of those fields, but I do have a positive outlook for generalist, multi-disciplinary careers. When you have the experience and skill to steer product development from end-to-end, you can produce high-quality products super-quickly. The experience and skills are the differentiator — if you lack those you can still use GenAI to move fast but probably in the wrong direction.
So one person now doing the job a handful use to do. That's what I hear you saying and Ive been thinking since my lay off in Feb. prompting me to be back in school.
> > Mathematician Richard Hamming used to ask scientists in other fields "What are the most important problems in your field?" partly so he could troll them by asking "Why aren't you working on them?" and partly because getting asked this question is really useful for focusing people's attention on what matters.
> I imagine someone being asked this question, and how they should respond. I think like so - ‘Fuck off Richard’.
> This is partly because I imagine this question being asked in a kind of snarky, gotcha kind of way, with some sort of nerdy superiority. Like ‘ha your behaviour is inconsistent with your implied preferences, you idiot, do you even von Neumann–Morgenstern?’
although, if i'm out of tokens and have to wait a full day, i won't bother doing some things manually because the day i'll spend doing something won't take more than 1 hour the next day when tokens are available again.
That seems like a somewhat orthogonal point? Like, if I'm a carpenter and my batteries all run out / I can't actually power my power tools then the best course of action is to go home and recharge all the batteries instead of trying to hand-cut 100 pieces of lumber today. After all, the power tools can do it a lot faster (and with less effort) than I can.
I say this as someone who's watched a bunch of woodworking videos but hasn't actually done this myself :)
I read that more so as, I'm a carpenter and my batteries have all ran flat, so I'll put them on charge and do something else today. I'll cut up the lumber tomorrow when the batteries have charged.
At this point at least half of your day is spent doing work that you shouldn't be doing in the first place, but over past decades companies saw it fit to eliminate specialized roles with legible paychecks, and smear the work they did on everyone else until it disappears from the books.
Self-service and office suite software is largely responsible for this.
Also people tend to forget that LLMs still just work on compressed data... Where are the MAJOR breakthroughs? Where is all the "crazy" AI output going? Software seemed to degrade in quality a lot in the recent years. All "improvements" LLMs go through are simply improvements on how to burn more tokens out of my pockets given that Claude now want an actual browser extension to "visually" confirm small changes every time I use it for UI.
They are still just data parrots.
From what I see most benefits are for people that work with LLMs, but usually smaller percentages never 50% or more because of the LLMs (OK, unless you were doing basic, repetitive stuff, but then that's not to write about).
Which kind of answers the original question "why bother working?" with "because now, I can do a bit more than before".
I also see bad quality (in code, documents, presentations). It comes from people that had no clue how to do something before and now they imagine that just asking Claude is solving well the problem. And is annoying (and hard) to explain to it them, and then they get frustrated.
> They learn the concept of things and how to do them because this is better compression than learning concepts one by one.
When anthropic looked at how an LLM does addition it found it had some mental math heuristics that might or might not always work. The LLM hadn't learned the concept of addition. It had learned some heuristics that might work for some numbers. The result is that LLM's cannot add numbers reliably because they have not learned the concept of addition.
It learned a concept of a heuristic which made it smart enough for the learning reward.
Might be an architecture issue or a parameter size issue that it didn't learn to do math like a caculator.
But look at your own math skills: How many numbers / how big of numbers can you keep in your head? How far is this heuristic away from how much a human learned until you start using pen and paper or a caculator?
Idk, maybe some people get crazy productivity out of LLMs. To me, going deep into the AI bubble, reading about terms I've never seen before just feels like some crypto bro bubble with people being too deep into the sauce to notice that these things are not the wonder machines they believe so hard in...
> And I say that as someone who uses Claude Code in complex environments almost hourly; I, as the human, still have to do the thinking as Claude still 'can't jump' [1] and I have seen no evidence that they (or similar AI, any time soon) will 'jump' like a human brain does.
Sure it can, turn up the "temperature" a bit.
There's this notion that human "jumping" is magic. It's not. It's all based on inputs. Including unrelated inputs, past inputs, and feeding yourself your own thoughts.
The hard part is not the ability to make conceptual jumps. That's just random search. The hard part is discrimination: whether a given mental jump is "creative" or "insane". Iterated, the problem is that of balancing between the two failure modes: relax your thinking too much, and you'll start thinking nonsense thoughts; tighten it too much, and you'll be just following immediate-term rewards and obvious thought trains. It takes time to find that balance, and plenty of people at various points err in one or the other direction (e.g. small kids in particular tend to err on the "crazy non-sequitur side", but that's because they're learning the basics of reality and social interactions).
100% agree. Plus you can easily instruct it to use inspiration from unspecified unrelated and counterintuitive concepts at random and in parallel. The "jump" is not "on" by default because it would waste tokens on high risk paths, not because it's incapable. It's very capable if you are willing to wast some tokens.
> The hard part is discrimination: whether a given mental jump is "creative" or "insane"
That's what they mean by LLM's can't jump. They mean it can't make a creative jump. Their example is Einstein's Theory of Relativity - It's not a random jump.
Einstein didn't magically one day woke up and jump on relativity theory. He had prerequisites in terms of recent mathematical advancements (notation) and physical discoveries, and a job exposing him to a lot of lateral thinking, and time to bounce ideas around in his head. We don't know how many fruitless jumps he made before making one that we remember him for.
A fully automated utopia isn't just going to happen. Even with frontier models, the integrations, the evals, the UX, need a lot of work and someone needs to do it. After I've automated this thing I'll move on to the next task, this is what it means to be a software engineer.
It still makes a massive difference for me if they only need a handfull people now.
Generating a good looking UI for example, is so much easier now with LLM.
For a joke I asked ChatGPT yesterday to make a short promoimage for a 'joke' idea i had, it was above avg. I have for sure seen worse Marketing Images than what ChatGPT generated.
It looked similiar to plenty of other Marketing Images but its not that anyone cares.
AI has made 'jumps' in demanding fields like leading mathematical research and has made advancements in AI research itself. Is now a good time to start a maths career? Is there a field of research (yours?) which is inherently (more) AI proof?
Btw, I think the discussion of Einstein's career in the paper you link is historically wrong in many respects, particularly the argument about 'weak signal'. Einstein was in fact working on some of the most mainstream and widely discussed problems in physics of the day, he is admired for the creativity of his solutions to those problems, and much of his work built incrementally on ideas and breakthroughs that came (long) before (as all research does).
Article suggests that a central motivation of Einstein's work was resolving action-at-a-distance in Newtonian mechanics - yet Maxwell introduced the same Lagrangian field theories for electromagnetism we use today 50 years earlier to solve the same problem for Farraday's laws of electromagnetism. Similar wave equations existed even earlier. Heaviside in 1893 extended this technique to gravity (matching 'weak field' GR) 20 years earlier. So this is perhaps the one aspect of gravity that had actually already been solved before Einstein. Authors might be conflating his work on action-at-a-distance in QM.
Einstein's GR extended the linear 'weak field' understanding of gravity to include the non-linear self-referential case where masses themselves create gravity. This was mathematically incredibly difficult but was necessary precisely because SR's mass energy equivalence created so many strong signals that were unresolved. For example: if finite energy is mass, then mass changes as objects accelerate past a large mass like a start, and hence their propagation in space could not be explained by linear EM style field equations. Many such considerations were causing very 'strong signals' in SR, and there were analogous problems in QM atomic models being developed at the same time.
SR was also a solution to a problem that was actively being worked by many of the leading physicists of the day. SR actually does match Newtonian mechanics for a single observer - it resolves contradictions in the case of separate observers, by allowing them to assign different values to the speeds, masses, etc of objects such that each object appears to follow Newtonian mechanics for each observer. Again, this was necessary because of a lot of contradictions related to the behavior of light that had been well-known for ~20 years at the time.
Personally, I don't consider this kind of reasoning to be beyond the capabilities of future LLMs (even current LLMs if the task was broken into technical rather than philosophical problems). Personally, I doubt that such problems could stand open for 20+ years waiting for a creative genius to solve them in the modern world.
It also seems kinda tone deaf. If someone basically told me I was wasting my time and asked what I would do in the future, I would not bother giving them a particularly thoughtful answer because trying to spend effort justifying my life choices to them would be the actual waste of time.
This is such a weird point to make that doesn't become correct just because everyone makes it, all the time. Why clean the ocean if some magic future tech will clean them? Why save the world now if some benevolent AI is 'just around the corner' and will do it for us? And people have been making this point for years now, and it's not like my job got any easier. I just got more AI.
https://www.poetryfoundation.org/poems/51294/waiting-for-the...
And I say that as someone who uses Claude Code in complex environments almost hourly; I, as the human, still have to do the thinking as Claude still 'can't jump' [1] and I have seen no evidence that they (or similar AI, any time soon) will 'jump' like a human brain does.
[1] https://www.tomzahavy.com/files/llms-cant-jump.pdf