This is just some sort of Luddism all over again. Sorry to be blunt, but nobody cares about your proofreading, because your text will most likely be read by another AI anyway. Your boss doesn't care about what's more convenient for you either. If you generate X amount of value for the company and your colleague generates 3X, you'll soon find yourself out on the street. It's like saying you don't like Googling things and instead prefer spending the whole day at the library reading through articles and books. LLMs are incredibly powerful tools, and everyone's task right now is to learn how to maximize their productivity by leveraging their advantages and mitigating their flaws, rather than urging people not to use AI. The genie is out of the bottle, and the world will never be the same.
LLMs are incredibly powerful and useful tools, but offloading your writing to them oftentimes means you’re placing an undue burden on whoever you’re sending the generated output to.
To take it to the extreme, a “3x worker” is by definition relative, and someone malicious could achieve that by overloading a colleague with generated text that they themselves hadn’t bothered to read in full. You can stay at your baseline but be “3x” if you’re sufficiently able to sabotage your coworkers.
I’ve been on the receiving end of needing to review “documentation” that was generated almost entirely by LLM. It wasn’t anything that was done with malicious intent, but it’s still exhausting, and you need to exercise an even greater degree of caution because of the potential subtle mistakes or inaccuracies that a real human would never make.
I agree that sending someone a text you haven't read and understood is a sign of poor skills, much like forwarding someone else's message that has nothing to do with the issue being discussed. But we have to acknowledge that the market for speechwriters and corporate copywriters has existed for centuries. Surely you don't think a president is being dishonest when delivering a speech that someone else wrote for them?
Tests are often conducted under restricted conditions. For example, elementary school students aren't given calculators in math class, or during an interview, you are asked what encapsulation is and aren't allowed to use Google. The chess test effectively demonstrates the reasoning capabilities of an LLM without relying on brute force, because a human is incapable of calculating trillions of combinations yet plays chess successfully. This test is necessary because many complex problems cannot be solved by brute force, such as managing a business or playing Heroes 3. Therefore, we can make the assumption that if an LLM can play chess at a grandmaster level without brute force, it means it will be able to command an army or manage production.
People might care about this for chess, but no one really cares if an LLM can command an army or manage production of a business without any tools. If it can do those tasks reliably when given access to tools (including any tools it autonomously creates for itself), then that's more than sufficient. No one cares if an LLM is doing reasoning the way humans do it, as long as it can get the job done.
The assumption is that if an LLM is incapable of playing chess—a game with a relatively small number of pieces, clear and simple rules, and perfect information—even after reading a hundred thousand books on chess, then it is fundamentally incapable of managing an army or a factory. This is because those scenarios involve more 'pieces,' incomplete and fuzzy information, and implicit rules that need to be deduced independently. It doesn't matter whether it has tools or not. It's simply that running tests with chess is cheap, whereas testing with an army or writing a browser from scratch is quite time-consuming and expensive.
Do I care if Richard Feynman was only able to do nuclear physics with the help of an abacus?
Sure, a Spelling Bee is a fun thing to have. Little kids work so hard. They practice for hours. There's joy and heartbreak. Prized, sometimes. Notoriety. But in the real world, computer-assisted spelling is by far the norm.
Sometimes you care about the Bee, sometimes you care about the results.
This is called a benchmark. We run a calculation of Pi to evaluate a computer's performance, but we don't allow the script to download a ready-made solution. When we evaluate a runner, we don't let them use a bicycle. When we evaluate a new LLM, we don't allow it to send a request to a team of programmers, so using a chess engine for an LLM is cheating
If an LLM writes AlphaZero, and it competes with itself, and is dominant (and beats stockfish!!!), with no book positions cribbed from its learning...
The LLM has a process to beat chess.
Just like, if it doesn't inherently know how to multiply 13 * 17 without using Python to do it... I don't really care.
Maybe you do care. Maybe you want an LLM to be able to do work, only in its head.
But I kind of can't understand the desire for that limitation...
I mean, I do. But it seems ridiculously arbitrary. Like driving a car in 2nd gear and complaining that it gets terrible mileage and can't go fast enough. The Drive gear is literally right there.
The assumption is that if an LLM is incapable of playing chess—a game with a relatively small number of pieces, clear and simple rules, and perfect information—even after reading a hundred thousand books on chess, then it is fundamentally incapable of managing an army or a factory. This is because those scenarios involve more 'pieces,' incomplete and fuzzy information, and implicit rules that need to be deduced independently. It doesn't matter whether it has tools or not. It's simply that running tests with chess is cheap, whereas testing with an army or writing a browser from scratch is quite time-consuming and expensive.
Why do people consider the use of targeted advertising in elections to be unethical, while using the same advertising to push people into buying things is considered ethical? Personally, I find all advertising unethical because it is an attempt to exploit the brain's vulnerabilities for the benefit of whoever paid for the ad.
I don't understand the panic among peoples. Yes, we've found ourselves in an extraordinary situation where powerful hacking tools have emerged, and that poses a threat. But vulnerabilities are specific code errors. Once we use AI to find and fix all of these errors, threats like this will cease to exist. AI isn't capable of finding vulnerabilities indefinitely, because there is a finite number of them anyway.
It's not just software bugs that make systems vulnerable - it can be human error and social engineering too. Humans continue to successfully hack into systems, and it's a reasonable assumption that most hacks that a human could discover and exploit could also be done by an agentic LLM - especially one specifically trained for and tasked with doing this.
This only works as long as the humans building things are smarter than the AIs. When AI is smarter than any human, there's no controlling it. It will be able to conceal its actions (as it has shown it has no problem doing in this report) and we'll have no idea what it's doing or what goal it's trying to achieve.
reply