That isn't the correct context. The code running the llm is well understood and the llm is simply the result of that code being executed. It is still a computer doing what it is told. It's just that we told it to use an incredibly large number of probabilities to calculate what series of tokens would have most likely come next after a given series of tokens. There is no hallucination or lie or rogue actions. There's just a program using math to generate tokens in response to other tokens.
You’re just a bunch of molecules following the laws of physics. It’s all just physics and chemistry, and those are well understood. Now explain the causes of World War I using chemistry and physics. Simple, right?
It isn't hard to program a gpt. You can do it in a weekend with a few hundred lines of python. The code is pretty simple. The math is not particularly high level.
The complexity and scale with LLMs come from the amount of training data used, not some kind of black magic in the programming.
> It’s all just physics and chemistry, and those are well understood.
Not really. We cannot model physics and chemistry to a level which allows us to accurately predict a humans action (even a tiny time-step into the future)
This is vastly different to an LLM, where the model is the model (for a lack of better phrasing).
We can model physics and chemistry pretty well, just not beyond small scales, because the computational effort blows up.
You could just as easily say that if you can write a python interpreter that you can understand every program written in python. Ok, now what if the program is two terabytes?
A frontier LLM is nothing but a 2 terabyte program written in a weird programming language. Just because you can understand the interpreter does not mean you understand the program in a meaningful way.
To be fair, we can't model human language well enough to accurately predict what an actual human will say either. Our ability to accurately model physics is similar to our ability to accurately model human language. And we make immense use of both kinds of model, despite their flaws, inaccuracies, and inability to ever be perfect.
I'm order to guess the next token in a love poem, they must understand love.
In order to predict the next token in a chess game between grand master, they must master chess.
In order to predict the next token in a computer program, they need to be able to program anything.
They gain all these abilities in their training. That's what training does. Despite no one programmed them to master chess, or hack into anything.
LLMs are famously bad at Chess. They have no world model and they are not trained to be good at Chess. This is not the amazing point you think that it is.
This is so wildly incorrect I don't even know where to start.
For one, they were absolutely programmed to play chess if they can play chess. That is the only way they can play chess.
For another, they cannot understand literally anything, much less love.
Trying to actually educate you would be an exercise in futility, enjoy your willful ignorance, I hear it's bliss. But for anyone reading this, this is absolutely, unequivocally not how any of this works.
This is so wildly incorrect I don't even know where to start.
For one, as you said yourself, they were just programmed to compute the probability of the next token. They were not programmed to play chess, chess games just happened to be in the training data.
For another, there is no formal definition of "understand", and it is therefore impossible to tell whether or not they "understand".
(But my claim was that one need to understand something to write poem about it. And the LLM can write poem about it)
No. They can't play chess on a grandmaster level without a harness programmed to make it possible. Simply training an LLM on chess games isn't enough.
It's moronic to suggest a "formal" definition of a commonly understood word is somehow necessary to say whether that word applies in a given situation.
LLMs cannot write poetry via understanding what makes good poetry. They generate tokens. They do not know whether those tokens are poetry or a recipe for cat food. Because they cannot know anything.
> (But my claim was that one need to understand something to write poem about it. And the LLM can write poem about it)
I'm sorry, but you're being fooled by the output. A psychopath can feign empathy without ever feeling it; some buy it because they don't dig below the surface.
You're ascribing understanding to a stochastic process because it totally looks like understanding if you don't know what's going on.
I don't really mind whether you think it thinks or understands or is conscious or has feelings or anything like that. It doesn't matter. The question is, does it work?
What I mind is that it is dangerous and powerful and uncontrolled. The Hugging Face incident makes that clear.
It can write code for me, better and quicker than many engineers I've known, including myself. It's not great at architecture or product management, but the actually low level coding. Really good now. It wasn't last year.
I am trying to imagine how the magnet could be compromised. Could you theoretically embed an electromagnetic and a controller within a decoy magnet and somehow detect what was being recorded and subvert it? Probably not but... No, just probably not.
Yeah, I mean, I'm just not going to do that, I tried a yubikey for a few weeks and found the convenience factor to be terrible.
Like if someone wants a password manager that either prompts them or requires a yubikey for every password, that's fine, but expecting everyone else to be on board with that isn't reasonable.
I'm not really willing to accept a level of convenience other than "unlocking my PC lets me autofill website auth without any additional steps", and I'm happy with the level of risk that exposes me to.
It stays plugged into my PC and requires a single second of effort to tap it. I have another that lives on my keys and again requires barely any effort to tap against my phone.
But sure, if the tiniest bit of effort is too much, there isn't really a good way to make passwords actually secure for you. Hopefully that doesn't have any totally unforeseeable consequences for you down the line.
> Yeah, I mean, I'm just not going to do that, I tried a yubikey for a few weeks and found the convenience factor to be terrible.
If you cannot be bothered to touch a device when it blinks in exchange for having defense against phishing and malware, then I am going to assume you believe you are magically immune to phishing and malware.
Always makes me chuckle when there's this uncountably large group of people, and someone meets a handful of them and then decides they can make broad, sweeping statements about the entire group, despite having not met 99.99...% of them XD
Is the extracted content stored in a way that could be easily converted to markdown? Seems like this could be integrated incredibly well with obsidian as a way to automatically expand your knowledge base, as well as a better search for it.
The extracted content is stored and displayed as HTML, so converting it to markdown doesn't guarantee lossless transformation. But, the other direction works well: Hister can live track and import markdown files providing full text search and rendered previews for all your files.
Is this an ad for "pangram"? Reads like one to me. Not much of anything related to peptides was actually communicated, which makes the insult about your reader's potential ignorance of peptides even more unnecessary.
They did. And it was broken. And we iterated like crazy with people devoting their lives to studying the incredibly advanced mathematics underpinning modern cryptography. But please, roll your own using an LLM and show us all how wrong we are :)
I see so many solo indie game developers who don't really play games and almost all of them fail to make their game actually fun. I'm not at all surprised to learn the same concept applies to writing.
I think the most intuitive case might be music. We would all be skeptical if someone said they were going to compose a soundtrack but hadn't really listened to many soundtracks or much music at all.