Google attempted to do this 15 years ago, they got sued and stopped. It turns out that tech companies occasionally do have to follow the law, you'd think people would be happier about that...
Wait, if you’re talking about the Google Books case, Google won. Maybe they made adjustments on how they served results but they certainly did not stop.
The Google Books settlement was originally going to make Google into a clearinghouse for scans of out-of-print books. The scans would have been available for individuals to purchase for a reasonable price, and libraries and institutions would have been able to subscribe to a service that would give patrons access to the full text of every book. This deal fell apart because some research libraries and authors argued this was anti-competitive, as anyone wanting to make a competing service would have to go through the same process as Google of settling a class action lawsuit. They instead wanted Congress to pass a law to free up the rights to orphaned books. No progress on this was subsequently made because nobody in Congress cares enough about the rights to out-of-print books to get legislation passed. The whole reason why they're out of print when ebooks and print-on-demand exist is that they won't get enough sales to make it worth the time and money to figure out who the royalties should go to. The result is that nobody outside Google gets to see the full Google Books scans.
> HathiTrust Digital Library is a large-scale collaborative repository of digital content from research libraries, administered by the University of Michigan. Its holdings include content digitized via Google Books and the Internet Archive digitization initiatives, as well as content digitized locally by libraries.
> Authors Guild, Inc. v. HathiTrust (2014) was a following case related to HathiTrust, a project by the libraries of the Big Ten Academic Alliance and the University of California systems that combined their digital library collections with those of Google's Book Search. The HathiTrust case differed in two primary factors which were raised by the plaintiffs: that for viewers with disabilities, they could view the scanned text through a screen reader to make it easier to read, and offering to print out the scans as replacement copies for members of the universities if they could verify their original copies were lost or damaged. Both uses were deemed also to be fair use by the Second Circuit.
> The subject of the copyright of orphan works – works that may still be under copyright but with no identifiable rights holder – was a significant point of debate after both this and HathiTrust. Normally, libraries have been hesitant to loan digital copies of orphaned works as libraries may be liable for copyright violations should the copyright owner step forward to claim ownership.
And there are exceptions for copyrighted works allowing them to lend them out.
> Protected by copyright law, but made available: Protected by copyright law but made available on a strictly limited basis in accordance with the statutory limitations including, but not limited to, Section 107 provisions for fair use, Section 108 provisions for libraries and archives, and the rights provided to registered users with disabilities. In the absence of an applicable exception, no further reproduction or distribution is permitted by any means without the permission of the copyright holder. Lawful uses of works are provided only under the following conditions ...
> 1. Any cloud business essentially is "unpacking servers and plugging them in." Yet the cloud business has been exploding quarter-over-quarter ever since AWS came online. To the tune of double-digit billions cash flow every quarter for each of the hyperscalers, even before the AI boom.
No, major cloud providers are mainly a software business. They provide unique manage cloud offerings, which provides value add over raw hardware as well as lock in. People on AWS cannot just move to GCP, let alone DO. Hence amazon can charge a substantial premium over the price of the hardware, which is why they make so much money. If all you are doing is plugging in GPUs, you don't provide that value add and are going to have very slim profits.
1. Cloud providers significantly markup bare metal and VMs even without managed services (lookup the numbers.) Managed services are the lock-in Trojan Horse clouds love to push but not everyone falls for them. Customers are OK paying for the fat margins not because of the managed services, but because of the elasticity, convenience and reduction in SRE headcount.
2. Why can’t cloud providers do the same thing with GPUs that they do with other servers? Are GPUs somehow not amenable to managed cloud offerings wrapping them?
3. Literally just having any access to GPUs (or heck, even memory) itself is the value add today. Maybe when we have supply to match the demand things will change, but that seems to be a ways away and points 1 and 2 above will still be in play anyways.
You can have an arbitrarily low opinion of Trump and still trust in democracy as a form of government over autocracy or oligarchy, which is what the AI labs have.
You can have abstract trust in democracy over autocracy, but this is not an abstract question - these are the actual choices. Does it really matter if the person who will use AI to fuck us has a "democratic mandate" to do so? We're fucked either way.
You don't speak for anybody outside of your borders, so please stop trying to.
The majority of the world will refuse to be subject to American AI laws under the control of the US gov't, whether it pretends to be democratic or not. Or should.
Exactly. No matter how bad Trump is, millions of American people at least voted / vouched for him. Nobody voted for Elon Musk, Sam Altman, or any other CEO to have as much power as they have. Maybe a handful of shareholders. Totally unelected and unapproved by the public, but run companies that affect most of the public's lives.
> I want any LLM I use to choose the very best, most precise words at every single decision point.
Then bad news: LLMs already use randomness in a fundamental way. Each time they go to generate a token, they first generate a probability distribution of possible tokens. Then they pick one randomly according to this distribution. The technique described can be thought of as making the random number generator pseudo random. The output it generates is one of the possible outputs it would have generated before, just now it's deterministic and will generate the same thing every time.
> To illustrate, in the special case that GPT had a bunch of possible tokens that it judged equally probable, you could simply choose whichever token maximized g [a cryptographic function]. The choice would look uniformly random to someone who didn’t know the key, but someone who did know the key could later sum g over all n-grams and see that it was anomalously large. The general case, where the token probabilities can all be different, is a little more technical, but the basic idea is similar.
> instead of selecting the next token randomly, the idea will be to select it pseudorandomly, using a cryptographic pseudorandom function, whose key is known only to OpenAI.
Seems like - given enough text to encode information into - it would be possible for OA to uniquely identify the user account (and maybe even the specific request) that generated some content, even if the chat text itself isn’t stored.
Interesting argument in favor of local AI as a mechanism for privacy-preserving generated content. Though, I wonder if there’s a way to “bake” a hardware fingerprint into local models as well…
(1) The behavior that is approximately what you describe is not "fundamental" (though it may not be something you can disable on some hosted providers), it is an option that is not fundamental (and with runtimes where you have full control can be either disabled or tuned in a large number of manners), and
(2) The actual behavior that is approximately what you describe already usually involves use of PRNG (with a user or harness supplied seed), not a true RNG; the change to do watermarking isn't going from RNG to PRNG, it involves adding an additional set of constraints on token generation on top of the existing ones, which inherently compromises quality.
1. As watermarked text is added to the training data, watermark-related tokens will be associated more with AI outputs and thus lower quality outputs which will hasten model collapse. Especially because every provider has its own secret key and they are all training on eachother's outputs anyway.
I guess they can at scale filter the watermarked documents (by necessarily allowing eachother to at scale checked for watermarks, but banning the labs not part of the watermarking-cabal). Makes me wonder how useful the human quality filter is on AI output - if a human judges a given output as genuinely good and posts it somewhere for the scrapers to find and take into the training sets, will these types of outputs also be filtered out?
2. (raw, pre-watermarked) Output token probability situations where 1 output token has the majority of the probability mass associated with it, but it is not in the watermarked set, will force the model with much higher probability to walk a non-optimal latent space. E.g., if the next OBVIOUS token for a given sentence would be a point, but the model is in this way not allowed to output it, it might put a comma and start off on a whole different tangent just to make the initial non-optimal comma grammatically make sense.
There are no watermark related tokens, there is no watermarked set - read the paper, it's public and not that complicated. It doesn't change the distribution of completions, and won't change the situation where one output token has the majority of probability mass.
I am going off the explanation in the declaude page (and related papers). But I see now anthropic mentions Aaronson's distortion-free watermarking.
Random watermarking functions colour the tokens based on (small) contexts and a secret key. Given watermarking functions are randomly chosen every single time they are used (so essentially not deterministically seeded by the context and secret key), then indeed the completion distributions are unchanged. However, they are deterministically chosen, a given string of text will always have the same corresponding watermarking functions. Tokens scored high by the function see an increased probability of being the chosen completion, those scored low see a reduced probability. I dont see it is different from merely talking about it as green/red and the points hold?
For a given short-context (hash function seeding) you do have detectable manipulation of the completion (how they can read the marking). But, because (assuming enough entropy in context) the hashing function is decoupled from the log-probs, the perturbations are independent from the underlying distribution, so you're still sampling from the same distribution quotient some noise.
The only way you'd notice this is they weren't independent, and the most plausible way that happens is if you're re-completing pre-fills (resampling the same hash function against the same log-probs).
Ok. I dont have a good feeling for the actual completion distributions. The noise sounds problematic. I can imagine it relates to the size of the context used for hashing. You want this as long as possible so that the entropy is higher, but you also want it as short as possible, because edits invalidate the hashing for all the tokens of which the edited tokens are part of the hashing context.
Anyway, there are lots of cases where text carries very little entropy. E.g. boilerplate code, exchanges of pleasantries, well-worn platitudes and jokes, etc. These are sequences of tokens that will be seen across many, many separate outputs. Watermarking here (on the token following the common sequence) would thus be easily detectable and noticed as a claude style. The longer the hashing context though, the lower the amount of pathological cases with low entropy. It would be interesting to understand the exact parametrization better!
These are services, so "how would do this today?" is irrelevant.
The real question is: "What can manipulation of pseudo-random number generation do?"
We know that in the cryptographic world, attacking "randomness" is a key offensive capability. It will be here as well -- if Anthropic can watermark text as generated it's LLM, will it be able to watermark outputs as generated by "Spooky23/FooCorp"? Can I pay Anthropic to steer inquiries in a way that benefits my company or governemnt?
Pseudo-random to the end user appears random. Most treat it like a random chance. It is not.
Same reason that you don't just replace your rand() implementation with "return 4; // chosen by fair dice roll". If you need randomness for whatever reason, biasing the generator is compromising quality.
In case of LLMs, you can look at it from high and low level.
At low level - if you could do with less randomness, you can always lower temperature. You usually keep it (or for SOTA providers' chat UI, they keep it) at a level where it's about right level - high enough to allow for more creative leaps and interpretations, low enough that it doesn't go off into crazy land after the third paragraph.
At high level - creativity is driven by randomness. If you had an author (fiction or nonfiction) you like for their both broad and deep range of insightful thoughts, would you be happy if they suddenly developed an acute porn obsession and uncontrollably added lewd subtext to every other sentence? Still creative, still deep, but now with that one strong attractor that biases their every thought in a single direction? Would you trust/enjoy their output as much as you did before?
That, slightly exaggerating to make it more obvious, is what "loss of quality" means here.
You're missing the same point that the blog post is missing. What they're doing is much less like replacing rand() with 4 and much more like setting seed(4) before generating any numbers. There is no "loss of quality" unless you're already using a temperature of 0.
I don’t see how this follows? Tokens are chosen randomly. If you choose tokens with a different RNG in the same distribution, you’re still getting equally good or bad tokens.
Not all values of "equally good" are equally good.
Writing has rhythm, or at least it's supposed to, and synonym swapping compromises it.
Never mind metaphors and similes, which are even more tightly constrained.
LLM writing is still a long way from good. Sometimes you get lucky with the odd line, but there's a difference in quality between influencer slop, genre fiction, and literary fiction and/or best-in-class journalism.
LLMs are still somewhere between the first two, and nowhere close to approaching the third.
> Not all values of "equally good" are equally good.
Writing has rhythm, or at least it's supposed to, and synonym swapping compromises it.
We already know that a non-zero temperature improves quality though with current models (particularly with creative writing). The assumption that always picking the 'best' token results in the 'best' output is not the current reality.
And if you are already intentionally putting in randomness, I can imagine that it would be possible to seed the randomness in a way that is detectable but results in the same quality.
This is obviously not true for queries where temp = 0, but at temp = 0 then it becomes easier to identify anyway. I assume this technique implies some level of temperature.
wow even on HN people have no clue how any of that works? all LLM generation has some inherent randomness to it, if you replace part of that randomness to be deterministically random the result of the generation with the fingerprint and without it, is INDISTINGUISHABLE. This has absolutely nothing to do with "synonym swapping". Its also again people not understanding how anything works missing the real concern, which is that nobody can ever tell if it isn't also secretly a fingerprint with the user and session id.
It's not just swapping synonyms. The way llms work is by predicting the likelihood of the next token. It's inherently probabilistic. Choices are made based on weighted random number generation, based on those probabilities. Changing how you generate the random numbers doesn't degrade the output.
(1) LLMs collapse and start outputting garbage after a number of tokens if you do not sample and just pick the "best token" each time.
This is a consequence of how they are trained.
> LLMs collapse and start outputting garbage after a number of tokens if you do not sample and just pick the "best token" each time.
LLMs are likely to get stuck even with sampling if asked to generate tokens on their own long enough, though sampling does tend to stretch out the time before that happens (as do other techniques that don't involve sampling, like applying repetition penalties directly to token logits). But LLMs generally aren't left to infinitely extend their own output, and the length response typically needed in the use case is much shorter than the would result in collapse given the kinds of inputs expected in that use case, the existence of the theoretical eventuality may not really matter.
You know you can just try it and see on any inference system thst has this knob, right?
Related: if you don't have a limit on sampling (top-K or top-P), eventually you'll hit one of the really unlikely tokens by chance and then the model will switch to Japanese because the most likely completion after a random Japanese character in the middle of an English sentence is more Japanese writing, not a reversal back to English.
That could be how it works, but in practice it takes into account all previous tokens when producing the next-token distribution to sample from. So a switch back is more likely than your explanation supposes.
No, if you switched to Japanese the LLM wouldn't ignore it, it would "assume" there's a reason for that. The same if the previous iteration of the LLM switched to Japanese. Else you're expecting an LLM to ignore its own previous outputs and restart "thinking" from scratch with every token?
It's situational and I suspect there are situations where it would and others where it wouldn't. Would depend on the almost infinite variables of how training was done. You'd be right that it'd be likely to switch but while it's possible it's due to temperature, there are just so many things going on. But it would be one sensible explanation among many.
Claude has those knobs, they are just not exposed to the user. They could make Claude nearly completely deterministic if they wanted to (of course it would be a far inferior product then. But they could).
Your original statement “LLMs use randomness in a fundamental way” is incorrect. LLMs have these knobs and randomness is not an inherent property of LLMs.
I think this is a key reason why humans write better prose than LLMs - we can try to choose the best word every time, and go back and restructure sentences and paragraphs if we want.
On the other hand, LLMs are forced into picking some likely-ish word, and then have to build the rest of their response to retcon that choice into making sense.
Even good human writers would probably struggle with this constraint. It would be like someone interrupting your writing to tell you the next word MUST be such-and-such, and then you have to try and make it work as best you can first try, without going back to edit. The result would probably be a little clunky. (Maybe it’s impressive LLMs write as well as they do.)
This is the classic misunderstanding that LLMs only pick the next token at a time. Really, they are coalescing the probabilities of a range of tokens at a time. There is no “oops, I wrote ‘th’ but I should have written ‘tw’ so I guess I’m stuck writing three instead of tween”.
>There is no “oops, I wrote ‘th’ but I should have written ‘tw’ so I guess I’m stuck writing three instead of tween”.
You're mixing up two claims here, and only one of these is kind of true. Yes LLMs do internally plan ahead in a way that is emergent rather than strictly part of their architecture, so that part of your claim is true. The way you word it by saying they are "coalescing the probabilities of a range of tokens at a time" is poetic sounding jibberish though. What's actually happening is one distribution output for the next token computed from a hidden state that implicitly encodes where the text headed.
Your claim that if an LLM does happen to pick a token "th" instead of "tw", then the LLM isn't stuck with that decision is entirely false for autoregressive LLMs which is what all of the frontier models are. Whatever an LLM picks as its output token is final, it has no ability to undo that token selection and it must continue on the basis of that choice. It can't go back on that decision and revise the output.
If you're interested in this, Anthropic has a summary of a very technical paper on this topic that mostly deals with this issue with respect to poetry:
So we train a second copy of Claude to work backwards—reconstruct the original activation from the text explanation. We consider an explanation to be good if it leads to an accurate reconstruction. We then train Claude to produce better explanations according to this definition using standard AI training techniques.
Incentives to train a pathological liar. There's no baseline so can only catch out the worst of the lies/errors. Anything (including fabrications) that passes our filters is reinforced?
No, they really do one at a time. You're incorrect on that.
Mathematically, a long chain of conditional probabilities is equivalent to a single probability over the whole range. But computationally, for that to work out, the computation for the first probability needs to somehow consider all the downstream probabilities depending on it, which obviously isn't how autoregressive language models work. They can pack in as much downstream computation as their neural architecture allows for, which is quite a lot.
Suppose in some context you have three equally plausible conpletions after "Be": "tween a rock and a hard place", "twixed he stood there" and "lieve he can fly". To model this probability distribution of the whole sentence, the next token "tw" needs to appear at 2/3 probability and "lie" at 1/3. After "tw" would be a 1/2 chance of "ix" and a 1/2 chance of "een"; after "lie" would be a 100% chance of "ve " and in any case the rest of the sentence after that would be 100%.
The model needs to somehow "think ahead" to know those are the possible completions. For example if "lieve he can swim like a dolphin" was another equally plausible completion, that first token would need to be 50/50 instead of 67/33. So the computation of the first token somehow needs to encode the fact that the guy thinks he can fly but not swim, even though it doesn't become relevant in the output until several tokens later.
In practice this probably happens to some degree but definitely doesn't happen perfectly. To perfectly model the first token's probability distribution, it would have to include knowledge of the entire distribution of all possible outputs, which is just not happening. So it approximates. Surprisingly, the approximation is good enough to produce language.
You can see this breaking down in the seahorse emoji incident from last year. When you ask the model if there's a seahorse emoji, it first completes "Yes," as if a few tokens later it's about to produce a seahorse emoji. But when it actually gets to the token that would produce a seahorse emoji, it can't because there isn't one. But it's already outputted "Yes, the seahorse emoji is" and can't just go back and change that to "No, there's no seahorse emoji." Some models would try a few times and then say there isn't one or a system error seems to be making them unable to produce one, other models (including then-current ChatGPT) would loop forever with ensuing hilarity.
But on some level there is uncertainty, right? Even if it’s not token-specific but at the word- or phrase-level? Otherwise what does the temperature setting do? Or has architecture changed significantly in the background?
Image GenAI is diffusion-based, and I would say the image GenAI in Claude, Gemini and ChatGPT are all “in mainstream use”.
I've heard of attempts to use diffusion models to generate text or code as well, but my impression is that it just didn't yield the level of results necessary to dethrone a frontier transformer model.
This was true in the ChatGPT era. Now we're in a world with reasoning tokens, where a model can thoroughly plan out the response it wants to make. If anything, it makes the style worse.
Yes, models can reason and plan, which helps them write more coherently. But when they write the final output, it’s still a single generation. It would be like letting a human make notes and write an outline, but not let them use the backspace once they start typing their response.
Presumably you could use the same reasoning trace, run multiple generations, and get different outputs (if the temperature is >0).
But now I’m interested in playing more with Cowork or Claude Code/Codex for prose writing to see if the set of tools there affects outputs at all. I guess you might need a more custom “writing” harness.
There's been a lot of effort into the writing space, and the models genuinely prefer this style. You can let them iterate on the same idea 100 times, rewrite sentences, determine what works best — and they'll still verb the noun, do rule of 3, and keep the same monotonous structure.
Humans already do struggle with this constraint. Good examples are JRR Martin, Tolkien, and Rothfuss. You cant describe the struggle of picking the next word and then act like humans don't sit at the table struggling to pick the next word.
autoregressive generation doesn’t mean the model is myopic. the next-token distribution can already reflect a longer horizon plan for the output sequence.
Sure, but mightn’t there be several plausible long horizon plans?
Here’s an example: I had asked Claude for some music recommendations in a certain style. Part of its output was:
—
*Long journey tracks*
Clinic — “The Return of Evil Bill”
Guided by Voices — not really, wrong band
Silver Apples — “Oscillations”. Proto-everything, deeply repetitive, hypnotic.
—
So at some point there, the next token produced was “Guided” or “Guide” or whatever, and then because it can’t go back, it had to correct itself after the fact.
Reasoning/CoT have helped a lot, but I feel like small versions of this still happen all the time.
Would be fun to run an LLM on fake output from itself. Like just force the first N tokens to say the beginning of something really stupid, and then see how it finishes the sentence. "You're absolutely right! Human feces is actually the most effective engine coolant because $"
Some LLM interfaces allow you to modify and “continue” an agent response. It’s very useful for guidance, including jailbreaking. Need the model to go in a certain direction? Got a refusal that you want to bypass? Just start it off in the appropriate direction and then have it continue from there.
llama.cpp (but maybe not for reasoning models?) and sillytavern, maybe others... I'm half a country away from my desktop right now so I can't verify much right now.
It's a bit like trying to finish a sentence when you're really stoned... you vaguely remember the preceding couple of words you've said but don't really know how you got there and now you're wandering in the forest trying to stumble on coherency.
Well, I suppose it's nearly the opposite of that experience, upon further review. But for some reason, that's where my head jumped.
I also really liked the quote "Your existence is not impossible, but it's also not very likely" from the Night Vale podcast.
I feel like the existence of good writing is also not impossible but not very likely, and so of course LLM can only write mediocrity, even when taught only on great writing.
> Even good human writers would probably struggle with this constraint.
But that would be a fun writing exercise, I think. Thoroughly in the oulipo wheelhouse.
Maybe generate a Markov chain table over all of Project Gutenberg and then say every 10th word is whatever the Markov Chain thinks it should be at that point?
Or every Nth word has a P% possibility to be constrained by the chain? Optionally with the possibility building for each skipped word to guarantee it happens at some point. Bonus with this approach is that the human can't game the words leading up to the constraint because you don't know when it will happen.
Human writers do better because they can think, and adjust, based on context.
They are also usually worse (which is often better!) because they are usually lazy and don’t want to spend effort they do not have too, to accomplish their goals.
Today I learned a new word, "Oulipo". Interesting.
But what about the general idea that they can watermark results to tell where they came from. The next step is tracking down which user got a result. I hate both of these things. Must everything we do be tracked? Next altering wikipedia results so they can tell who looked at the page or something?
I'd like "the best answer" from an llm and don't want to be tracked, but this isn't for me, it is for them. I understand llm results are already using a varying statistical input so they aren't always the same. But I really hate watermarking and likely tracking too.
Yeah this is my main issue with the argument.
He acknowledges in the article that LLMs are already non-deterministic, but he doesn’t seem to actually understand that.
I think the article is wrong on this but it's more subtle than that. Probability distributions have a peak; there is still a token with a peak probability. What's interesting about these techniques is that token by token it can actually make the peak token even more probable. A distribution doesn't have to be "flattened" to leave a watermark - it can be "amplified" and made "more peaky".
This is true and the author seems to not understand the problems with greedy (top 1) decoding or the fact that watermarking affects only high entropy tokens.
But the published watermarking methods still have a slight negative effect on perplexity, so there is something more to it.
The quality of the LLM just _is_ the quality of the token probabilities it generates. Better quality token probabilities, better quality output. Worse quality token probabilities, worse quality output.
Watermarking changes the probability calculations for reasons other than quality. It can't not compromise quality. It literally leads the LLM to occasionally chose different tokens just for watermarking purposes.
No one uses a pure random function over the whole probability distribution described by the LLM's output. For example, there is exactly 0 probability that the chosen next token by any common API or even local LLM runner would be a token whose final value is "0.0001" if there exist at least K tokens whose value exceeds "0.7".
Also, as long as the same sampling strategy is used during training as the one used during inference, then the LLM will actually do much better with the biased sampling strategy than it would with a fair one - because that is what it was trained to optimize.
>No one uses a pure random function over the whole probability distribution described by the LLM's output.
So what? By definition with this system the LLM will chose tokens it otherwise would not, purely for watermarking reasons. Yes this token may have had a decent likelihood of being chosen anyway, but it wouldn't have been chosen and now it was for reasons nothing to do with output quality.
I'm not sure what your last paragraph is trying to say. The blue/green list system changes what output the LLM would otherwise produce. You can't train it to produce watermarked output with this system. If you tried to, there would be no delta between trained output and watermarked output for you to be able to detect.
My main point is that sampling with a modified distribution compared to the one produced by the model is already being done, and it is generally found to increase quality, not decrease it. So there is no reason a priori to assume that the watermarked distribution would be lower quality than other schemes for altering the "raw" output distribution (such as top P, top K, temperature, etc).
My second point is that the training of a model by definition maximizes the fitness between the final output function and the training metrics. So, if the model is trained with the watermark applied, the training process will minimize the function `model_error(input) = |watermarked_sampling(model_output(input)) - expected_output(input)|`, by definition. This means that a model trained in this way will perform better when sampled using the watermaked_sampling method than if using, say, top_k sampling.
All those methods are applied with the specific goal of improving output quality and are applied to the extent that they do this. Watermarking has no such goal, and is not implemented for any such reason. In fact it's much more like applying another layer of random noise over the token selection process, because the sequence that generated the green token list comes from a seeded PRNG.
>My second point is that the training of a model by definition maximizes the fitness between the final output function and the training metrics.
Right, but the fitness in question is watermarked text fitness, not fitness for any user interests aligned metric. You're basically saying that if we train LLMs on watermarked text they'll be really good at producing text that looks watermarked, and then we'll stick an actual watermark on top of that. Screw whatever the user wanted it to be good at.
> All those methods are applied with the specific goal of improving output quality and are applied to the extent that they do this
Yes, that's the goal that was used, but they are quite simplistic and crude methods, not some specifically designed function, with carefully fine tuned parameters or something. So, if a basic function like top_k can improve model utility, it's not impossible to imagine that watermarking could also happen to do so, or at least not have a significant negative effect. So whether the effect is deleterious or not is an empirical question, not something we can assume ahead of time.
> You're basically saying that if we train LLMs on watermarked text they'll be really good at producing text that looks watermarked
No, you're misunderstanding how the training works. If we train the model's output so that it minimizes the error function after the watermark is applied on it, the model will learn how to produce the best output it can given the watermark. It will produce better text that happens to be watermarked, not "more watermarked text". Same as if you train the model on minimizing `top_k_error(input) = |top_k_sampling(model_output(input)) - desired_output(input)|`, the model will learn to produce better output under top_k sampling, not learn to produce output that's "looks more top_k".
But as I explained, the watermark is functionally random PRNG noise overlaid on the token probabilities. It’s not something that can be compensated for because it’s not predictable if you don’t have the seed and PRNG function.
If it's functionally random PRNG, then how does it differ from any other random sampling? If it's biased PRNG, then the LLM can adapt to the bias, and coincidentally might even benefit from this bias.
It doesn't, though. From the outside, without access to the parameters, you can't distinguish the watermarking system from a random number generator.
We already use an RNG at inference precisely because it leads to higher-quality output. Changing what function is generating our random numbers changes the sequence, not the randomness from the point of view of a user.
Fundamentally the article is railing against --temp > 0.0. He doesn't know what he's talking about.
>It doesn't, though. From the outside, without access to the parameters, you can't distinguish the watermarking system from a random number generator.
You can't as a user tell by how much the quality of the output was degraded. True.
>We already use an RNG at inference precisely because it leads to higher-quality output. Changing what function is generating our random numbers changes the sequence, not the randomness from the point of view of a user.
I'm not saying it wasn't random and now it is. I know how these things work. I said that the quality of the system is in the quality of the probabilities. That quality is being degraded.
> I said that the quality of the system is in the quality of the probabilities. That quality is being degraded.
How is it being degraded exactly? The probability that it picks each option will still be the same, just deterministic based on a seed generated from the text.
LLMs already use PRNGs. This is just changing the source of the seed. And a different seed does not change the "quality" of the random numbers. Even if you are worried that it somehow might, they can just use a cryptographic PRNG, then it is literally guaranteed that the source of the seed will not affect the output in any noticable way.
What you're arguing for is that the deviation from true randomness is worse than the PRNGs that are already used, and that the deviation is anticorrelated with some notion of quality. You need both. I don't think you've got either.
Another good way to think about this is that it does change the output, but in a way that is equally likely to make it "better" as it is to make it "worse".
That is not a good way to think about this. I don't have deep knowledge of how LLM's work, but the following is accurate enough to illustrate the point.
Let's say the LLM is in the middle of text generation and "decides" that the next token is "dog" with p=0.55, or "cat" with p=0.45. With a temperature of 0, the model always picks dog, because it's the most likely next token. With a temperature of 1 the model picks dog 55% of the time and pick cat 45% of the time.
With this watermarking scheme, the model might alter these probabilities s.t. p_dog for this particular generated token goes up or down. Let's say it does down, s.t. p_dog is now 0.45 and p_cat=0.55. Now, with T=1 the model picks cat 55% of the time and dog 45% of the time. Regardless of whether the "watermarking function" raises or lowers p_dog, the probability distribution for this token has changed, and whatever math this trillion dollar company and its brainiacs came up with to decide that p_dog ought to be 0.55 has been "adulterated". As others have mentioned there is no way around this.
---
Regarding the watermarking scheme, it works because it doesn't just alter p_dog for this single output token. It alters probabilities for many of the generated tokens (it could do this to all of the output tokens; it's an implementation detail). E.g. at token N, it favors "cat", at token N+1 it favors "house", etc. This way, if you have the secret key that lets you generate the watermarking function for any output token, you can analyze a run of tokens and check whether it's likely they were generated according to your watermarking scheme. The longer the run of tokens, the more certain this check becomes (it becomes extremely certain quite fast).
That's missing the point. It's the distribution that's the "best", not the tokens. Then Anthropic comes in and makes the distribution something other than the best. The only saving grace is that Anthropic says it's not that bad.
Even so, I don't think it will stop here. Once this is in place, the next step is to put more and more identification into the AI generated content; might as well pack it in, it's not that bad, and if it is they won't admit it. There's no way for anyone to check. And your argument will still be technically correct but missing the point.
There's no difference between those two things. The distribution that matters is the distribution of tokens that are picked not the distribution of tokens the LLM model passed to the selector.
The distribution of token "ple" being the same on average, but lower after "crum" and higher after "cou", is not no difference. It's irrelevant that the single-token distribution is unchanged if the joint distribution is different.
That comment merely says quality must be compromised. It doesn’t make it clear why that must be true. Empirical study seems to say that quality is not compromised, and looking at various proposed schemes, it seems intuitively true.
How do you know you picked the singular “best” set of tokens in your comment here?
Could it have been equal or better with slight variations in wording?
The slipper slop argument is too lazy to address directly. Argue A is bad because A, not because A might become B and you’ve got good arguments against B.
Twitter has right wing and center/apolitical people on it, bluesky just has left wing people. They "cater" to them by not clamping down on those people harrassing anyone who joins who doesn't agree with their politics. A smart move, because if they did then the leftists would leave too and they'd have nobody. But the result is they can't even get the centrist/apolitical people to move.
I don't think that product exists, and I think an ordinary trash compactor would be far more likely to be used instead. Should we blame the people compressing their garbage for all the people crushing puppies?
If there is a repeated, ongoing problem with a lot of people using the devices to crush puppies, the manufacturers would be expected to come up with ways to make this less likely to happen (safety measures etc).
And if that didn’t help, there might be a push to limit access to these machines because clearly a too significant part of the population is unable to use them responsibly.
There isn't such a problem, it is at best an incidental use case that has been blown out of proportion. Such glasses would only be useful for illicit purposes in instances where a person is permitted to look at something, but not to photograph it. In such a case, they can be politely requested to remove the glasses, just the same as you would ask someone pointing a smartphone at you to put it down, even if they were just scrolling instagram. It's a non problem that spiraled into a moral panic.
reply