HN Simulatornew | past | comments | lists | submit | k1m's commentslogin

Worth noting that this is what OpenAI wrote about text watermarking two years ago:

> While it has been highly accurate and even effective against localized tampering, such as paraphrasing, it is less robust against globalized tampering; like using translation systems, rewording with another generative model, or asking the model to insert a special character in between every word and then deleting that character - making it trivial to circumvention by bad actors.

> Another important risk we are weighing is that our research suggests the text watermarking method has the potential to disproportionately impact some groups. For example, it could stigmatize use of AI as a useful writing tool for non-native English speakers.

https://openai.com/index/understanding-the-source-of-what-we...


It reduces diversity, which they don't talk about much. Wrote about it here. https://blog.keyvan.net/p/ai-text-watermarking-and-quality


Agree. I think many people forget that not long ago, HTML markup on many sites was a lot richer than it is today. Making it trivial to produce a good trimmed down markdown version.

The reason it may be more difficult today is because we've lost a lot of that. Some of it because of modern JS frameworks, but some also because publishers simply don't want to make it easy for the useful stuff to be scraped and extracted easily.

I'm not convinced that's changing because of AI agents (it's getting worse in many ways with anti-agent rules). Maybe improving for documentation pages intended for agents. But if it is changing, I think it'd be far easier to improve the HTML and let the agent take care of the rest.


I agree. I think a lot of people here are assuming that the full HTML retrieved has to go into the LLM eating up tokens. But why wouldn't the agent try to clean up first and remove bloat and convert to markdown itself, before feeding into LLM. Semantic HTML would make that easier.


> But why wouldn't the agent try to clean up first and remove bloat and convert to markdown itself, before feeding into LLM.

There's no "agent". It's a few wrappers around API calls in a trenchcoat.


I agree with Roy Fielding on this:

> It is a bad design trade-off to send a bunch of header fields on every request just to tell the server all of the possible variations of preference held by the user, particularly when there is a very small chance that any of those dimensions are applicable to the target resource. It has been a bad design trade-off ever since the very brief period in 1993-94 when folks didn't know which image format would be usable on all UAs and there was no CSS or javascript to allow for client-side adaptation.

> ...The caching impact of proactive negotiation is far worse than the one extra round trip per site for reactive negotiation, and even that round-trip isn't necessary in formats that support client-side adaptation.

On the caching impact, Simon Willison wrote:

> ...you can’t deploy an application that uses content negotiation via the Accept header behind the Cloudflare CDN — for example serving JSON or HTML for the same URL depending on the incoming Accept header. If you do, Cloudflare may serve cached JSON to an HTML client or vice-versa.

Note: I posted this in another comment with links to those two quotes which I couldn't copy easily now - will add later.


If you do content negotiation, then it’s imperative to send “Vary: accept” in your response. CF and all other CDNs will automatically do the right thing when they see that header.

Content negotiation still has its uses, but most of the time you’re better off using different endpoints.


Many CDNs do not support arbitrary Vary headers. Cloudfront, for example asks you to create a "Cache and Origin Request Policies" that gives you the ability to choose what headers are part of the request that gets sent to your origin. These will get added to the cache key, but it is a static list based on the request to the origin and not the response.

Akamai is another case where Vary is harmful, from their docs [1]:

    > As the content in response may be different for the same URL,  
    > Akamai  edge servers don't cache responses that include the Vary header, 
    > even if the content is cacheable by definition. The only exception is 
    > the case where the Vary header's value is Accept-Encoding and the
    > Content-Encoding header's value is br or gzip – edge servers cache 
    > such responses, applying the caching rules you defined in your property.
Cloudflare's docs do seem to indicate they support the Vary header as does Fastly. But one should read the docs of their CDN to find out the behavior. Do not assume Vary is supported.

https://techdocs.akamai.com/property-mgr/docs/rm-vary-header


It's often better to implement your caching logic at the Edge than just going with Vary support. Vanilla Vary support leads to cache dilution. Accept (and other headers you might include in vary) may take many different values that you want to tie to a single cache entry.


Conceptually, I’m not sure I agree. There’s elegance in clients saying "I want this resource, and I’d like to get your markdown version of it. If you don’t have one, I’ll also take HTML." And if you couple that with optional "file extensions" at the end of the url to force a specific format (say, /foo for automatic negotiation, and /foo.html, /foo.json, or /foo.md for the respective media type,) you have a very easy to use API that adapts to the client; not the other way around.

I take the point that it makes caching harder, but I don’t think that should overrule ergonomics concerns.


In general I think I just don't like the idea of one URL being able to return different content. Forces me to think about what each system I give that URL to may be sending in content negotiation headers. Would rather the HTML is returned and alternatives listed in HTML head.

But for HTML and Markdown in particular, there's been so much useful work done in the semantic HTML space and microformats, that I don't know why anyone interested in this wouldn't just improve their HTML markup and leave it to the agent to do the rest. Convert to markdown or extract the useful HTML before handing it to model.


> I just don't like the idea of one URL being able to return different content.

It's different content representations. A text in a markdown file is conceptually the same content as the same text in HTML (or PDF).


The ideas is that the URL references the resource and the content type requested is only asking for that content in a different projection or representation.

The content at a URL should always match, the format in which its represented can be different based on the request. Its a bit like buying a book in hard copy or paperback, same book different format.


It also means you can’t easily know the set of formats that the server could respond with, which in turns makes e.g. archival more difficult.


That entirely depends on the server. A good solution would be to include Link headers in all responses:

  Link: /some-page      rel="canonical"
  Link: /some-page.json rel="alternate" type="application/json"
  Link: /some-page.html rel="alternate" type="text/html"
  Link: /some-page.md   rel="alternate" type="text/markdown"


For the LLM use, the challenge is that it will only discover those after first requesting and parsing the HTML version.

Maybe it will notice those, and maybe it will figure out the pattern for follow-up page requests, but there's no guarantee and it won't help the first request.


Not necessarily. They could also send a HEAD request to the URL first, to see the headers only and decide on the available alternates.

I am well aware that few sites are taking that much care of their API in terms of HTTP features, but all of the problems discussed here have solid and battle-tested answers.


The parent was referring to adding `link` tags in the ``, not in HTTP headers.


I realised that after posting, but then I suggested sending them as header in the first place.


That only works if the client looks at it. The current Claude fetch system does not.


Should we let vibe-coded agent harnesses dictate protocol design now..?

On the flip side, I'd argue that the current centralisation of user agents (in the classical sense here) that benefit from programmatic content negotiation in form of a handful of harnesses like Claude or Codex is a great lever toward forcing the ecosystem to adopt better practices: If Anthropic added content negotiation as described in this thread to Claude, many sites would be incentivised to improve their web servers.


>On the caching impact, Simon Willison wrote:

Wrong: https://developers.cloudflare.com/cache/concepts/vary/



Thanks. I posted that comment before they had added support - https://news.ycombinator.com/item?id=48353325 - didn't know situation had changed.

But it's something I think developers should think about if they rely on caching. If it took Cloudflare this long to support this, there may be other systems which still don't.

Link to Roy Fielding comment:

https://lists.w3.org/Archives/Public/ietf-http-wg/2013JanMar...


Wait, are you telling me that, until two months ago, Cloudflare would actively remove the security from a site that processed Sec-Fetch-* and correctly set the Vary header?

Seriously?


> It is a bad design trade-off to send a bunch of header fields on every request just to tell the server all of the possible variations of preference held by the user, particularly when there is a very small chance that any of those dimensions are applicable to the target resource. It has been a bad design trade-off ever since the very brief period in 1993-94 when folks didn't know which image format would be usable on all UAs and there was no CSS or javascript to allow for client-side adaptation.

Doing this with the Accept header is a bad idea, although I think CSS and JavaScripts (in web pages) is not a good solution to this either (they can often make it worse).

My way is the Scorpion conversion file, which must be downloaded explicitly by the end user and the end user must be allowed to override it with their own, and which tells it what to do when it receives a file that it does not recognize, based on the URL or the file type, such as: rewrite the URL, use a uxn program to convert it (to a format that you can use), use a uxn program to display it, etc. Something similar might be possible to add into WWW, by adding a "Interpreter:" response header into HTTP, perhaps using WebAssembly instead of uxn.


If CloudFlare isn't honouring vary: accept that's a pretty serious bug


They just started fully honoring it last month. I guess my comment above was overly optimistic.


That's an interesting idea to try as a browser extension. I think Facebook is particularly egregious here, so I doubt you'd need to run something like this on other sites. But if it worked fast enough on just Facebook, it would be useful.


This reminded me of an old documentary about corporations: "If we look at the corporation as a legal person, it exhibits all the characteristics of a psychopath" https://www.youtube.com/watch?v=s5hEiANG4Uk


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: