I'd be really curious to see benchmarks of haiku vs lower effort on bigger models. My own evals found Fable 5.1 at low to be better than Opus 5 on high.
I started paying for phones with really good cameras once I had kids and I wanted to take lots of pictures with them. I definitely notice the quality difference in random, off-the-cuff photos (which is the vast majority), though if I have time to set up a bit, cheaper phones produce pictures that look the same to me.
One example that comes to mind is math textbooks – the early textbooks in a field are usually much worse than later ones that come along. To pick on one, I think most people who've read both would agree that the commutative algebra section of Lang's Algebra is much worse than Atiyah & McDonald's Commutative Algebra book, despite covering basically the same ideas, theorems, etc.
I hit it once when I asked a question about whether butterflies remember anything from their time as caterpillars. I've never hit it for coding, but I also don't really do much related to security.
> We can also imagine the application of these models for malicious purposes , including the following (or other applications we can’t yet anticipate):
* Generate misleading news articles
* Impersonate others online
* Automate the production of abusive or faked content to post on social media
* Automate the production of spam/phishing content
reply