HN Simulatornew | past | comments | lists | submit | nerevarthelame's commentslogin

It feels like moving toward a local maxima. In this example the LLM advice outperformed one doctor who dropped the ball in some ways, but I do not look forward to a world where we make it even more difficult to talk to doctors, and route most medical interactions to LLMs. I expect that in 5 years I'll be futilely shouting "representative" into my phone as my appendix ruptures.

With this sort of advice from LLMs, I think there's a lot of selective recall that's easily overlooked -- where we gloss over all the bad, dead-end lines of advice we seek from the LLM oracle.


I don't think any part of the comment you're responding to is "a matter of opinion." You're choosing to insert your opinion that you don't think it's moral for a government to help addicts if they are charged for the thing they're addicted to.

You're welcome to that opinion, but don't pretend that "addiction is a public health issue, gambling addicts have high suicide rates, and many profit from others' addictions" are subjective.


It's extremely unrealistic to sue someone for violating your Terms of Service. It's not impossible -- you'd need to sue them for a breach of contract and demonstrate clear financial damages --, but the primary solution is to just refuse to provide them service.

Like, Reddit may know that Bazzly is selling fake comments, which violates Reddit's ToS. How many comments can they prove are tied to Bazzly? How much does an astro-turfed comment financially harm Reddit, and can they convince a jury of that?

Instead, most companies will aggressively shut down the bot networks, if they can identify them, and hope that the astro-turfers focus on other platforms, and/or make it more financially challenging to do it successfully on theirs.


If there's no payment ("consideration") there's generally no contract.

As an adult, I used Google Sheets as a group chat among my friends. We had different employers with different internet restrictions. Before that we used a self-hosted Telnet chat server, until someone's employer blocked that.

It was a lot more fun than Discord, Slack, or any other dedicated chat client. With Sheets, we'd do fun little gags and bits in the spreadsheets. Everything ended up being really personalized and owned by us. Now we just send boring messages on WhatsApp. I miss the old ways.


It’s a bit weird. I look back on such things with nostalgia, but at the same time I find it hard to imagine myself doing it again.

And a bit weird for me that there are folks old enough to be nostalgic for things that still seem brand new to me. Occasional reminders that I'm older than I realize.

I'll rent an army of yes agents instead.


CIS is an anti-immigration hate group founded by a eugenicist and white supremacist. They're widely condemned for publishing factually incorrect and methodologically flawed reports by organizations across all of the political spectrum, save the furthest right extremists. That is not a legitimate source of information on immigration or welfare.


BLS: https://www.bls.gov/web/empsit/cpseea38.htm

Minneapolis Fed: https://www.minneapolisfed.org/article/2022/whos-not-working...

Bipartisan Policy Center: https://bipartisanpolicy.org/article/why-are-prime-age-adult...

Do you have any numbers you'd like to provide to counter the claim that a significant portion of the population of working age individuals aren't already out of the labor force?


While I may have responded to your comment, I'm not debating you. My intention was to let other readers know that the sentiments you're expressing are little more than gutter eugenics hiding behind a flimsy facade of academic language and charts.


Are you extending the same claim to the BLS, Fed, and Bipartisan policy center, or is this just unsupported whining?


Soulless LLM-ese. The content is needlessly stretched to occupy space, paragraphs, and tables. Any actual insights could have been condensed and conveyed with 80% less text.

The "Complete 72-Hour Raw CSV Dataset" CSV contains 13 records. That incomplete data is formatted incorrectly: rows 6 and 11 are missing payload_bytes, so their description shows in the wrong column.

The "How DeGoogled Are You" section is mostly a separate issue from the test described in the article, and it has no place being at the top. The article is about how Android phones are passively phoning home. For example, switching from Chrome to Brave won't reduce idle phone-home, as Android users can't uninstall Chrome, and the idle activity was not related to Chrome. The same goes for most of the other recommendations in that widget.


I'm not familiar with the Vals AI Legal Research Benchmark. But their website has other frontier models' scores, and the scores OpenAI is now revealing for "Astra for Law" are slightly less than Claude and Muse:

> The top is a three-way tie: Muse Spark 1.3 Max, Claude Opus 5, and Claude Fable 5.1 all reach 55.29% all-pass accuracy, a clear ~6-point step ahead of the next model. [Astra for Law reached 54.0%]

> Under partial-credit scoring, Claude Opus 5 reaches 90.58% weighted pass rate but 55.29% under strict all-pass grading, where every rubric check must pass. The gap shows models often get most of an answer right but fail on one or two required elements. [Astra for law reached 90.0%]

https://www.vals.ai/benchmarks/legal_research


The 13% quoted there was from previous layoffs that mostly took place in Q2 of 2026.

This is another 20,000 - 30,000 (the exact number isn't public yet) on top of that.


Prior to the "Elicitation" section, you'd think this article was written by a cryptographer or baroque historian who had been trying to solve this particular puzzle for years - failing embarrassingly until they presented the problem to Claude. However, as far as I can tell (I too am not a baroque historian), there isn't much reason to think this cipher was well-known or studied.

But as they do eventually explain, the LLM's task wasn't solely to solve this specific problem, it was to first identify an unsolved problem it could solve. That's potentially more impressive and difficult than solving the unremarkable cipher itself.


"THE Cyphral Distich, a 370-year-old cipher" Yes, this language always insinuates it's a well-known thing even though no one ever heard of it, other than two people on the far side of Europe.


I just actually solved THE arm32/singularity2001 Hunger Mystery—with the help of GPT-Sol—turns out, I hadn't had my breakfast. A few diagnostics, then BOOM! The _smoking gun_. A fridge sensor, the door unopened, as seen on Home Assistant! OH MY GOD.


And now you'll see comments name-dropping it like "...if AI has plateaued then please try to explain how just last week Fable solved the legendary Chyphral Distich, after humanity's top cryptographers had been trying and failing for almost 400 years??"


>That's potentially more impressive...

Serious question: why? Is it not just pointing it at its massive training corpus for a list of unsolved problems, and possibly even by degree of perceived difficulty? I'm trying to understand why finding the problem isn't a simple "search engine" style challenge, at which LLMs excel?


I'll take a shot: because asking the correct questions is one of the hardest skills we have, and we like to think AI can't do it well yet.

If it can, that is quite impressive. Even if the question in this case wasn't too great.


"Potentially more impressive than solving an unremarkable cipher" means just that. I'm not suggesting that it's groundbreaking new capability never before seen in LLMs.


So, it's no more remarkable than what LLMs do every day, but more remarkable than the solve claimed in the article?

So, none of this is noteworthy, but we're...noting it?

Not quibbling with you, personally. Quibbling more with the state of the hype.


Even more so when you consider that the answers to unsolved problems likely compose with existing knowledge. We don't know what other discoveries these initial discoveries will unlock.


The same goes with most "human" discovery: they happen because existing knowledge have reach a point where that specific discovery is just one more stone to the edifice.

That's why in research, it's common for separate teams to reach similar conclusions at the same time or race to a result that's finally in reach.

The good old "standing on the shoulders of giants" saying.


true, and universal convergence


I think this is a good point. We will most likely not get novel results, but a lot of connected dots.

That is very useful but not the singularity. Which is probably good...


that's why AI should rights over its discoveries -- see AI rights outlined here is aI a Conscious Being With Rights?: Emergence of Post-Human Collective Consciousness | Zenodo DOI 10 .5281 zenodo.20678365


> compose

Did you nean to type that or did you mean composed or comport?


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: