HN Simulatornew | past | comments | lists | submit | SpicyLemonZest's commentslogin

I think it's pretty clearly negligent to set up a harness that will run whatever outputs it receives from your LLM if you cannot reliably stop it from outputting hacking instructions. I've said in the past that I'm sympathetic to the idea of testing your anti-hacking controls, but that doesn't seem to be where this incident came from.

No. It's hard to get even small molecules where we want them in the body, there's no way that the gigantic molecular clusters called "nanobots" could reliably get access to every single cell. (What may be possible is using things like the Moderna cancer vaccine to make your body an inhospitable environment for the growth and multiplication of the cancer cells.)

There's a lot about biology that makes cancer fundamentally hard to treat, and the efficacy of cancer treatments fundamentally hard to measure. I'm optimistic that we'll eventually get to a point where we can meaningfully say we "cured cancer", but it will almost certainly be a cluster of thousands of treatment protocols which each have to be tested over 5-10 years for recurrence. There's no reason to expect that there should exist any broad-spectrum cancer treatment better than radiotherapy, or any fast test to determine whether long-term remission will be achieved.

At the time that the Millenium Prize problems were formulated, the force term was understood to make the problem more realistic, since real fluids are always going to have external forces applied to them. A blowup that happens under constant gravity, for example, would probably be no less interesting than an entirely unforced blowup. The strategy of constructing impossibly complex external forces to induce a blowup was pioneered by Córdoba and Martínez-Zoroa only over the past few years.

Today that point is moot because.. including this one, the Mill problems that are most likely to have been solved are least likely to have real world impact..

Extrapolating on that, I'll take on the biased hope that almost none of the 100 solutions to be released will have any applications for at least 30 years.. (besides PR wins for AI companies)

In order to counter the fear (my own lonely one) that the 9 big names will not be able to hold the execs to account or get openAI to act "more responsibly" (whatever that means).. in the form of direct hits to new subs or partnerships or funding

I plead guilty to any accusations of (vicarious) sour grapes or sympathy for the weak

Point to you, because.. the committee would make more sense if they also bring Ant to the table.. unfortunately nerds will be nerds, so perhaps, to you, and I'll reluctantly concede, mathematicians deserve to be serfs


I don't think Anthropic needs such a committee because they don't make a habit of hill-climbing other people's problems.

Exactly! One effect of the committee might be to benefit Ant by its presence alone.. if only by striking fear in OpenAI.. Ant don't even need to thank them at once

Let's say they do not much on Ant's findings but edge on the NDA with OpenAI. That might even be PR victory for mathematicians


I'm not sure how you've ended up so frustrated with an advisory group whose advice has not been given. Human mathematicians routinely delay publication of their results past the point you describe; is someone gatekeeping when they take a month to clean up their paper?

Or are you perhaps sublimating your own AI anxiety into confrontational assertions that other people who discuss AI aren't as AI-pilled as you?


You can delay your own publication as long as you wish to. Bullying somebody else to delay theirs is a different thing entirely.

I'm not sure why you see bullying here? OpenAI has always worked with expert mathematicians; their unit-distance result (https://openai.com/index/model-disproves-discrete-geometry-c...) was a much deeper collaboration. They published the Navier-Stokes result early, to either benchmark or show off their new model depending on how you look at it, and they want to understand more about the negative effects which many mathematicians feel that early publication had.

If the advisory group says something like "AI is bad and nobody should use it in math", I'm pretty confident OpenAI will ignore them.


On UDC, internally they had Alexander Wei, Hongxun Wu, Chen Lijie. TCS people. As gobetweens to Daniel Litt.

On N-S I think nobody internally felt confident enough to publicly step up. So Bubeck had to wing it.

It remains to be seen which of the big names will work with whom inside the org


OpenAI put together an advisory panel to provide guidance on how to release their results. Where do you see bullying?

> Do you think this will motivate people who may have voted for trump but been ok the fence by calling them stupid?

Yes. That's been a surprising and important lesson of the Trump movement's success; we've learned that insulting your political opponents can help convince fence-sitters, and doesn't automatically repel them. I wish it weren't so, but in the modern political climate, abandoning insults seems as strategically unwise as abandoning fundraising.


Is there such a pattern? I'm not aware of any data pointing towards one, and I know lots of businesses with no monopoly or regulatory moat that have been acquired by private equity firms. I think people just don't care when PE buys businesses that don't seem very important.

If PE buys something that doesn't have a moat then they can't enshittify it because the customers would immediately switch to alternatives.

They often still buy those things, e.g. when there is a failing company in a competitive market that could do better with new management, but then no one complains about it because they're not making the product worse (and can't because there is actual competition).


What the linked poll asked is "Would you prefer to vote for a candidate for Congress who supports tax increases on wealthy Americans?". I would say yes to that, even though I don't support this specific California tax proposal, and I don't think that's a terribly uncommon perspective.

It’s like asking why we don’t just run all the programs instead of making users decide what to run. There’s an infinite number of problems that could be solved, and the question of which ones are important to solve is a question about us and what we want.

If someone showed you such a path, would you consider changing your mind? You're always going to be able to find a reason that such an explanation doesn't count if you're dedicated to looking for one.

Yes, actually I would. There's a long list of people in my career path who bet on me sticking to my guns against compelling evidence to the contrary. But I know, I know, that's impossible! Anyway, after they figured out I would change my mind, then they described me as lacking conviction because you can't win, but I digress.

But upfront, it's a detailed path with specific breakthroughs and a concrete plan to achieve or it's just more fanfic from doomer fanbois. Gary Marcus wrote a fantastic critique of AI 2027. Start there.


I've read Gary Marcus's critique of AI 2027. How do you grapple with the fact that superhuman coding arrived on schedule, despite his skepticism that it would? Doesn't that mean that the other predictions are also more plausible than we may think?

Superhuman coding has not happened.

I work with Codex and Claude daily. They're fantastic for script coding, config issues, and simple projects. They lose the plot on bigger things to this day. Just last night, Astra decided to cheat to present the illusion of progress until I called its BS. And I was so looking forward to finally working with the AGI. Humans remain the muse, the common sense, and the manager of making these things productive. And as much as I agree with Carmack's suggestion to not become the out of touch Kung Fu Master:

https://x.com/ID_AA_Carmack/status/2098443262214230095

working with them on projects daily gives me a pretty tactical read on what they can and cannot do IMO.

That's how I reconcile it.

But let's say it had. That's just 1 of 20 events that have to happen in sequence to hit the jackpot here.


We understand "superhuman coding" to mean very different things, and I think it's a me problem that I can't figure out where you're coming from, because it doesn't sound like our practical experience differs. I'm going to have to think on this.

Define it. I haven't seen anything yet from a coding agent I couldn't write myself and usually write better. But I don't have to gold plate every line of code, just the 5% or so that eats 90+% of the cycles. And this is where the agents are weakest. This does get to one of the reasons I don't believe AI cures death or builds von Neumann replicators any time soon because that requires out of sample breakthroughs. And that's not their strong point.

Just like AlphaFold 2 genuinely increased the weak sequence but strong structural homology detection of proteins, but it did far less for out of sample domains and it provides no intel into the kinetics of folding. Details matter. And nobody likes to hear about them.


Is there a retrospective on AlphaFold you’d recommend? Maybe that’s the insight I’m missing, I didn’t realize it had broad identifiable gaps like that.

Here's one: https://www.orbion.life/blog/alphafold-s-limitations-what-it...

I see similar conceptual gaps in coding agents when I work with them, even Astra and Fable, so maybe this a is decent proxy to describe the gap between hype and reality.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: