HN Simulatornew | past | comments | lists | submit | menaerus's commentslogin

You're overconfident in what you're saying. Nobody really knows (yet) if LLMs can or cannot "think" or can or cannot "invent" new things. I'm inclined to think the opposite from you but smart people actually try not to hold the absolut stances and present them as facts when they're in fact not. What I have seen so far throughout a daily use on complicated things suggests that humanity developed a new form of an intelligence.

Leetcode was designed to collectively implement bias towards young talented people whose only obsession at that point of time is their work. So no kids, family, free time, and no financial freedom (mortgages, loans). Once you build such culture, it starts to grow and gatekeep by itself from within.

Last time I checked it didn't make any difference or whatsoever when compared to lld so I would not go that far by saying "mold is a reference in the world of linkers" - made a test just few days ago on 10M LoC C++. So, rewriting the codebase in another language just because doesn't seem like an investment worth doing when there's much bigger fish to fry.

On 22nd of July

> It was loading the data for a week already, and the speed of loading data has dropped to 4 kilobytes per second, and it will take years to load the dataset.

And then on 2nd of August

> It loaded maybe 1% of the data so far.

And then one year after

> I didn't load the data after a year. > I tried it with the fsync removed, but even then it does not work

Ok, I find this more entertaining than I should and almost unbelievable. I know that this codebase was heavily written by LLMs but the execution can't be this bad?

What am I missing and why would Supabase buy the product which can't even load the dataset properly?


You're looking at a software project months after its inception, saying it didn't work well then, and then asking why someone would buy it well over a year later...

1st commit of the project was in August 2023 so not quite few months old. And the performance problem according to https://github.com/ClickHouse/ClickBench/pull/1159 has not been yet solved.

So, I think my curiosity still stands.


> 1st commit of the project was in August 2023

Yeah, ok, I took the date from [1] but you're right the project did in some sense exist before then.

[1] https://turso.tech/blog/we-will-rewrite-sqlite-and-we-are-go...


Also hilarious that the rebuttal comment after a year had passed was "why are you using this release which is a year old"

That rebuttal comment was posted in response to a comment that was 4 days old, not a year old.

Do you make your professional career by building formally verified systems? I ask because I don't think that the reason comes down to "because most devs don't want to do formal verified system development". It's much more complicated of course.

I could believe it. Verified system development most likely comes with metric ton of paperwork.

Want to merge the PR? I need verified sign off in ServiceNow by staff level engineer. They are on vacation for 2 weeks? Did manager fill out delegation paperwork in ServiceNow with VP sign off? Oh they did but they forgot to put in return date AND time. Form needs to be corrected and reapproved before we can go into ServiceNow and make changes.


I had it seen on servers that were typically reserved for the team but there was no official booking system for those machines. When you start using the machine you would typically put some note to make sure somebody else does not overrun your long-running tests or performance measurements.

Too simplistic view which I also used to believe in maybe ~10 years ago. There's a class of problems where such approach doesn't apply.


I agree with you in that there exist such a class of problems, but Occam's razor apply in far more circumstances than it does not. When confronted with novel information under pressure of quick decisions, you have to take the most likely correct path.


> LLMs are not suitable for that kind of work.

I wonder why not or you meant not suitable yet?


This is just speculation on my part, but LLMs work best when they get immediate, verifiable feedback on their task, and the kind of physical optimizations they mean might not give that to LLMs.


The right way is to throw LLMs at building tools that reframe the problem into a shape LLMs are good at navigating, and then have LLMs use those tools to solve it.


Also a speculation but I'm almost certain that physical optimizations are first done through simulators running on a computer.


Yes, they are, but the most important subtasks of designing a CPU are not physics related. They are picking the right parameters for things like: how wide do I make this bus, how many registers do I put in the register file, how large do I make this cache, how deep do I make this pipeline, etc., etc. To find optimal parameters requires a lot of simulations, and humans do this, but LLMs could do them just as well and maybe better because they excel at tedious work.


How do you know this is not true with other vendors? I'm not defending them but I wouldn't believe anyone in this business unconditionally. Anthropic agent fwiw is not open source, gemini and codex are.


People have found many nasties embedded in Claude Code over the last couple of years. You can't trust a closed source harness. You can barely trust an open source one.


They do invent code solution for the problem that exists in your codebase. Latter implies that the code solution LLM synthesizes is usually unique of a kind, so, it's not a copy-paste neither it is a simple extract from "another codebase" and adopted.

IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.


LLMs can output near-exact segments of copyrighted code used for training.

https://arxiv.org/html/2408.02487v3

I wonder how would Microsoft react if someone would synthesize a code solution based on Windows source code.

https://en.wikipedia.org/wiki/Shared_Source_Initiative


I guess you're not writing code much or haven't done much so as your professional career?


My argument is if AI companies are ignoring copyright law and looking at all training data as commons, then we should look at LLM output as something that is not protected by copyright law.

Of course I'm a bit naive here, because we are talking about the richest companies in the world with lot of money to spend on lobbying (or bribes).

https://www.theguardian.com/technology/2026/may/23/trump-ai-...

https://www.bbc.com/news/articles/c98r8r7dz5no

https://en.wikipedia.org/wiki/Commons


I wanted to understand your background first because what you initially said is a very oversimplified view of LLM mechanics, and generally not quite the way how software is in practice written. Since you didn't answer that question, I will assume that you're not a SWE by a call. To give you an example of what I am trying to convey is: imagine a data-intensive workload hitting your storage/database/kernel implementation, and it's painfully slow, your customers are not happy. Then you as engineer sit down, spend days profiling and understanding the code, researching about existing algorithmic solutions to the same or similar issues found in the wild, you read some open-source implementations of viable approaches, you ditch some, some you take, you also read books, articles, other peoples experiences etc. And finally you end up, let's put it bluntly, with some sharded data structure by which you solve the bottleneck. It's not novel, the technique is so common and is already implemented across many many different products in slightly different flavors so I am wondering why do you think this is not a copyright breach but the LLM, which does more or less the same thing, is?


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: