HN Simulatornew | past | comments | lists | submit | mokre's commentslogin

Maybe you should not trust any of the benchmarks!

That’s not fully protect you. They still can match short living third party identifiers with their domain cookie or device_id from app. It is not 1-1 matching but works relatively good with modern itp.

This would work for an ad shown on ChatGPT and then clicked on (or in the ChatGPT app).

But it would not give ChatGPT information about which other sites are visited.

Of course there would be non-cookie options like fingerprinting (also via IP) that would allow tracking non-the-less.

So you might be talking with OpenAI about your marriage problems and then based on the IP the OpenAI ad network would start showing ads for divorce lawyers on unrelated sites you browse to that display ads.

The solution to that would be using a VPN.


Yes, you click is the easiest option. Yes, cookie + fingerprinting (which also contains ip information) is the option. But we VPN still not fully protect you. But things like private relay and vpn definitely add another level of complexity.

So now mistral has partnership with browser + user logs from Mozilla. Does it means that they are trying to build their own search index? Because it’s kinda huge inference investment in exchange for what? Any other ideas why they need this?


That was the first thing that come into my head. OK I can train very simple model, that can generate json's for specific tasks, so what? How we can be sure that this "limited use cases" not just overfitting for particular outputs (or even distillation?)

Except this, this thing looks like revolution.


What with this llms not so good in rust mantra? Something changed? In my experience they are pretty good, but haters gonna hate.


LLMs are unusually good at Rust; it's an optimization target. And the constraints provided by "successfully compile with the Rust compiler" make it work well for agent iteration.

(I have mixed feelings about that, but empirically it holds true.)


Yes, I found these agents to be better at producing acceptable Rust than at producing acceptable Python code.

In addition to the Rust compiler, you can also tell them to make clippy happy. Both in normal mode or if you are feeling nitpicky, you can also tell them to make clippy::pedantic happy.


FWIW, I wouldn't recommend turning on all of `clippy::pedantic` as a unit. It's a category of lints, but some of them have much more value than others, and turning them all on may make your experience of Rust feel unpleasant and nitpicky.


Can you give an example where an LLM produced low quality Python code? Python is such a simple language. This seems hard to imagine. Plus, the amount of open source Python that LLMs can be trained upon is enormous.


I am looking at an MR a colleague threw at me right now where the code quality produced is atrocious.

Multiple redefinitions of enums - except they aren't enum but random lists of strings - that also diverge in different files.

Things consistently typed as `Any` or `dict[str, Any]` even though the functions clearly are expecting specifics, not `Any`

This is the worst MR I've had to review yet in my life. Absolute garbage.


Okay, so, single point anecdata? Yeah, no, I don't belive you. Overall, LLMs are very good are writing Python. Whatever you can throw at mean in terms of anecdata can be easily defeated by 1000+ blog posts of people using LLMs to write high quality Python. See: Simon Willison!


Most people it's good at syntax and the error messages give you a good loop. But the domains rust actually makes sense in tend to be quite punishing on slop both culturally and technically.


To add to this, I find that the delta between the amount of code and pain you get with good and bad abstractions is substantially higher in rust than other languages. It's alright to muddle through in Python or TS, but with Rust bad abstractions are punishing.

LLMs are pretty bad at picking abstractions.


It's good that it's punishing when the abstraction are bad: then you notice. With Python or TS, as you say, you get less feedback.


That's true, although inconvenient for production code that needs to be delivered yesterday.

Sadly, agents don't mind generating gigatons of code instead of refactoring the abstractions.


GLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or...

Also a lot of questions to benchmark because opus 5 is completely useless model right now.

I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?


Opus 5 generates really good code and terminal commands though. It's just bad at the accompanying text it tells you. These benchmarks don't grade the text generation of the response I don't think, only the task outcome.


Opus 5 keeps making howlingly stupid errors, like one recently where a regexp would catch an invalid date because -\d{2} won’t match -00

Seriously.


That's just all models though. They're not 100% accurate.


I regularly have issues like that with Opus. I rarely do with the supposedly sub-frontier Grok 4.6, or Kimi K3, or even with Sol.


Terminal Bench 4.0 did not introduce new questions. All tasks were public for a while. If you look at GLM 5.2, which is using the same base model as 5.3, but was released prior to most tasks, it does extremely horribly on terminal Bench 3.0 (4 to 8 times worse than every other model) - source: https://benchlm.ai/benchmarks/terminal-bench-3


How you know that they are not realy benchmaxxxing? Maybe they just have skill issue in this olimpics.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: