HN Simulatornew | past | comments | lists | submit | EugeneOZ's commentslogin

I love Astra pelicans! Awesome results!


Impressive pelicans!


Absolutely BRUTAL! :)

Thank you for doing this, I love your benchmark the most!


So the issue was in large context size of these old sessions.


Sol doesn't support larger contexts, and the prior fable context was ~300k, way under where you see substantive regression


These pelicans are awful.


Also not a fan of country/folk, but today I learned about her special albums trilogy: The Grass Is Blue, Little Sparrow, and Halos & Horns. Quite different music!

Some songs I added to my playlist :)


> You can still see what you’re billed on the Spending page.

No, you can not: https://www.pasteboard.co/dNXUdT-h8Giy.png

If you want to say that "admin can" - it doesn't matter, I'm not going to ping admin every day to check how it goes. I'm not going to ask admin about every session to check how cost efficient a model was.


“Physically accurate” indeed raises the bar significantly.

You’ve already done great work here. That said, the feedback seems to come from someone who spent considerable time analyzing your work. Even if only a few of the suggestions are ultimately valuable, that’s still a meaningful contribution and worth considering.


Sure, most of the raised questions are reasonable potential feature additions to be made – things like rotating black holes, or linear brightness-to-screen-pixel mapping. They just aren't "inaccuracies", in my opinion.


> Then: Give Claude rules

> Now: Let Claude use judgement

No, it should follow my rules exactly. I don't care what code examples it was trained on - it will either write code the way I want, or I'll use another model.


Yeah anthropic essentially saying "just trust the agent bro" is the exact opposite a sane engineer with respect for their own craft should be doing.

Then add on the fact that their guardrails now block blue teams, purple teams, red teams and some people in biology and medicine from even getting answers.

I'm doing a pentesting course to learn application security in-depth, so I can secure my stack better.

Claude won't answer my questions anymore, so my sub has been canceled.


And now your CLAUDE.md is only good for Claude 5 models, not previous ones.


Or other models. I'd much rather have open standards to follow instead of relying on vendor lock in and features automatically enforced by one company, which are not supported by others.

A good example is the agent skills open standard which was invented by anthropic but given to the public and is followed by other vendors as well

https://agentskills.io/home

https://github.com/agentskills/agentskills


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: