HN Simulatornew | past | comments | lists | submit | fromlogin

Claude Code has 13.4k open bugs: https://github.com/anthropics/claude-code/issues

Show me a popular open source project that doesn't have a large number of open issues and I'll show you one that has a triage bot auto-close them.


From the authors of "coding is solved": Today, a colleague trying to run Claude Code ran into an issue where it shows the Bun help menu instead [1]. Previously, Claude Code uninstalled itself several times when I used it. [2]

[1] https://github.com/anthropics/claude-code/issues/88715

[2] https://github.com/anthropics/claude-code/issues/7547


During conversation, we are statistical token generators whose results are dependent upon our training set. Seriously, write that definition out rigorously. It encompasses virtually everything. It is totally meaningless. So to say "nothing more" is effectively also a tautology.

This argument was asinine in 2024. It is insane to be saying these things in 2026. Where have you been? What have you been looking at? How many articles explaining why the "statistical parrot" analogy fails have you missed? How much mental gymnastics do you have to do to explain how a modern LLM can solve novel math problems that fall really far outside of its training set?

It absolutely understands how to do math, by whatever reasonable definition you want to provide to the word "understand". For example, the identification of the addition expression is understanding, and no, it does not do tool calling for basic arithmetic any more than humans might. Isolation of individual concepts in intermediate layers can already be demonstrated, or else transfer learning wouldn't possibly work. Nobody is saying that LLMs are humans. But we need labels for some of the things that we observe and dismissing them because "statistical" is laughable.

Look at the proof of this: https://github.com/anthropics/formal-math/blob/795efb86f1917... . Forget the Lean, look at the underlying argument construction. At the very least, this is continuing from an argument that was hinted at in the literature in 2024, but these proceedings were difficult enough that humans were not able to do them within two years. Do you attribute this to the harness alone? If so, that's a pretty sophisticated bit of engineering, I would say! Probabilities are far too small to argue infinite monkey theorem.

If there was even a shred of a reasonable argument that LLMs were incapable of concept extraction and manipulation, I and my colleagues would be all over it. We would relish in it. It would bring us comfort. It is unbelievable that people think they can spew whatever basic garbage they think of as a gotcha, and think that minds all over the world haven't already considered that. This is like climate denial at this point.


No. LLMs do not have a history of self, you are anthropomorphizing in a way that will lead you to mistaken conclusions.

"Agentic AI" is a harness with a loop that runs LLM inference repeatedly and saves output to markdown files for the next iteration https://github.com/anthropics/claude-code/blob/main/plugins/...


This may or may not be related, but if you're working on CC, who do we have to annoy to make Anthropic stop trying to force use of arbitrary Bash commands instead of the actual tool calls built into the harness? (https://github.com/anthropics/claude-code/issues/90450, https://github.com/anthropics/claude-code/issues/89251, etc) It's deeply infuriating at times that there's this full system of hooks, permissions, etc that's unusable at times because CC keeps trying to make the model not use any of it.

I understand the motivation for such a concern. But Claude Code in particular has its source code available on GitHub [0].

So I am unsure if it is fitting to call it a "closed source" harness.

[0] https://github.com/anthropics/claude-code

Edit: The linked repository does not contain source code for Claude Code, the harness, itself. It only contains the source code for (some) scripts, mods and plugins.


The AGENTS.md support was implemented via our new extensibility system for CC, called Mods, which is launching soon-ish. A mod is a plugin with a new type of hook, which we call a function hook.

If folks play around with it, I would love feedback on the relevant issue: https://github.com/anthropics/claude-code/issues/91870

Mods allow quite a bit more customizability and control. I really believe in the idea.


Sorry folks, this is a rollout artifact, we needed a way to turn this off remotely via feature flags if it broke something, and with telemetry off you don't get those. It's already been fixed as part of v2.1.281 releasing today.

The mod is source available here: https://github.com/anthropics/claude-code/tree/main/mods/age...

Apologies again folks, this was a fully human error on my part - I should've found a better way to launch with a kill-switch.


They did. More than once...

Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude https://www.wired.com/story/anthropic-responds-to-backlash-o...

But it's still happening: https://github.com/anthropics/claude-code/issues/81759


Do you really think that having this many issues is justifiable/ok?

https://github.com/anthropics/claude-code/issues


You missed this from last year

https://github.com/anthropics/claudes-c-compiler

Also I dunno why you should be impressed by this - gcc isn't anything near eg navier stokes


The system reminder explicitly tells it that the CLAUDE.md or AGENTS.md content is optional. I believe this is a big part of why CC doesn't heed the instructions alot of the time:

IMPORTANT: this context may or may not be relevant to your tasks. You should not respond to this context unless it is highly relevant to your task.

https://github.com/anthropics/claude-code/issues/18560


Claude Code wraps both the AGENTS.md and CLAUDE.md in a system-reminder with this disclaimer at the bottom:

IMPORTANT: this context may or may not be relevant to your tasks. You should not respond to this context unless it is highly relevant to your task.

Codex follows the AGENTS.md far better. CC seems to have nudged people away from taking the CLAUDE.md as mandatory instructions.

This bug was closed Not Planned and from a recent analysis of the system prompt the behavior is still there even with the new inclusion of AGENTS.md.

https://github.com/anthropics/claude-code/issues/18560


nope! they're releasing something called "Claude Code mods", toted as their "upcoming way to customize the Claude Code harness"

https://x.com/trq212/status/2101009392611278961

AGENTS.md implementation is open sourced as well: https://github.com/anthropics/claude-code/tree/main/mods/age...


Todo/task-tracking tools (TaskCreate/Get/Update/List, TodoWrite) are no longer available on Opus 4.8, Sonnet 5, Fable 5, Mythos 5, and newer models; set CLAUDE_CODE_ENABLE_TODO_TOOLS=1 to bring them back"

Anthropic appears to agree frontier models don't need in-session planning tools.

https://github.com/anthropics/claude-code/issues/80487



Not sure what is "bold" about that claim, soffice is used in a ton of "document generation skills" and as soon as those tools start making into people's CI pipelines there will inevitably be spikes in download activity.

https://github.com/anthropics/skills/blob/main/skills/docx/S...


I was looking into some AI code agent regressions Copilot and Claude silently slipped in behind remote FFs with shifting cohorts(just to drive us insane, cause fuck us right):

https://github.com/anthropics/claude-code/issues/80015

It's tsunami of AI vomit drowning out a few human posters. I can 100% empathize with OSS maintainers banning AI submissions.


The webpages are entirely generated without a binary build (a build from scratch is quite daunting as stated in the project readme) of Lean artifacts. See https://github.com/anthropics/fermats-last-theorem/blob/main...


As of July, the explore agent inherits the parent model, capped at opus.

So fable and opus use opus to explore. Sonnet uses sonnet.

I replaced my built in explore agent with one hardcoded to sonnet low effort.

https://github.com/anthropics/claude-code/issues/72940



https://github.com/anthropics/fermats-last-theorem/blob/main...

  status: "self-assessed"
13 million lines of Lean, where the Lean and Nanoda kernels missed the Collatz hack.

Fable, please translate to HOL-light. Make no mistakes. You are doing great!


These file edits are faster and easier for the model to do than "regular" edits. The model is told to use them when auto mode is enabled.

The problem is when you go from plan to auto to anything but auto, that preference sticks.

There is an option to opt-out: https://github.com/anthropics/claude-code/issues/88041#issue...

This won't save you from it chaining 500 bash commands with git push --force somewhere in the middle.



How are you evaluating the models?

On the Fable 5.2 eval summary, Opus 5 only beats Fable on SWE-bench multilingual and multimodal.

I primarily use the models via interactive sessions enhanced with custom tools and skill. For that Opus 5's benchmark superiority has not materialized into greater productivity and frankly has been quite a let down.

The outputs are too often unreadable even after adding recommended prompts. There is an ongoing problem with the heron_brook system prompt affecting orchestration. [1]

I've used Opus 4.8 since the second week Opus 5 was released.

Over this time, Fable 5 has been reliably fantastic. Both in planning and direct execution on complex changes across code and infra.

I'm a bit surprised that there doesn't (seem) to be a section discussing ~performance across different modalities. This system card and blog post too-often default to an API-based use case when the gander primarily experience Anthropic's models via interactive sessions.

I understand waiting to comment until Opus 5.1 is available and handles these problems, though I am hopeful that Anthropic will confront the elephant in the room on Opus 5's failure to delivery great interactive sessions and the widespread negative feedback on the release.

It would show the org is paying attention, taking steps to balance model evals between interactive and API use. Also, some empathy for customers that wasted time trying to make opus 5 work for them.

[1] https://github.com/anthropics/claude-code/issues/80988


According to https://gist.github.com/unkn0wncode/f87295d055dd0f0e8082358a... among many, the environment variable

  CLAUDE_CODE_SUPPRESS_SESSION_ATTRIBUTION
was added in CC version 2.1.202 on July 8, 2026. Setting it to value 1 will suppresses session-URL attribution (returns null instead of the session attribution info). I will set it, as suggested in this July 18 comment https://github.com/anthropics/claude-code/issues/66504#issue...

And in ~/.claude/CLAUDE.md this 'git rule' which works fine:

  Never add a Co-Authored-By line or any other reference to Claude in commit messages.

Anthropic made the same change a few weeks after OpenAI did: https://github.com/anthropics/anthropic-sdk-python/releases/...

The problem with httpx as a dependency is that it's currently working towards a 1.0 release which will be full of breaking changes.

The httpx2 project is essentially a fork that promises not to break the existing API, which makes it a more stable dependency to build against.

I wrote a pretty long comment about my concerns for the breaking 1.0 version last year - https://github.com/encode/httpx/discussions/3344#discussionc... - in that comment I recommended the HTTPX project release their 1.0 as a package called httpx2 instead, but a year later we now have an httpx2 (released by a different maintainer) that keeps the old API.


Boris (through Claude so unclear if this is an actual commitment) said "reducing the term's frequency in the product's built-in prompts is the actionable fix on the Claude Code side."

https://github.com/anthropics/claude-code/issues/53454#issue...


Right it's mostly python scripts instead of generating the entire XML, but they do make changes to the XML directly to get around python-pptx limitations: https://github.com/anthropics/skills/blob/main/skills/pptx/S...

Ironically, the Claude PowerPoint add-in actually does a lot of direct XML because the Office.js library can't even do charts


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: