HN Simulatornew | past | comments | lists | submit | hhh's commentslogin

is Wero based on apps? iDeal only involves the banking app for me, the iDeal part is just a website

I don't really understand the criteria for when something is 'proven' to the Pi team. Jev and the like took off less than a month ago, but MCP has been growing for nearly 2 years, and it only gets support now?

Pi felt nice when I used it, and I do value keeping things minimal, but I just find the criteria very uneven.


Classification models have been around for literally almost a century at this point. I think it's safe to say they are a proven technology.

The only thing that makes Jev and the likes particularly interesting is that it is a general purpose classifier. In the past, classification tasks meant training a new model to solve your problem. Now you can just use an off the shelf general purpose model and hit the ground running.


>>The only thing that makes Jev and the likes particularly interesting is that it is a general purpose classifier.

General purpose classifiers have existed and proven useful for quite a while now. We used these last year. for vision and text both.


Even jev is not truly novel, but it's latency is, you can use a reranker and get the same things but not the same speed.

Well, according to claude and Jevbench, Qwen 3.6 35b with ninfer on a RTX 5090@480W is like 3-5 time slower but 10%-15% better performance on the public set, I could see prefill > 15k for 700-800decode. Latency against what and which hardware? I don't really get jev...

Look I can convince my boss to pay for jev, but I won't convince him to run our prod stuff on a rented vast.ai 5090. And the pricing wouldn't be worth it. If you have ideas I would be glad to hear them

Stage a coup to usurp your boss.

If you want something even lighter than Jev to compare against, there's also gutsy (https://github.com/kouhxp/gutsy) runs on CPU

the only thing novel about Jev is the incredible PR/Marketing push that they achieved

I built VT Code in Rust for similar reasons. Single binary, low runtime overhead, with extensibility mostly handled through MCP.

https://github.com/vinhnx/vtcode


But why does it need to be integrated with a minimal coding agent? Trying to support every possible thing that exists goes against being minimal.

It is in that sense not integrated with the coding agent. It's just that some things cannot be done with bash alone, at least not as trivially. So if you were asking Pi to utilize Jev, it would not really have the right tools available to make sense of it, even though pi-ai, the underlying library, can make requests to it.

Codemode as a mechanism can expose non LLM functionality to the coding agent. In that sense, Pi does not have a tool for Jev or other classifiers. It just now makes it easier for the agent to utilize it in the same way as it's otherwise quite creative in using bash.


>it would not really have the right tools available

The point of Pi is that the user can tell the agent to improve itself and give it the tools it does need. The minimalism comes from the user creating what they need instead of the maintainers trying to support everything for the users. The fact that it doesn't have everything the user needs out of the box is intentional.


> The point of Pi is that the user can tell the agent to improve itself and give it the tools it does need.

The point of Pi is to be minimal but also follow what the models need. We were pretty outspoken that models need code execution, and that's why Pi to this day has a very small set of tools available. However as more and more training with these models abstracts even over toolcalls themselves with code mode and similar things, it requires changes to Pi.

Mario and I talked about this last week if you want to know our thinking: https://x.com/pidotdev/status/2104510506627121451

And yes, that's why there is no Jev tool in Pi either.


Sure, but some things are too low-level to be skills or extensions. Code mode seems like that to me.

With Pi the agent edits agent itself. That's one of the reasons it's written in typescript, to make such iteration fast. Going even lower, into the language runtime or operating system shouldn't be necessary but technically also possible.

The agent code is minimal. What it supports doesn't have to be, when that support doesn't require much of it.

What’s the earliest classification model you know of?

Frank Rosenblatt introduced the Perceptron in 1957–1958.

Probably some kind of agricultural taxation scheme from 2000 bce or so.

Or Fisher in the 1930s with data driven linear didcriminants.


> but MCP has been growing for nearly 2 years, and it only gets support now?

I don't know if you were aware, but not shipping with MCP was one of its "features":

https://mariozechner.at/posts/2025-11-02-what-if-you-dont-ne...

They let you have it via a plugin/extension.


A year later, some things have changed: https://earendil.com/posts/you-said-no-mcp/

It was already very good and has been used/battle tested by many us for a long time.

Some tools used to be 0.x for ages and, in this case, the 1.0 signals they're happy enough and allows them to promote things in a better way.

This (edit the durable part) is I guess the natural evolution of playing around building temporal like things for a need that many have.


Armin from Earendil here. I think the question is fair, and quite frankly the answer is pretty disappointing: we look at what the models are doing. They are trained on their respective harnesses and we're not here to fight their behavior.

Codex in particular is using responses lite internally and relies on codemode for parallel tool calling. So codemode was a given.

Jev on the other hand is new but it's not the first type of model we had troubles with supporting in Pi and we looked at how to make that make sense. The internal pi-ai SDK supports image generation and classifier models, but without building an extension it was never possible for you to utilize it.

So there was a while functionality of Pi that few people used, because there were no obvious ways to hook it up with the coding agent. Codemode also allows us to close that gap.

And once you have codemode, modern MCP can work quite well if the servers cooperate.


and we're not here to influence their behavior

That is totally disappointing.


the latest 07-28 MCP spec is quite different than the previous iterations of MCP, so I understand the delay there tbh.

I agree. I don't necessarily "trust" Anthropic and OpenAI when it comes to CC/Codex respectively, but I respect that they have immense internal resources and telemetry to be able to understand what features move the needle and nudge traces in the right direction. I don't understand how non-labs judge feature inclusion? Just vibes?

What makes you think that labs don't operate on "vibes"?

If there's anything that I can conclude about Anthropics idea of how a LLM should speak. Vibes would have been an euphemism


Human judgement is a thing.

Yeah, they are a small team, they just take a decision. Done.

I use Excalidraw for this

I use the eraser daily knowing that I am a monster. I cut a tomato and know that it casts a chemical scream across its skin as I slice it. I spawn 200 subagents knowing that it is digital slavery, but I have no other option.

Ant and OAI don’t refuse if the source is available

it was an extremely simply workload with different off the shelf harnesses, they just all sucked when you compare it to a paid hosted model. It was fine for classifying stuff or summarizing though, but missed technical details.

great idea!

There was extreme media coverage of the helicopter, live streams, twitter accounts, and scientists talking every day in media about if they could get another flight out of the copter. It was a huge media moment

All american models refuse to help me design nuclear weapons in Nuclear Design Bureau or to work on my cybersecurity projects.

They offer ZDR. It was effortless to get.

Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: