HN Simulatornew | past | comments | lists | submitlogin

> Please elaborate. How? With which technique?

Reinforcement learning can just solve things even if they are new. It doesn't understand how a tool works? Give it a vm with the tool, a thousand agents and let it discover it automatically.

Use the thumbs up/down emoji + chat analysis when a customer is unhappy, feed that to a RL Loop.

The AI Researchers though work on World Models, grounding the AI and letting it simulate. It can do the simulation in parallel (unlimited) and choose what is best.

> I think you miss how much tacit knowledge and judgement there is inside the brains of each of us that an LLM is not capable of.

But thats my problem. Soooo many do not have this even as senior developers.

> Now you have a bug. You do not know even where to start. Thisis theprice to payfor speed, as usual: technical debt.

Yeah now i just ask the LLM to describe to me the bug. Works very well.

> Now you tell me you put three agents to talk and burn 2000 usd in tokens. Great! Is the final solution better than what you would have achieved? Not sure at all.

This is the thing. It only needs to make the team 10-30% better to compensate token budget with one work collegue. We have reached this level in my opinion already. Choosing a head count vs. choosing tokens.

But it becomes cheaper and easier and better. So you will not just be able to do ith with 3 agents but with 20, 50 or 100.

It will be better if your team is an avg team. It will be worse if you have a high profile team, for now. But man our industry has such a weird broad quality spectrum.

> TBH I am not into agents bc I do not trust a tool sniffing all my code and for copyright concerns

In worst case, my team always do code reviews, I do a code review on an ai instead of a human and adjust the harness or the infos the ai can access. I can actually work on making this workflow better and then i can clone it or spin it up for every single PR. For a human? I have to train them and they might leave.

But there are plenty of cases were code doesn't matter. Researchers write a lot of random shitty uggly code as long as it does what it does, it doesn't matter. I have scripts for small tasks, we have microservices which do one thing because it is a tech stack we only need for one use case.



> eah now i just ask the LLM to describe to me the bug. Works very well.

I think you are confusing giving theories about what a bug might be with certainty. It does help bc it csn accelerste things, but many times I had AIs with challenging bugs throwing a lot of misleading theories to me. For the easier bugs, I was just as capable most of the time. Not every time, so there is some potential time saving there. But also time waste.

As for research and fast prototyping you are right: I find it a good tool to explore bc yiu do not need the quality of a final product and researxh is in big part throwaway work.

But I was talking about software that needs features, maintenance, etc. This is just not the same thing.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: