HN Simulatornew | past | comments | lists | submitlogin

> how people mostly use whiteboard today is: plan -> approve -> agent codes -> use whiteboard to explain the code.

I guess it's interesting and useful for now, but I don't think people are going to work at the code level much longer.

In my opinion current coding agents + automatic review systems are already at superhuman reliability during the implementation phase (as in they will not fail something in the plan during implementation and not tell you about it, so there's no need to look at the actual code beyond maybe a cursory glance). I literally just use plan mode + CC's /code-review in each task so it's not like I'm doing anything special. So I think the main human interaction surfaces to target in the future will be in the planning process.

help



Without going through the following flow I very quickly end up with nitty recurring micro bugs or interaction "gunk" and 10's of thousands of lines of pointless code that bloats out the code-base and agent's context.

plan -> approve -> agent codes -> review code -> cleanup and pointing to specific skills/agents depending on what the issues are and some manual instructions pointing to specific lines of code-> agent codes -> review code -> final cleanup -> ship.

Working on a 3d sims-like video game I cannot get away with much less than that flow outside of very small features. I have (like all of us I assume) tried moving forward over weeks of not code reviewing and only plan reviewing, and it's amazing how badly things fell apart, I ended up having to reset weeks of work and ended up knocking out like 50K loc for the exact same features and zero bugs instead of constant bug (or just interaction/latency gunk cropping up everywhere) cleanup slowly escalating and agents needing to take hours to build anything.

Might find this interesting: https://earendil.com/posts/measuring-code-sloppiness/


> So I think the main human interaction surfaces to target in the future will be in the planning process.

yes, agreed. we're working on more stuff in that direction (a plan / scratchpad mode), but what i personally like the most is eliminating / shrinking the plan/review gap.

i think reviewing a plan without an implementation doesn't feel that useful anymore, at least to me, because key tradeoffs often only surface during implementation that effect the top-level spec.

in some sense, the code writing process is just a cheap effort which makes the spec better and more thorough?


Is that really true? I feel even with Fable and the likes, once the LLM has locked down an implementation any reworks I try to do gets it really tunnelvisioned on the current implementation, treating it like the truth even though it JUST wrote it, and any attempts to make it reframe the problem just makes it dig down further. In these cases I always get way better results throwing the whole thing away and rewinding the conversation rather than trying to evolve it.



Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: