I review the code that matters - anything security adjacent or that's an API that will be used by other code in the future.
I don't review code that either works or doesn't - most HTML and CSS layout code for example. There I test it on desktop and mobile and commit it if it works.
Ditto for stuff that's simple. A JSON endpoint that runs a SQL query and returns some JSON? If it works and a glance at the tests looks OK then I trust my agents wrote it properly.
I'm getting more confident with my judgement over what needs a close look and what doesn't over time, as so far I haven't been majorly burned my any mistakes that snuck through.
Honestly, it's similar to being an engineer on a larger team. You don't review every line of code written by every one of your coworkers.
I think this is THE issue of our time as programmers to be honest: do you review every line of code an agent writes?
An increasing number of expert programmers are moving in the direction of NOT reviewing every line. It's working out OK for a lot of them.
I've found that engineers on a large team do read every line, mainly due to the fact that the skill levels run the gamut from intern to lead, and only 1 or 2 people out of 12 might have knowledge of the application being modified.
It's actually getting worse due to "AI code bloat", for example I have 16k lines of code to review across 3 apps by the end of this week. Normally it would be a quarter of that, but what Claude produces is extremely verbose in some places and anemic in others, and I can't tell at a glance what's right and what looks right with that much ground to cover.
Goodness, how is that being tolerated? I guess it can’t be stopped without a lot of political capital; but 16k lines of code is HUGE, and I cannot imagine that it’s actually 16k lines of value - I’ve written whole new subsystems of a product in fewer lines. Are these all written in an exceptionally verbose language like Go or Java? Are they VERY well documented? Are they doing things they shouldn’t be doing???
I mean, I've done that for a couple of things, but probably 90+ percent was vendored libraries and javascript which could be ignored. Sadly, this is not one of those cases, and Im guessing yours isn't either?
> I don't review code that either works or doesn't - most HTML and CSS layout code for example. There I test it on desktop and mobile and commit it if it works.
Good example of what not to review if you're working on your hobbies. Also exploratory can sometimes be done this way. However, this ultimately boils down to how you approach programming as an engineering discipline, including your responsibility for the outcome.
> I'm getting more confident with my judgement over what needs a close look and what doesn't over time, as so far I haven't been majorly burned my any mistakes that snuck through.
This doesn't generalize well. If you drink raw milk, or if you don't wear your seatbelt, or if you don't escape your user input correctly, you'll probably be fine, but I really hope aspiring programmers/engineers don't take this attitude towards any serious task. One should always examine their biases, tools' failure modes, etc. regardless of how many times something didn't fail.
> Honestly, it's similar to being an engineer on a larger team. You don't review every line of code written by every one of your coworkers.
One [should] review the code they're responsible for. In a team, people usually assign you (or ask you) to review code, and the work is divided accordingly. If the code isn’t reviewed by the code owners, it’s a problem, not something inspiring!
> An increasing number of expert programmers are moving in the direction of NOT reviewing every line. It's working out OK for a lot of them.
Have you considered that the sheer amount of code being generated is what makes thorough review infeasible, not that it’s a desirable approach?
But if you observe that the agent day after day do handle user input safely; and also routinely run an agent that scans for security vulnerabilities and observe it finding cases where input is not handled safely in existing code, you may conclude that the chance of an issue is at the same level at, or probably lower than, if a human wrote it and a human reviewed it.
("Escaping" user input is not good practice though, use parameters, assuming you are talking about SQL.)
> Ditto for stuff that's simple. A JSON endpoint that runs a SQL query and returns some JSON? If it works and a glance at the tests looks OK then I trust my agents wrote it properly.
That is *exactly* the sort of area I *wouldn’t* blindly trust AI, there’s a huge security boundary there. What if the AI is doing string concatenation with user-provided data???
Is this really a rational strategy for something whose nature is to be right most of the time and then spectacularly wrong a much lesser amount of the time?
I agree with a commenter above/below (depending where this comment lands). For some time models won't do this. And any review from review agents would caught this.
For most of the AI programming there needs to be a more stricter (automated) review process now. Most SAST tools would caught this type of security issue.
I think it also depends on what you're building. Some solo project or basic html thing? Sure no need to review every line. It's a bit different when you're working on foundational libraries that a business relies on, anything touching a production database, etc.
> Honestly, it's similar to being an engineer on a larger team. You don't review every line of code written by every one of your coworkers.
We don’t because everyone is accountable for his or her own mistakes. So everyone is incentivized for their recklessness to not be the root cause of some bug.
> An increasing number of expert programmers are moving in the direction of NOT reviewing every line. It's working out OK for a lot of them.
Have you ever asked your users? What about bug reports? Is the amount and rate decreasing?
lol I think you’re setting yourself up for failure. Why? Just because something works doesn’t necessarily know you the boundaries of it.
What scale does it work for?
Will it crumble under load in prod?
It works but allows cross tenant access (security issue) because security checks weren’t in the location you thought… Dangerous!