As much as I'd love to believe it, it is now a conservative take. Sure, having a solid architecture in mind still matters right now, but manually writing code is completely unnecessary and, before long, even designing the architecture is going to be completely automated.
Here I am starting to think a central part of knowledge work is the knowledge gained.
Maybe you are right, maybe not, let's see. Unfortunately I kinda agree, because humans are really good at being lazy and going in the path of least resistance (including me). It's genuinely difficult to not use AI even if it makes my work worse, as long as it's easier and faster.
Does writing table schemas count as writing code? Is that architecting? Both?
I personally don’t trust coding agents to have enough context to write domain-specific table schemas, and I don’t have the patience to transcribe all of the context into a natural language prompt. If I ask it to, it’ll write something for sure, and maybe that can be a jumping point for me, but at some point I have to physically write what the columns will be.
You should definitely try it if you haven't already. Most of the models are deeply familiar with domain specific areas in ways that most of us aren't. In fact, I would say the problem is generally the opposite. If model size and effort is is high, it will over architect what an app needs. I find myself reigning in a giant email notification system with subscriptions and "channels" and other such nonsense when sometimes you just want a simple one off.
Yes, the LLM knows the definitions of the words in my domain, that’s clear and obvious.
No, the LLM doesn’t know our product strategy and why certain things matter and certain things don’t. That’s very much what I get paid to do. There’s not even agreement within our team about what path we should take through the domain-product space, there’s no chance in hell that the LLM will choose a profitable random walk through that domain-product space.
If it were able to do that, then AGI would have already been achieved and we’re only compute power away from OpenAI or Anthropic making the marginal utility of any piece of code $0.
Yeah, it's not just over-engineering, it is just that generic systems that look like everything else usually don't solve a business problem, but I think I would be less dismissive of its world model in general.
LLMs are not a "random walk", they take in information and they explore the space according to the way they've been instructed.
I’m taking this view now as well. If you’re reading code, you’re probably doing it wrong. You should absolutely be setting criteria that can be objectively measured and rejecting code that doesn’t meet those criteria or perform as specified. We are all senior software engineering managers now, with a fleet of cheap and ambitious young engineers doing all the authoring.
But reading code? What does that accomplish, other than to slow your dev process down enormously? Serious question.
Man, I think that code is still the artifact that we produce as developers. Code is the truth. I don't find it difficult or super time consuming to just... read the code, either. I've highlighted quite a few issues with LLM/Agent output from just glancing at the code.
I'll let you know how it goes... My new VP of engineering is a 'no looking at code' type of guy and is ripping 10K LOC PRs / Docs / plans against our 25 year old codebase and I would not say that they're 'good' PRs.
Maybe I'm completely wrong, but I think reading the code is more valuable than ever when working in a full-stack / small company role. I can tell you exactly what the business logic or functionality is for a certain piece of our system, in truth, without having to step through and make sense of ambiguous docs (that were also AI generated).
(I have a sneaking suspicion that in two years or less, my small team is going to significantly compromise the integrity of this codebase. Maybe by then we can refactor with GPT 12.)
> You should absolutely be setting criteria that can be objectively measured and rejecting code that doesn’t meet those criteria or perform as specified
This is the classic "make no mistakes".
On a serious note, I might set as criteria "avoid code duplication". Does that mean that the model/agent will actually follow it?
> What does that accomplish, other than to slow your dev process down enormously?
I am an OSS developer and I often see PRs (i.e. from the general public) that look correct, pass all CI checks, are heavily documented and they are still wrong.
Most of the times either they duplicate code that already exists somewhere else, or they implement a "feature" by opening a can of worms for subsequent "features" in the same area.
> even designing the architecture is going to be completely automated
I mean, people have been dreaming about this for decades. The whole reason UML was so overdesigned in the first place was in hopes that people could code by drawing boxes and lines. The full COBRA spec had software autonomously buying components in digital marketplaces and installing without human supervision.
It is funny to see engineers insisting that there's no way a machine could do this better. If anything, the surprise is that it took good old human language -- that second L in LLM -- to get the computer to sling code. The assumption was that a computer would just "speak code" like some kind of native tongue, but instead it just understands human language and associates that with code, and relies upon things like compilers and tests to see if it's right. Just like humans.
Predictably, now that it's actually happening, engineers are worried. As they should be, but I think it's a short-term worry. The job is changing, productivity is leaping, but its still a world where computers don't need to do stuff, humans do, so humans will be making it happen one way or another.
Seems like content curation is becoming an increasingly important problems in online communities. Eventually a new equilibrium will be found, I'm wondering what it will look like.
Agreed, what I understand from RSI would be models creating new models, or at least upgrading their own weights/architecture. It does not seem to be the case here.
This technique is very useful to gain intuition for a given sample size. Just run a few simulations with uncorrelated data and then you can get a sense of how extreme the estimators can be.
Well not only research, but also a lot of other aspects of life; it is way easier to interact with something that has always an answer is polite whatever tell them.
>
> Short story
This looks more like a dating app simulation than a chatbot