Yup that’s quite literally what ‘agency’ is and the whole point of agentic workflows. Personally I’ve had them running for days with good results and as you can see OpenAI had them running for months, and yes indeed the things made some very questionable decisions and assumptions… but they unquestionably did a lot of stuff correctly, for some definitions of ‘technically correct’.
Personally I don't believe in agentic workflow. I don't think that's a coincidence that both OpenAI and Anthropic chose math problems to test their long agentic workflows, they are well defined, with a clear finish line and with a 100% clear progress path, most of real life tech projects aren't like that.
No matter how clever the model is, most problems have multiple valid, invalid and unclear decisions to make, running it for a long time is just picking the first option on everything, which isn't usually what you want
Regardless of the model, running it for hours means that the model will takes decisions and assumptions alone instead of you.