Whenever I try to make 'AI agents' do anything it seems like a struggle for even basic tasks. How on earth are people having it do anything successfully, let alone really complex tasks like loading and navigating a website to buy tickets?
This article just sounds like an Ad for Muse. It's painful reading and just sounds awful.
What model do you normally run the subagent on? You mentioned flash as well for that, but I wonder if a more 'strict' model would do a better job at pushing the main back on track.
For cost reasons, I’ve been using Flash for the reviewer, too, but I plan on trying to use Pro for that. Thus far, however, Flash has been doing well at reviewing. I’m cheap as I’m paying for all the tokens myself.
reply