DHH · 2026-08-26 · source ↗ · whole document (8)
Programming with AI agents
I give it a couple of tasks, and I realize that the quality of the output is uncannily close to what I would've written. And I remember just leaning back and thinking, "What just happened?" "How did we go from this autocomplete mess that I was talking to you about in the summer to this just a few short months later? How did we get both the increase in intelligence and then also the increase in usability?" This agent harness question where I don't know if Opus 4.5 was that much smarter than Opus 4, which is what we had in the summer, but its ability to instrument your computer, to use tools, to check its own work, to apply its intelligence in such a way that you could get real meaningful work out of it, was completely different. And I think this is then the big change that happens for almost anyone who paid attention and started playing with it over the Christmas break. This is what I heard-... from Shopify and other places with lots of employees who suddenly had a breather, suddenly had a couple of weeks to lean back and just look at what was going on, gave it a try, and had the same mind-blowing experience that these agents were of a different genre than what we had before. And then what we get to is, I'm already excited at this point. Like, by December, I'm already revisiting sort of all my priors-... and going like, "Wow, if it can do this, can it also do that?" "Oh, yeah, it can." And again, Opus 4.5 now looks like a retarded model.
And this is the magic of this progress, is you think you've reached something... I remember thinking at the moment, "If this is the last model we get, I'll be set. I'll be happy." I could live with Opus 4.5 for the next 20 years-... and you would hear no complaints from me because it was just so incredibly amazing to see an agent do all this work in the way that I wanted it done. Because it was not just about it being able to solve a task, it was also that I could look at its path there and go, "Yep. Yep. Maybe not there, but almost," and it would give you two notes. You'll get to where I wanted to go in the way I wanted to go there. You could produce code I wanted to merge. You could produce code that actually looked beautiful if it was written in Ruby. Rust, different question. But this ability for the agents to truly become an extension of how I wanted to work was very novel. But then we wait just until early spring, and suddenly we get sub-agents. We get harnesses that can subdivide the task, and something that would take Opus quite a while suddenly took a fifth of the time, a tenth of the time, because it could get chopped up and suddenly you got eight sub-agents working for you. But both of those two first phases of the agentic age, to me, still felt like I had to drive. I could tell it what I wanted, where to look for it, and steer it a little bit when it went off, and then we'll get there, and I would go much faster. But I had to be in the driver's seat.
I had to tell it what I wanted, and I had to be the reviewer, the auditor of what was coming out. And then finally now, this summer, with Opus 5, Fable, and Sol, GPT Sol, and to a lesser extent, some of the open weight models, we've arrived at a new era where I'm not telling it where we're going. I'm telling it the problem I have. I'm telling it the fuzzy, vague idea I have. It tells me where we're going. It tells me which path to take, and I will still look at it because I'm a curious person and I like computers and I like the outcome of it, but I really kinda don't have to. I've become optional in the part that produces the code that picks the route. Remember when early GPS systems came out? They were amazing compared to looking at a map, but you still wanted to pay attention. Is it gonna drive you in the harbor? I remember these newspaper articles. "Ah, GPS, they're terrible because people don't pay attention, they drive in the harbor." When was the last time GPS drove anyone in the harbor? Like, that just doesn't happen anymore. In fact, the cars now just drive themselves, right? And this is where we've arrived at now, that I can trust. For the domain I'm working in right now, I can trust it, and feel completely confident that it's gonna have my back. It's not gonna do something stupid, and if it does something stupid, it's gonna be able to recover.
`0:14:12`
Yes, but that said, AI is insanely capable at both finding-... and fixing security vulnerabilities. This was the whole blowup about Fable. This model was so capable of finding holes that a an attacker could exploit that it was simply not safe to release. So the irony here is that when you look at that field, it seems like we've reached levels of intelligence that virtually no human can match because many of these security holes are about stringing combo moves together. You find one little vulnerability here that by itself might not be the worst thing in the world, but then you combine it with four others, and suddenly you have RCE, remote command execution. Humans who are able to do that are very rare. They usually work inside state-sponsored organizations or other clandestine operations. They're not just out and about finding things. So they've gotten so good at that. The Linux distribution is why I've gotten 100% AI pilled. Because I've been working on Omarchy for the last three months, this version that just dropped a few days ago called Quatro. And almost right from the beginning I'm working on that version, the agent acceleration neared 100%, and in the last two months it has been 100%. I have not written-
`0:15:43`
... any of the code that's shipped in Quatro by hand. I've reviewed the shape of all of it. I've reviewed the individual lines of anything that's critical in the model layer of the system, and I've not looked at a bunch of the UI code. I have not looked at a bunch of the auxiliary code, and I've not written any of the new functionality entirely by hand. But then the web part, actually evolving Basecamp and Hey, our professional products that have lots of users and are relatively large code bases, have proven surprisingly tricky to fully accelerate with agents. We just released Basecamp 5 not too long ago. That was the first product at 37signals that was really agent accelerated. Because we were in this final sprint phase from around February. By then, agents were already good. And we had this early surge of, "It's solved. We can just have the designers do the programming. They know what features they want. They know what shape they want it to take. Let's just-- Let them vibe." And we let them vibe. And we ended up with a lot of PRs that individually perhaps could have been justified for a hot moment, taken all together, destroyed the architecture of the system. And we actually had to clean up manually, mop it up by hand, by human hand, to get back to an architecture that felt cohesive and coherent. So we still have a bit of that-- that was February, by the way. Things are quite different now.