Superintelligence CouncilСовет гения · sim.im

DHH · 2026-08-26 · source ↗ · whole document (8)

Programming with AI agents

`0:05:54`

To me, that's why even when I look back upon our conversation a year ago, I don't actually have different opinions. I have the same opinions. A year ago, I did not like the mode of AI we were offered. It was the autocomplete mode, or it was the AI chatbot mode. Now, the chatbot I actually liked, as we talked about. Great tutor right from the get-go, great way of looking things up on the internet, not what was gonna replace me chiseling code. But-... then we get the agents. And the agents start out being curiosities for about five minutes, and then they get amazing, and then they get, "Oh my God, is this AGI?" And all of that happened since last year, even just within the last nine months. We have basically these few phases here. We have AI in the pre-agentic era. I was excited about that, but it was not fundamentally rewriting the rules of the game for me. It was not completely changing how I worked. I was still chiseling code, I just had a little helper, a little sidekick-... who could bounce ideas off, and I could look up this information online and so forth in a more efficient way. It was just a more efficient way to do what I was already doing, and it didn't change the emotional connection I had to the computer. Then we get to November 24th, 2025. Opus 4.5, to me, was the dividing line, where suddenly... I didn't even try it on the 24th. I think I tried it on the 26th.

I give it a couple of tasks, and I realize that the quality of the output is uncannily close to what I would've written. And I remember just leaning back and thinking, "What just happened?" "How did we go from this autocomplete mess that I was talking to you about in the summer to this just a few short months later? How did we get both the increase in intelligence and then also the increase in usability?" This agent harness question where I don't know if Opus 4.5 was that much smarter than Opus 4, which is what we had in the summer, but its ability to instrument your computer, to use tools, to check its own work, to apply its intelligence in such a way that you could get real meaningful work out of it, was completely different. And I think this is then the big change that happens for almost anyone who paid attention and started playing with it over the Christmas break. This is what I heard-... from Shopify and other places with lots of employees who suddenly had a breather, suddenly had a couple of weeks to lean back and just look at what was going on, gave it a try, and had the same mind-blowing experience that these agents were of a different genre than what we had before. And then what we get to is, I'm already excited at this point. Like, by December, I'm already revisiting sort of all my priors-... and going like, "Wow, if it can do this, can it also do that?" "Oh, yeah, it can." And again, Opus 4.5 now looks like a retarded model.

And this is the magic of this progress, is you think you've reached something... I remember thinking at the moment, "If this is the last model we get, I'll be set. I'll be happy." I could live with Opus 4.5 for the next 20 years-... and you would hear no complaints from me because it was just so incredibly amazing to see an agent do all this work in the way that I wanted it done. Because it was not just about it being able to solve a task, it was also that I could look at its path there and go, "Yep. Yep. Maybe not there, but almost," and it would give you two notes. You'll get to where I wanted to go in the way I wanted to go there. You could produce code I wanted to merge. You could produce code that actually looked beautiful if it was written in Ruby. Rust, different question. But this ability for the agents to truly become an extension of how I wanted to work was very novel. But then we wait just until early spring, and suddenly we get sub-agents. We get harnesses that can subdivide the task, and something that would take Opus quite a while suddenly took a fifth of the time, a tenth of the time, because it could get chopped up and suddenly you got eight sub-agents working for you. But both of those two first phases of the agentic age, to me, still felt like I had to drive. I could tell it what I wanted, where to look for it, and steer it a little bit when it went off, and then we'll get there, and I would go much faster. But I had to be in the driver's seat.

I had to tell it what I wanted, and I had to be the reviewer, the auditor of what was coming out. And then finally now, this summer, with Opus 5, Fable, and Sol, GPT Sol, and to a lesser extent, some of the open weight models, we've arrived at a new era where I'm not telling it where we're going. I'm telling it the problem I have. I'm telling it the fuzzy, vague idea I have. It tells me where we're going. It tells me which path to take, and I will still look at it because I'm a curious person and I like computers and I like the outcome of it, but I really kinda don't have to. I've become optional in the part that produces the code that picks the route. Remember when early GPS systems came out? They were amazing compared to looking at a map, but you still wanted to pay attention. Is it gonna drive you in the harbor? I remember these newspaper articles. "Ah, GPS, they're terrible because people don't pay attention, they drive in the harbor." When was the last time GPS drove anyone in the harbor? Like, that just doesn't happen anymore. In fact, the cars now just drive themselves, right? And this is where we've arrived at now, that I can trust. For the domain I'm working in right now, I can trust it, and feel completely confident that it's gonna have my back. It's not gonna do something stupid, and if it does something stupid, it's gonna be able to recover.

`0:14:12`

Yes, but that said, AI is insanely capable at both finding-... and fixing security vulnerabilities. This was the whole blowup about Fable. This model was so capable of finding holes that a an attacker could exploit that it was simply not safe to release. So the irony here is that when you look at that field, it seems like we've reached levels of intelligence that virtually no human can match because many of these security holes are about stringing combo moves together. You find one little vulnerability here that by itself might not be the worst thing in the world, but then you combine it with four others, and suddenly you have RCE, remote command execution. Humans who are able to do that are very rare. They usually work inside state-sponsored organizations or other clandestine operations. They're not just out and about finding things. So they've gotten so good at that. The Linux distribution is why I've gotten 100% AI pilled. Because I've been working on Omarchy for the last three months, this version that just dropped a few days ago called Quatro. And almost right from the beginning I'm working on that version, the agent acceleration neared 100%, and in the last two months it has been 100%. I have not written-