DHH · 2026-08-26 · transcript · source ↗
Best AI coding models
`2:21:11`
Don't let the perfect be the enemy of good. I mean, if you're suddenly enabling something people wouldn't be doing before at all, it's okay to take five seconds longer. I think this is one of the areas that Steve Jobs was so good, that he realized, like, the first iPhone was terrible in all sorts of ways, right? Absolutely horrendously slow data connection, very underpowered, all the things. But the product was so compelling that it just didn't matter. You'd be shipping too late if you waited until everything was just right, just perfect. If you're solving a problem that people currently don't have a solution for, they're willing to give it five seconds. If you're in a competitive market, yeah, okay, it's different. But that means the problem's already solved, and you don't need to solve it.
`2:22:02`
Yeah. Yeah, we're using Voxtype. But others have-
`2:22:07`
It is. It's just using one of the open models. I forget which parrot model it's using or something like that. It works quite well for short commands, so I have used it for that. Uh, I think you just hold down F9-... and it starts a dictation.
`2:22:23`
You just have to get online. It'll offer you to install-
`2:22:26`
Because the Voxtype package includes a model that's 150 megabytes, and I was like, "Ah."
`2:22:35`
... we don't have that in there by default, but we have the option set up. And I actually just yesterday was thinking we really need to combine all of these things. So Omarchy Quattro is truly amazing as a malleable operating system where you can add functionality to it just by talking to your agent. But that's exactly what we should be doing. There's a lot of people who would find it very natural just to talk to their computer and say, "Hey, can you make me a stock panel? I wanna track Apple and Dell," and just see the agent go off through, entirely through voice, both in terms of the input and in terms of the output, and then your operating system just changes. This was the vision of Iron Man and Jarvis and-
`2:23:23`
... bespoke applications that just appear magically. I'm thinking the area of AI I find truly intriguing are these live gaming models, I don't know if you've seen that, where the AI's inventing literally the next frame, but you can actually play the game. And I think if they're able to do that, why can't we do that with our operating system? Why can't we just talk to it and ask it to be a different way or change, and it'll all just change? I think this, by the way, is one of the other breakthroughs that is going to lead to the total domination of Linux. The agents have taken all the hardship out of diagnosing Linux systems and turned the fact that Linux produces these overly specific, totally arcane error messages into its greatest advantage. When Linux has an issue, the agent can take that very specific error message that makes no sense to a normal human and correlate it with the fact that the agent was pre-trained on 40 million lines of Linux code, so it knows exactly where to look and dial it down.
`2:24:28`
I have not had a single problem on my Linux machine since the beginning of this year that an agent could not diagnose. That was not true a year and a half ago. A year and a half ago, I was still searching on forums-... to find answers to esoteric questions. Now the agents know the source code of not just the Linux operating system, but every single piece of software I have on that box. The agent has access to the source code of all of it. In fact, this was one thing I built in just before we, uh, we shipped. So once you set up your default agent-... Omarchy Quattro has a crash watcher. If any app on your machine crashes, it'll pop up a little thing, ask you whether you want your AI to diagnose the problem.
`2:25:14`
Unbelievable. I've seen things where there's an error in some subsystem. Some, some... Not even Linux, just some app I installed. The agents start digging through the logs and looks up systemd log. Then it checks out the damn source code. Of the application that crashed, pins down that it's in this Rust file line 472 that there's an unbounded, unwrapped variable that overflowed or whatever it is. Then offers you whether you wanna file a bug report with all this detail. I'll give you one amazing anecdote. So I was working with this guy JDX, who's working on mise. This, this package manager for fast-moving development tools. This is how we manage all the agent software and so forth. Normal package managers are not built to be updated seven times a day. And all of these agent harnesses are updated about seven times a day, so we needed an out-of-band package manager, and mise just turns out to be perfect for this. Anyway, it had a small bug. There was some race condition which agents are very good at finding because when you start running multiple agents at the same time, they will suss out all these race conditions you have in your underlying infrastructure that was never triggered by a human trying to do manual thing. So it finds this issue, right? I tell I tell the agent, "Hey, can you post this as a bug report to JDX's GitHub?" And unfortunately, right before that, I had had to do a QA run on Omarchy itself with eight different agents. It had found 28 real issues-...
that it needed to file. Well, it went to GitHub and tried to file all 28 issues at the same time, which it did in about, uh, I don't know, 12 seconds. GitHub, not unreasonably, marked that as probable spam and banned-... my Omarchy bot. And then I couldn't do that. So I was blocked from the bot having access to GitHub, so I just told it, "Hey, do you know what? Just email JDX. Here's his email address." I had already set it up with hey.com-... an email address. We have a CLI that's in beta right now. So I set it up with that so it can send to email. That's how it gives me reports about outstanding issues and PRs. So it sends, um, JDX this email about the bug it found, and he was like, "Wait, this is the first time I've gotten a bug report in unreleased software." Because it had downloaded the source code to mise, saw that he had already fixed the bug, but that it would not fix the problem entirely in software they had not shipped yet. Pinpointed the problem, and he was like, "Damn it-... I got a bug report before we even cut a release."
`2:27:54`
... believable AGI levels of mind-blowing stuff.
`2:28:10`
They are so incredibly good at this, and this is actually interesting because I was on the perhaps same side of that, a little skeptical about can it reason about all these things. And Mikhail, who's the CTO at Shopify, ran a scientific study on this. They had... And this was, I think, late last year or early this year, where he had agents go back through all the incidents, both outages and, and other problems that Shopify had had in production, trace that back to the PR that was merged, and find out whether PRs that had been reviewed by human or PRs that had been reviewed by agents were of higher quality. Well, lo and surprise-... the PRs that have been reviewed by agents caused far fewer issues in production. And this was with models we had six months ago. At this point, 100% in the majority of domains we work in today, agents are better at finding bugs.
`2:29:23`
What's amazing to me is that you could rattle off so many different contenders. That this market is so wide open, that it does actually change back and forth, that we have real competition, that there are so many labs that are able to get either to the frontier or close to it. That, by the way, is remarkable. I still don't fully understand that. But to answer your question, the best model in general right now is Fable. The second-best model, in my opinion, is Opus 5. But the tier just below Opus 5, and it's not even that they're always below, sometimes they're ahead. I'll get to that in a second. I would rank GPT Sol very good. Grok 4.6 I just started testing a few days ago. I had this wonderful test that I've set up by accident where I translated this Python library into Rust. Uh, the screensaver you just saw with all the cool animation? That's powered by a Python library called Terminal Text Effects. Really cool library. We've been using it since the first day of Omarchy. The problem with that is it's written in Python, so when it starts up, especially on a laptop, and it runs in Python, it uses all of your CPU to do these effects, and therefore, it means it uses about 30 watts of energy, and it spins up your fans, and it drains your battery. Doesn't really matter on a local computer, but it does matter on a laptop. So I thought, "Do you know what?
This sounds like a problem for Rust." So first I gave Fable the challenge, and all I told it was, "Here's the source code for the Python library," this TTE library that had a bunch of dependencies and so forth. "I want a Rust version of this with no dependencies, a single executable." Like, that's what Rust does. So basically, I want it in Rust. I want it to be pixel perfect, frame by frame, do a full analysis, don't stop until you're finished.
`2:31:29`
I kid you not, in just under 45 minutes, it was like Am I heard all? I'm finished. I've checked everything. I have reduced the startup time from 86 milliseconds to two milliseconds. I have sped up the execution time by 9.6 times, I believe it was. The executable is three megabytes. Do you wanna run it?
`2:32:02`
... and I don't know why I'm surprised, because this translation job is something we've known for a while that AI is pretty good at, but it was still staggering to me that I could one-shot a full translation of a Python library I had been using for a year, that others had been using for much longer, and turn it into a Rust executable without knowing any Rust, without looking at the Rust code at all, and produce this executable that I then told the agent right after, "This is great. Ship it." It packaged it up as a new package. It told me, "What do you want it to call?" "Uh, let's call it TTFX. Let's create a new Git repo." It sets up the Git repo. "Let's create a new package, build package for our build system." It puts that up. "Let's push it out. Let's open the pull request to the Omarchy itself so we switch around from TTE, the Python implementation, to TTFX." It does all of it, and I'm just sitting there. And again, I'm already at this point fully delirious with agent acceleration, and I still had to lean back and I go like, "This is AGI, isn't it? This is what AGI looks like. If we have... If these moments are just what it is all the time, this is AGI."
`2:33:21`
I did it on all the models. First, I did it... Actually, funny thing. So I ran out of Fable tokens about two-thirds through, and it just automatically switched over to Opus 5 and kept going and finished the job. Part of the reason why I think it was able to do that was that the first thing Fable did was create a plan, and it was a really detailed plan. I think it had eight separate steps of, "Here you do it, and here, how you analyze it, and here you run the effects," and so on. I didn't review the plan at all. I didn't change the plan. It just made the plan so Opus could take over. And then I thought, "Well, if Opus can finish the job, maybe some of the other agents could finish it, too." So the first thing I did was I gave it to... I think I gave it to Sol. And Sol finished the job, too. It took twice as long, so it took, I think, about an hour and a half. But here's the kicker. The per token cost, I didn't pay per token. I have a max subscription to Claude, right? So it did it within the subscription. It used all my tokens, so I had to switch over to Opus 5. But if I had paid per token, it would've been 550 bucks, I think, to do the whole thing. And I thought for a second, "Holy shit. What a steal." If I personally had to learn Rust well enough to be able to do this translation, I'm looking at a nine-month job here. I can pay 500 bucks to have this translation happen, and suddenly I get a 10x execution speed up. This is ama- I would totally pay 500 bucks for this.
But competition. So I give it to Sol. Same plan. To be fair, I didn't ask Sol to do the plan. I just took the Fable plan, gave it to Sol. Sol, in an hour and a half, and I think $46 worth of token, repeated the task. Did the same thing. And I thought, "Well, blimey, that's amazing." Then I got greedy. So I asked GPT Luna, which is this crazy cheap model that OpenAI has as well, "Can you do it?" Absolutely not. First of all, it didn't even wanna start the task. I think something happened I don't know. I think it was in the spring, where we didn't need these slash goal things anymore. The agents could just automatically keep going in a loop if you told it not to stop-... and so forth. So Sol could do that. Fable could do that. But Luna couldn't. So I had... I think I did 12 prompts. Kept telling it to do it, and eventually I sorta got it started. The first thing it did was to cheat. So the first thing it, it looked outside its own directory, saw that there was already another imple- implementation, and just did a short wrapper around that and said, "I'm done." Hilarious. But it couldn't finish because it made just a couple thing. But I mean, okay, so it can't do that. Then I gave it to Grok-... 46. And I had used Grok 45 a little bit, and I thought like, "Ah, I mean, it's cool that there's others trying, but, like, I'm not gonna use it," because it felt quite far behind. Grok 46 fucking completes the task. 10x speed up, same size executable.
$55 Worth of per token cost, I think it was. Absolutely unbelievable. Then I repeated, too, with, uh, Kimi K3-... which took forever. I forget how long Kimi actually spent on it. And then I also did it with DeepSeek V4 Flash first. And Flash failed the same way that Luna did. It couldn't do it. And then I did it with Pro, and Pro also completed the task. It took 2 hours 45, $23. So here we are, right? Like, Fable, clearly the best. It was the fastest. It was the one that wrote the plan, but 550 bucks, and the output the same. Uh, the others, Sol, Grok, about the same 1/10 the cost. DeepSeek, 1/20 the cost, but you have to wait a little longer. Absolutely gobsmackingly incredible. And now, by the way, by the way, so I had Fable finish the first job, and then Opus finished it. That was the first one shot, right? It's 10 times faster. I did two auto research runs, which isn't even auto research anymore. You don't have to do the slash. You just tell it to keep going until you tell it to stop. It ended up... I think we ended up with a 46 time execution improvement over the original.