DHH · 2026-08-26 · source ↗ · whole document (17)
Best AI coding models
`2:33:21`
I did it on all the models. First, I did it... Actually, funny thing. So I ran out of Fable tokens about two-thirds through, and it just automatically switched over to Opus 5 and kept going and finished the job. Part of the reason why I think it was able to do that was that the first thing Fable did was create a plan, and it was a really detailed plan. I think it had eight separate steps of, "Here you do it, and here, how you analyze it, and here you run the effects," and so on. I didn't review the plan at all. I didn't change the plan. It just made the plan so Opus could take over. And then I thought, "Well, if Opus can finish the job, maybe some of the other agents could finish it, too." So the first thing I did was I gave it to... I think I gave it to Sol. And Sol finished the job, too. It took twice as long, so it took, I think, about an hour and a half. But here's the kicker. The per token cost, I didn't pay per token. I have a max subscription to Claude, right? So it did it within the subscription. It used all my tokens, so I had to switch over to Opus 5. But if I had paid per token, it would've been 550 bucks, I think, to do the whole thing. And I thought for a second, "Holy shit. What a steal." If I personally had to learn Rust well enough to be able to do this translation, I'm looking at a nine-month job here. I can pay 500 bucks to have this translation happen, and suddenly I get a 10x execution speed up. This is ama- I would totally pay 500 bucks for this.
But competition. So I give it to Sol. Same plan. To be fair, I didn't ask Sol to do the plan. I just took the Fable plan, gave it to Sol. Sol, in an hour and a half, and I think $46 worth of token, repeated the task. Did the same thing. And I thought, "Well, blimey, that's amazing." Then I got greedy. So I asked GPT Luna, which is this crazy cheap model that OpenAI has as well, "Can you do it?" Absolutely not. First of all, it didn't even wanna start the task. I think something happened I don't know. I think it was in the spring, where we didn't need these slash goal things anymore. The agents could just automatically keep going in a loop if you told it not to stop-... and so forth. So Sol could do that. Fable could do that. But Luna couldn't. So I had... I think I did 12 prompts. Kept telling it to do it, and eventually I sorta got it started. The first thing it did was to cheat. So the first thing it, it looked outside its own directory, saw that there was already another imple- implementation, and just did a short wrapper around that and said, "I'm done." Hilarious. But it couldn't finish because it made just a couple thing. But I mean, okay, so it can't do that. Then I gave it to Grok-... 46. And I had used Grok 45 a little bit, and I thought like, "Ah, I mean, it's cool that there's others trying, but, like, I'm not gonna use it," because it felt quite far behind. Grok 46 fucking completes the task. 10x speed up, same size executable.
$55 Worth of per token cost, I think it was. Absolutely unbelievable. Then I repeated, too, with, uh, Kimi K3-... which took forever. I forget how long Kimi actually spent on it. And then I also did it with DeepSeek V4 Flash first. And Flash failed the same way that Luna did. It couldn't do it. And then I did it with Pro, and Pro also completed the task. It took 2 hours 45, $23. So here we are, right? Like, Fable, clearly the best. It was the fastest. It was the one that wrote the plan, but 550 bucks, and the output the same. Uh, the others, Sol, Grok, about the same 1/10 the cost. DeepSeek, 1/20 the cost, but you have to wait a little longer. Absolutely gobsmackingly incredible. And now, by the way, by the way, so I had Fable finish the first job, and then Opus finished it. That was the first one shot, right? It's 10 times faster. I did two auto research runs, which isn't even auto research anymore. You don't have to do the slash. You just tell it to keep going until you tell it to stop. It ended up... I think we ended up with a 46 time execution improvement over the original.