Superintelligence CouncilСовет гения · sim.im

DHH · 2026-08-26 · source ↗ · whole document (17)

Best AI coding models

`2:25:14`

Unbelievable. I've seen things where there's an error in some subsystem. Some, some... Not even Linux, just some app I installed. The agents start digging through the logs and looks up systemd log. Then it checks out the damn source code. Of the application that crashed, pins down that it's in this Rust file line 472 that there's an unbounded, unwrapped variable that overflowed or whatever it is. Then offers you whether you wanna file a bug report with all this detail. I'll give you one amazing anecdote. So I was working with this guy JDX, who's working on mise. This, this package manager for fast-moving development tools. This is how we manage all the agent software and so forth. Normal package managers are not built to be updated seven times a day. And all of these agent harnesses are updated about seven times a day, so we needed an out-of-band package manager, and mise just turns out to be perfect for this. Anyway, it had a small bug. There was some race condition which agents are very good at finding because when you start running multiple agents at the same time, they will suss out all these race conditions you have in your underlying infrastructure that was never triggered by a human trying to do manual thing. So it finds this issue, right? I tell I tell the agent, "Hey, can you post this as a bug report to JDX's GitHub?" And unfortunately, right before that, I had had to do a QA run on Omarchy itself with eight different agents. It had found 28 real issues-...

that it needed to file. Well, it went to GitHub and tried to file all 28 issues at the same time, which it did in about, uh, I don't know, 12 seconds. GitHub, not unreasonably, marked that as probable spam and banned-... my Omarchy bot. And then I couldn't do that. So I was blocked from the bot having access to GitHub, so I just told it, "Hey, do you know what? Just email JDX. Here's his email address." I had already set it up with hey.com-... an email address. We have a CLI that's in beta right now. So I set it up with that so it can send to email. That's how it gives me reports about outstanding issues and PRs. So it sends, um, JDX this email about the bug it found, and he was like, "Wait, this is the first time I've gotten a bug report in unreleased software." Because it had downloaded the source code to mise, saw that he had already fixed the bug, but that it would not fix the problem entirely in software they had not shipped yet. Pinpointed the problem, and he was like, "Damn it-... I got a bug report before we even cut a release."

`2:27:54`

... believable AGI levels of mind-blowing stuff.

`2:28:10`

They are so incredibly good at this, and this is actually interesting because I was on the perhaps same side of that, a little skeptical about can it reason about all these things. And Mikhail, who's the CTO at Shopify, ran a scientific study on this. They had... And this was, I think, late last year or early this year, where he had agents go back through all the incidents, both outages and, and other problems that Shopify had had in production, trace that back to the PR that was merged, and find out whether PRs that had been reviewed by human or PRs that had been reviewed by agents were of higher quality. Well, lo and surprise-... the PRs that have been reviewed by agents caused far fewer issues in production. And this was with models we had six months ago. At this point, 100% in the majority of domains we work in today, agents are better at finding bugs.

`2:29:23`

What's amazing to me is that you could rattle off so many different contenders. That this market is so wide open, that it does actually change back and forth, that we have real competition, that there are so many labs that are able to get either to the frontier or close to it. That, by the way, is remarkable. I still don't fully understand that. But to answer your question, the best model in general right now is Fable. The second-best model, in my opinion, is Opus 5. But the tier just below Opus 5, and it's not even that they're always below, sometimes they're ahead. I'll get to that in a second. I would rank GPT Sol very good. Grok 4.6 I just started testing a few days ago. I had this wonderful test that I've set up by accident where I translated this Python library into Rust. Uh, the screensaver you just saw with all the cool animation? That's powered by a Python library called Terminal Text Effects. Really cool library. We've been using it since the first day of Omarchy. The problem with that is it's written in Python, so when it starts up, especially on a laptop, and it runs in Python, it uses all of your CPU to do these effects, and therefore, it means it uses about 30 watts of energy, and it spins up your fans, and it drains your battery. Doesn't really matter on a local computer, but it does matter on a laptop. So I thought, "Do you know what?