Superintelligence CouncilСовет гения · sim.im

DHH · 2026-08-26 · source ↗ · whole document (17)

Best AI coding models

that it needed to file. Well, it went to GitHub and tried to file all 28 issues at the same time, which it did in about, uh, I don't know, 12 seconds. GitHub, not unreasonably, marked that as probable spam and banned-... my Omarchy bot. And then I couldn't do that. So I was blocked from the bot having access to GitHub, so I just told it, "Hey, do you know what? Just email JDX. Here's his email address." I had already set it up with hey.com-... an email address. We have a CLI that's in beta right now. So I set it up with that so it can send to email. That's how it gives me reports about outstanding issues and PRs. So it sends, um, JDX this email about the bug it found, and he was like, "Wait, this is the first time I've gotten a bug report in unreleased software." Because it had downloaded the source code to mise, saw that he had already fixed the bug, but that it would not fix the problem entirely in software they had not shipped yet. Pinpointed the problem, and he was like, "Damn it-... I got a bug report before we even cut a release."

`2:27:54`

... believable AGI levels of mind-blowing stuff.

`2:28:10`

They are so incredibly good at this, and this is actually interesting because I was on the perhaps same side of that, a little skeptical about can it reason about all these things. And Mikhail, who's the CTO at Shopify, ran a scientific study on this. They had... And this was, I think, late last year or early this year, where he had agents go back through all the incidents, both outages and, and other problems that Shopify had had in production, trace that back to the PR that was merged, and find out whether PRs that had been reviewed by human or PRs that had been reviewed by agents were of higher quality. Well, lo and surprise-... the PRs that have been reviewed by agents caused far fewer issues in production. And this was with models we had six months ago. At this point, 100% in the majority of domains we work in today, agents are better at finding bugs.

`2:29:23`

What's amazing to me is that you could rattle off so many different contenders. That this market is so wide open, that it does actually change back and forth, that we have real competition, that there are so many labs that are able to get either to the frontier or close to it. That, by the way, is remarkable. I still don't fully understand that. But to answer your question, the best model in general right now is Fable. The second-best model, in my opinion, is Opus 5. But the tier just below Opus 5, and it's not even that they're always below, sometimes they're ahead. I'll get to that in a second. I would rank GPT Sol very good. Grok 4.6 I just started testing a few days ago. I had this wonderful test that I've set up by accident where I translated this Python library into Rust. Uh, the screensaver you just saw with all the cool animation? That's powered by a Python library called Terminal Text Effects. Really cool library. We've been using it since the first day of Omarchy. The problem with that is it's written in Python, so when it starts up, especially on a laptop, and it runs in Python, it uses all of your CPU to do these effects, and therefore, it means it uses about 30 watts of energy, and it spins up your fans, and it drains your battery. Doesn't really matter on a local computer, but it does matter on a laptop. So I thought, "Do you know what?

This sounds like a problem for Rust." So first I gave Fable the challenge, and all I told it was, "Here's the source code for the Python library," this TTE library that had a bunch of dependencies and so forth. "I want a Rust version of this with no dependencies, a single executable." Like, that's what Rust does. So basically, I want it in Rust. I want it to be pixel perfect, frame by frame, do a full analysis, don't stop until you're finished.

`2:31:29`

I kid you not, in just under 45 minutes, it was like Am I heard all? I'm finished. I've checked everything. I have reduced the startup time from 86 milliseconds to two milliseconds. I have sped up the execution time by 9.6 times, I believe it was. The executable is three megabytes. Do you wanna run it?