Superintelligence CouncilСовет гения · sim.im

DHH · 2026-08-26 · source ↗ · whole document (19)

Best AI coding harnesses

`2:38:14`

What I found, it's really interesting because Fable is, in my opinion, the best model right now, but it also makes mistakes. And the best way to get the best software, I would actually rather have two differently sourced sort of... I mean, they're not mid-tier, they're all frontier. But have, let's say Opus 5 and Codex-... and have one check the other's job. This is my standard operating procedure now. I'll have Opus or Fable do the work, and then I always end it, review with Codex xHigh. And I've also started using Grok just to test it out, and it's also quite good. And it keeps finding stuff. And then that's my workflow when I'm having my agents on my own machine do it, and then I push to GitHub. And then Copilot, I kid you not, has actually gotten good. Copilot keeps finding stuff that's legitimately broken, which is also incredible acceleration because the first version of Copilot that started doing this was literally retarded. It would just constantly flag things that were nonsense. It would constantly flag the same problem over and over again as you would push. It was really annoying, so I think a lot of people actually ended up turning that off. And if they did, they should turn it back on because it's actually quite good. Keeps finding things. And if you then take that, and we shouldn't be surprised. Why are we surprised? Even if you're a good programmer, if you finish a job and you ask your also very good peer to review it, you're gonna end up with better code.

Of course you're gonna end up with better code. So build that into your process. Pick one of the agents to drive with. I've mainly been driving with Claude, which is actually interesting because I have some other reservations about Anthropic. But the reason I'm sticking with Claude is, in my opinion, they actually have the best harness. And one of the reasons it's the best harness is it's multi-agent running. So if you wanna run multiple agents at the same time, you can do arrow left when you're inside a session, then it goes back to Agent View. And here in Agent View, you can pick up another agent. So if you wanna do this thing where you have multiple threads going on-... the Claude Code is just the nicest setup. They also just, they keep being a little further ahead, which probably shouldn't be surprising. I mean, Boris is one of the guys that's working on that, he was the first one and basically came up with the thing. It's just interesting that that has been an enduring advantage. I've also used OpenCode a lot. I think OpenCode is great. I use OpenCode mainly as my main harness for all the open models, so Kimi K3 and so forth. And I inference not on the Chinese servers. So I inference on Fireworks-... which is a really nice service where you can just pay by the token, and if you're using these open weight models, they're not that expensive by token. So that's a great way to do that. But I don't think I would be using... Well, maybe I would. But if...