Superintelligence CouncilСовет гения · sim.im

DHH · 2026-08-26 · transcript · source ↗

Best AI coding harnesses

`2:38:14`

What I found, it's really interesting because Fable is, in my opinion, the best model right now, but it also makes mistakes. And the best way to get the best software, I would actually rather have two differently sourced sort of... I mean, they're not mid-tier, they're all frontier. But have, let's say Opus 5 and Codex-... and have one check the other's job. This is my standard operating procedure now. I'll have Opus or Fable do the work, and then I always end it, review with Codex xHigh. And I've also started using Grok just to test it out, and it's also quite good. And it keeps finding stuff. And then that's my workflow when I'm having my agents on my own machine do it, and then I push to GitHub. And then Copilot, I kid you not, has actually gotten good. Copilot keeps finding stuff that's legitimately broken, which is also incredible acceleration because the first version of Copilot that started doing this was literally retarded. It would just constantly flag things that were nonsense. It would constantly flag the same problem over and over again as you would push. It was really annoying, so I think a lot of people actually ended up turning that off. And if they did, they should turn it back on because it's actually quite good. Keeps finding things. And if you then take that, and we shouldn't be surprised. Why are we surprised? Even if you're a good programmer, if you finish a job and you ask your also very good peer to review it, you're gonna end up with better code.

Of course you're gonna end up with better code. So build that into your process. Pick one of the agents to drive with. I've mainly been driving with Claude, which is actually interesting because I have some other reservations about Anthropic. But the reason I'm sticking with Claude is, in my opinion, they actually have the best harness. And one of the reasons it's the best harness is it's multi-agent running. So if you wanna run multiple agents at the same time, you can do arrow left when you're inside a session, then it goes back to Agent View. And here in Agent View, you can pick up another agent. So if you wanna do this thing where you have multiple threads going on-... the Claude Code is just the nicest setup. They also just, they keep being a little further ahead, which probably shouldn't be surprising. I mean, Boris is one of the guys that's working on that, he was the first one and basically came up with the thing. It's just interesting that that has been an enduring advantage. I've also used OpenCode a lot. I think OpenCode is great. I use OpenCode mainly as my main harness for all the open models, so Kimi K3 and so forth. And I inference not on the Chinese servers. So I inference on Fireworks-... which is a really nice service where you can just pay by the token, and if you're using these open weight models, they're not that expensive by token. So that's a great way to do that. But I don't think I would be using... Well, maybe I would. But if...

Using Claude with a subscription does feel like a bargain, like a crazy bargain.

`2:41:06`

I just signed up for my second one this morning-... because I got up at 4:00 with jet lag, and I started working with the agents right away, and now I'm three days away from limits resetting on Fable, and I ran out of tokens.

`2:41:22`

I mean, I don't... Why is that so complicated? Can't you just stack one subscription? Why do I need to log in multiple times?

`2:41:33`

Oh, we're building that into Omarchy, by the way. So the next version of Omarchy is gonna ship with multi-sub support. I wish that the labs would just make you buy a max 100 times.

`2:41:57`

Yes. That's actually the deal. The deal with Grok is that their fast mode is cheaper than the regular mode on the others. And when you've used Grok 46 on fast, it's kind of addictive 'cause it almost gets you to this single thread mode-... where the agent is able to keep up with you.

`2:42:34`

By the way, one of the reasons why I'm not sure what the final form of these harnesses is gonna take, because one of the things we've started experimenting with at Base Camp is that we put the agents inside of Base Camp, and we treat them as coworkers.

`2:42:47`

... inside of Base Camp, so they will do work on a to-do in Base Camp, or they'll work, do work on a card. And it turns out that a collaboration tool that's optimized for asynchronous communication is actually the right format versus these harnesses are a little more like chat. And chat, which was, I mean, the original format for these things, is not the right thing because you're sitting around waiting. It entices you to sit around and wait, versus when you're dealing with something like Base Camp and you can just give an agent a to-do item, you don't expect to hear back immediately. So I don't know what the final form is gonna take, but I've been using all of the harnesses. So I use them inside of Herdr, and Herdr has panes. So oftentimes I'll run Claude up top, and then I'll run Codex down below, and then maybe I also have an OpenCode set up. But I will say just very recently, I found that the human in the loop is the limit here. So I've started setting up more automated systems. I've been building an Omarchy bot that can do more autonomous development on Omarchy that is just On a regular schedule process, all depending to do or PRs and issues, and then send me an email using the Hey CLI, and then I just get this email. Here's 12 PRs that are either ready to go or I think you should close. And then I just make the final determination there.

`2:44:10`

Yes, but in an exhilarating way. Like, one of the things I always loved about race cars was when I would stumble out of the car absolutely smashed and barely able to hold my head up, and I'd lay down on the garage floor and just think, "Holy fuck, I'm alive." That's the kind of exhaustion, mental exhaustion I'm feeling at moments with the agents right now.

`2:44:34`

It's physical. It's just... I mean, I bet it's the same thing with jujitsu, right? Like, it's actually satisfying to feel exhausted. It feels like you've applied yourself. And I feel that with the mental exhaustion you can get from agents.

`2:44:49`

Yes. If you drive in difficult conditions, especially if you drive in the rain, where you're constantly managing the thing right at the knife's edge. But most of the time, I just enjoyed the physical exhaustion.

`2:45:25`

It's funny. It's actually very similar with racing. There's some tracks, like Le Mans, where you get time to relax. There's the Mulsanne straights where it's kinda long. You're going fast, but in a straight line, you can take a breath. And then there are other tracks where it's just coming full on all the time, and you don't have any moment to relax, and you stumble out of the car after an hour in each, and it's very different how exhausted you are. And this is exactly what I'm finding with the agents too, that when you're running at max human capacity-... you're just constantly in a corner.

`2:46:01`

I think the moment we're in right now is gonna pass. That this need to... It's not a babysitting motion, but constant interaction with the agents is gonna fade. From what I've seen on Omarchy, I think we can automate a ton of the development and debugging of the system-... to the point where I can just review that email once a day, and then make the decisions once a day. Yep, goes, no, goes, in, out, and whatever. We're not there yet, but I think we're gonna get there.

`2:47:14`

It's the birth of a new paradigm. It's always messy.

`2:47:17`

It's always exhausting. There's no other way around that. And by the way, hasn't San Francisco always been like this? I remember the dotcom boom years, and everyone was just the same, and then it was all mobile, and before that it was the gold rush. So I think that's probably-... just San Francisco. It attracts-... the kind of individuals-

`2:47:34`

... who would, who would think, "I don't have time to sleep." I've not been on that track, generally speaking, but I will admit that the last three months have felt more exhausting than any other project I've done in the last five years, since the Hey launch. That was the last time we had a crazy exhausting launch.

`2:47:57`

No, no, no, no, no. This is not sustainable at all.

`2:48:00`

But I also, I can see the end of the tunnel. Because the automation, which by the way, everyone is building. This is so hilarious. Like, everyone is building their little gas town, their little agent coordination, their setup. This is all gonna be solved. We're not all gonna have to build our own coordination harnesses. Of course we're not. And in fact, I'm a little surprised that it's gone this long, that there's not more of it has been sucked up by the major labs. You'd think that they'd just built this stuff in. But some of it is also this is what a new domain looks like. I remember for a hot moment when JavaScript kind of came to realize its own power, and there was basically a new JavaScript framework every five minutes, right? Like, a lot of churn. This is what happens at the advent of a new paradigm. So I think it's natural, and I think it's gonna settle down, and I can already see the light at the end of the tunnel. So I don't mind a sprint. In fact, I welcome it now. I like this notion that most time is calm, but then occasionally you gotta climb a mountain.

`2:49:04`

What a-... thriller that this has been. I mean, if you would've scripted this, I'd be like, "This is so far-fetched. Get out of here. Unbelievable," right? I wouldn't wanna know anything. The fact that this rollercoaster and this acceleration has been absorbed in real time by everyone, no one knew, right? Like, some had premonitions that were a little better than others, but no one knew, not exactly the way-... it was gonna go and how fast it was gonna take off. I mean, it's the show of a lifetime. I mean, the mean cinema-... absolutely applies-

`2:49:52`

That all our dreams are coming true. Like, this is the AI maximalist abundance argument. That once Elon's robots have the level of intelligence and AGI that I'm experiencing on some of this development stuff, the world is gonna be so unrecognizably different that we can't even imagine it. Now, I don't spend any time thinking about that because I think that is the way to the AI psychosis. So therefore, I just immerse myself in the moment and get absolutely the maximum out of it and-

`2:50:30`

Just have fun. I mean, I get it why people are finding this exhausting. Just keeping up. Not even using agents, just, like, being like, "What's the latest model? What did, what tool are we using now? Is it TMox? Is it a Herder? Is it none of it? Is it agents?" Whatever. I get it, but I also again think we should be so blessed. You are alive in this moment where decades are happening-... in weeks.