Superintelligence CouncilСовет гения · sim.im

Anton Gladkov · 2026-08-09 · source ↗ · whole document (21)

Market Scan Model Benchmark - Human Scorecard Generated: 2026-08-08T22:35:20.254

## Bottom Line

- Keep Qwen3.8 max as the production-safe baseline until a separate code change implements stage-specific model routing. The current production path is intentionally fail-closed there.
- Use GPT 5.6 Sol Max for research expansion and maximum sourcing-map breadth. It found the widest segment/arms universes on every client, but it is slower and can over-produce arms that need pruning/measurement.
- If Qwen is unavailable and we need to write letters, use GPT 5.5 xhigh first. It completed every cell fastest, kept copy controlled, and produced balanced maps without the Sol-scale explosion.
- Use Kimi k3 max as the second writer / voice challenger when capacity is healthy. It produced strong copy and wide maps, but the first Revopush attempt hit a Kimi profile quota/auth limit before the successful retry.
- Do not make GLM5.2 via Qwen CLI the first fallback. It completed, but the route/attestation is less clean and the first Connectro attempt hit Qwen API data inspection before a long successful retry.

## Recommended Model Routing

| Market Scan phase | Primary | Fallback | Why |
|---|---|---|---|
| Research map / adjacent discovery | GPT 5.6 Sol Max | GPT 5.5 xhigh | Sol produced the broadest segment maps on all three clients; GPT 5.5 is the faster, more controlled fallback. |
| Sourcing arms before provider measurement | GPT 5.6 Sol Max | Kimi k3 max if capacity healthy; Qwen3.8 max for conservative maps | Sol gives maximal coverage; Kimi is strong on arms/voice; Qwen is narrower but production-shaped. All arms still need measurement/pruning before purchase. |
| Copy packs / letters | Qwen3.8 max in current production path; GPT 5.5 xhigh if Qwen is unavailable | Kimi k3 max | GPT 5.5 was the best operational writer fallback: fast, complete, controlled. Kimi is a strong alternate voice but capacity-sensitive. |
| Whole-run emergency fallback | GPT 5.5 xhigh | GPT 5.6 Sol Max for quality-max, Kimi for second pass | GPT 5.5 had the cleanest end-to-end operational profile. Sol is better for breadth, worse for latency/overproduction. |
| Judge / acceptance | Not evaluated here | Existing Kimi judge remains current contract | This benchmark evaluated writer cells only. |

## Aggregate Matrix

| Model | Completed | Avg min | Segments | Arms | Packs | Touches | Arms / segment | Note |
|---|---:|---:|---:|---:|---:|---:|---:|---|
| Qwen3.8 max | 3/3 | 41.1 | 50 | 575 | 50 | 250 | 11.5 | Production-safe baseline; conservative, doctrine-compatible, not the broadest discovery engine. |
| GPT 5.5 xhigh | 3/3 | 16.4 | 48 | 661 | 48 | 240 | 13.8 | Best first fallback when Qwen is unavailable: fastest complete cells, balanced maps, controlled copy. |
| GPT 5.6 Sol Max | 3/3 | 41.8 | 73 | 4407 | 73 | 365 | 60.4 | Best research and breadth generator; use before pruning/measurement, not as default final writer without review. |
| Kimi k3 max | 3/3 | 29 | 39 | 1275 | 39 | 195 | 32.7 | Strong voice and sourcing challenger; use as second writer/reviewer when capacity profile is healthy. |
| GLM5.2 max via Qwen CLI | 3/3 | 31.1 | 39 | 236 | 39 | 195 | 6.1 | Useful compact independent check, but not first fallback due route attestation, provider-screen, and latency caveats. |

## Cell Matrix

| Client | Model | Status | Min | Segments | Arms | Packs | Touches | Self-fit research/sourcing/copy |
|---|---|---:|---:|---:|---:|---:|---:|---|
| connectro | Qwen3.8 max | completed | 42.1 | 22 | 175 | 22 | 110 | good/acceptable/good |
| connectro | GPT 5.5 xhigh | completed | 16.7 | 27 | 463 | 27 | 135 | good/acceptable/acceptable |
| connectro | GPT 5.6 Sol Max | completed | 36.8 | 29 | 2301 | 29 | 145 | good/good/good |
| connectro | Kimi k3 max | completed | 39.3 | 21 | 886 | 21 | 105 | good/good/good |
| connectro | GLM5.2 max via Qwen CLI | completed | 27.4 | 18 | 114 | 18 | 90 | good/acceptable/good |
| masha | Qwen3.8 max | completed | 38.7 | 11 | 82 | 11 | 55 | good/acceptable/good |
| masha | GPT 5.5 xhigh | completed | 15.2 | 9 | 54 | 9 | 45 | good/acceptable/good |
| masha | GPT 5.6 Sol Max | completed | 40.2 | 16 | 757 | 16 | 80 | good/acceptable/good |
| masha | Kimi k3 max | completed | 36.7 | 7 | 114 | 7 | 35 | good/acceptable/good |
| masha | GLM5.2 max via Qwen CLI | completed | 32.3 | 8 | 59 | 8 | 40 | good/acceptable/good |
| revopush | Qwen3.8 max | completed | 42.5 | 17 | 318 | 17 | 85 | good/acceptable/good |
| revopush | GPT 5.5 xhigh | completed | 17.4 | 12 | 144 | 12 | 60 | good/acceptable/good |
| revopush | GPT 5.6 Sol Max | completed | 48.4 | 28 | 1349 | 28 | 140 | good/good/good |
| revopush | Kimi k3 max | completed | 10.9 | 11 | 275 | 11 | 55 | good/acceptable/good |
| revopush | GLM5.2 max via Qwen CLI | completed | 33.7 | 13 | 63 | 13 | 65 | good/acceptable/good |