Superintelligence CouncilСовет гения · sim.im

Anton Gladkov · 2026-08-09 · source ↗ · whole document (21)

Market Scan Model Benchmark - Human Scorecard Generated: 2026-08-08T22:35:20.254

Market Scan Model Benchmark - Human Scorecard

Generated: 2026-08-08T22:35:20.254Z

Run directory: /Users/antongladkov/.codex/worktrees/SLSBMB-Sender/market-scan-model-benchmark/var/market-scan-model-benchmark/20260808T201848Z-market-scan-model-benchmark

## Scope

This is an offline sidecar benchmark for model interchangeability on Market Scan writer work. It used production PostgreSQL client/project/input rows exported read-only, plus the current writer-visible Market Scan doctrine from this repo. It did not mutate production, queues, providers, or customer-facing state. It also did not run the production Kimi judge or paid/provider sourcing measurement.

## Bottom Line

- Keep Qwen3.8 max as the production-safe baseline until a separate code change implements stage-specific model routing. The current production path is intentionally fail-closed there.
- Use GPT 5.6 Sol Max for research expansion and maximum sourcing-map breadth. It found the widest segment/arms universes on every client, but it is slower and can over-produce arms that need pruning/measurement.
- If Qwen is unavailable and we need to write letters, use GPT 5.5 xhigh first. It completed every cell fastest, kept copy controlled, and produced balanced maps without the Sol-scale explosion.
- Use Kimi k3 max as the second writer / voice challenger when capacity is healthy. It produced strong copy and wide maps, but the first Revopush attempt hit a Kimi profile quota/auth limit before the successful retry.
- Do not make GLM5.2 via Qwen CLI the first fallback. It completed, but the route/attestation is less clean and the first Connectro attempt hit Qwen API data inspection before a long successful retry.

## Recommended Model Routing

| Market Scan phase | Primary | Fallback | Why |
|---|---|---|---|
| Research map / adjacent discovery | GPT 5.6 Sol Max | GPT 5.5 xhigh | Sol produced the broadest segment maps on all three clients; GPT 5.5 is the faster, more controlled fallback. |
| Sourcing arms before provider measurement | GPT 5.6 Sol Max | Kimi k3 max if capacity healthy; Qwen3.8 max for conservative maps | Sol gives maximal coverage; Kimi is strong on arms/voice; Qwen is narrower but production-shaped. All arms still need measurement/pruning before purchase. |
| Copy packs / letters | Qwen3.8 max in current production path; GPT 5.5 xhigh if Qwen is unavailable | Kimi k3 max | GPT 5.5 was the best operational writer fallback: fast, complete, controlled. Kimi is a strong alternate voice but capacity-sensitive. |
| Whole-run emergency fallback | GPT 5.5 xhigh | GPT 5.6 Sol Max for quality-max, Kimi for second pass | GPT 5.5 had the cleanest end-to-end operational profile. Sol is better for breadth, worse for latency/overproduction. |
| Judge / acceptance | Not evaluated here | Existing Kimi judge remains current contract | This benchmark evaluated writer cells only. |

## Aggregate Matrix

| Model | Completed | Avg min | Segments | Arms | Packs | Touches | Arms / segment | Note |
|---|---:|---:|---:|---:|---:|---:|---:|---|
| Qwen3.8 max | 3/3 | 41.1 | 50 | 575 | 50 | 250 | 11.5 | Production-safe baseline; conservative, doctrine-compatible, not the broadest discovery engine. |
| GPT 5.5 xhigh | 3/3 | 16.4 | 48 | 661 | 48 | 240 | 13.8 | Best first fallback when Qwen is unavailable: fastest complete cells, balanced maps, controlled copy. |
| GPT 5.6 Sol Max | 3/3 | 41.8 | 73 | 4407 | 73 | 365 | 60.4 | Best research and breadth generator; use before pruning/measurement, not as default final writer without review. |
| Kimi k3 max | 3/3 | 29 | 39 | 1275 | 39 | 195 | 32.7 | Strong voice and sourcing challenger; use as second writer/reviewer when capacity profile is healthy. |
| GLM5.2 max via Qwen CLI | 3/3 | 31.1 | 39 | 236 | 39 | 195 | 6.1 | Useful compact independent check, but not first fallback due route attestation, provider-screen, and latency caveats. |