Anton Gladkov · 2026-08-09 · source ↗ · whole document (21)
Market Scan Model Benchmark - Human Scorecard Generated: 2026-08-08T22:35:20.254
## Model Notes
### GPT 5.6 Sol Max
Best pure quality/breadth candidate for research and sourcing-map generation. It produced 73 total segments and 4,407 arms across the three clients, far above the rest. That is useful for discovery and adjacent-market capture, but it should feed a pruning/measurement stage rather than go straight to purchase or final campaign without review.
### GPT 5.5 xhigh
Best operational fallback. It finished all three cells in 15-17 minutes each, with balanced segment counts and controlled copy. If Qwen is unavailable and Anton asks specifically who should write letters, this is the first answer.
### Qwen3.8 max
Best current production baseline because it is already the intended runtime lane and its outputs are conservative and doctrine-shaped. It was not the widest or fastest in this sidecar run, but it is the safest comparison anchor.
### Kimi k3 max
Strong second writer and style challenger. Good breadth on Connectro and Revopush and good copy self-audits. Operational caveat: one first-attempt quota/auth failure on Revopush before successful retry, so capacity preflight matters.
### GLM5.2 max via Qwen CLI
Completed all three validated outputs after retry, but should stay a tertiary fallback / compact independent check. It was narrower, route attestation was less clean, and Connectro first hit Qwen API data inspection before the successful retry.
## Operational Chronicle
- Full matrix was launched as a local sidecar with max parallelism; production stayed untouched.
- Primary full matrix initially completed 13/15.
- Revopush/Kimi first attempt failed on a Kimi profile usage/auth quota before final files; single-cell retry completed: 11 segments, 275 arms, 11 packs, 55 touches.
- Connectro/GLM first attempt failed with Qwen API 400 data_inspection_failed; single-cell retry completed after 27.4 minutes: 18 segments, 114 arms, 18 packs, 90 touches.
- Final roll-up after retries: 15/15 completed and JSON validation ok.