Anton Gladkov · 2026-08-09 · source ↗ · whole document (21)
Market Scan Model Benchmark - Human Scorecard Generated: 2026-08-08T22:35:20.254
### connectro
- Qwen3.8 max: 22 seg, 175 arms, 22 packs, 42.1 min; subjects: The license is approved. The bank still says no / Where the subscription money lands
- GPT 5.5 xhigh: 27 seg, 463 arms, 27 packs, 16.7 min; subjects: Named account rails for licensed gaming volume / A named account route for adult dating revenue
- GPT 5.6 Sol Max: 29 seg, 2301 arms, 29 packs, 36.8 min; subjects: When a licensed operator gets a blanket no / The fiat account after crypto registration
- Kimi k3 max: 21 seg, 886 arms, 21 packs, 39.3 min; subjects: When the bank letter arrives / The 0.4% behind your 1.5%
- GLM5.2 max via Qwen CLI: 18 seg, 114 arms, 18 packs, 27.4 min; subjects: Licensed, audited, and still declined / The checkbox declines the business, not the documents
### revopush
- Qwen3.8 max: 17 seg, 318 arms, 17 packs, 42.5 min; subjects: The fix is live, but not on their phones / Hotfixes for markets that never close
- GPT 5.5 xhigh: 12 seg, 144 arms, 12 packs, 17.4 min; subjects: Smaller OTA patches for finance apps / OTA patches for crypto app hotfixes
- GPT 5.6 Sol Max: 28 seg, 1349 arms, 28 packs, 48.4 min; subjects: When the banking fix is still downloading / A hotfix during a moving market
- Kimi k3 max: 11 seg, 275 arms, 11 packs, 10.9 min; subjects: The one-line compliance fix that weighs 25 MB / The day the market moves is the day your app gets judged
- GLM5.2 max via Qwen CLI: 13 seg, 63 arms, 13 packs, 33.7 min; subjects: OTA egress math for a banking app / The last mile of an incident fix
## Model Notes
### GPT 5.6 Sol Max
Best pure quality/breadth candidate for research and sourcing-map generation. It produced 73 total segments and 4,407 arms across the three clients, far above the rest. That is useful for discovery and adjacent-market capture, but it should feed a pruning/measurement stage rather than go straight to purchase or final campaign without review.
### GPT 5.5 xhigh
Best operational fallback. It finished all three cells in 15-17 minutes each, with balanced segment counts and controlled copy. If Qwen is unavailable and Anton asks specifically who should write letters, this is the first answer.
### Qwen3.8 max
Best current production baseline because it is already the intended runtime lane and its outputs are conservative and doctrine-shaped. It was not the widest or fastest in this sidecar run, but it is the safest comparison anchor.
### Kimi k3 max
Strong second writer and style challenger. Good breadth on Connectro and Revopush and good copy self-audits. Operational caveat: one first-attempt quota/auth failure on Revopush before successful retry, so capacity preflight matters.