Anton Gladkov · 2026-08-09 · source ↗ · whole document (21)
Market Scan Model Benchmark - Human Scorecard Generated: 2026-08-08T22:35:20.254
### Qwen3.8 max
Best current production baseline because it is already the intended runtime lane and its outputs are conservative and doctrine-shaped. It was not the widest or fastest in this sidecar run, but it is the safest comparison anchor.
### Kimi k3 max
Strong second writer and style challenger. Good breadth on Connectro and Revopush and good copy self-audits. Operational caveat: one first-attempt quota/auth failure on Revopush before successful retry, so capacity preflight matters.
### GLM5.2 max via Qwen CLI
Completed all three validated outputs after retry, but should stay a tertiary fallback / compact independent check. It was narrower, route attestation was less clean, and Connectro first hit Qwen API data inspection before the successful retry.
## Operational Chronicle
- Full matrix was launched as a local sidecar with max parallelism; production stayed untouched.
- Primary full matrix initially completed 13/15.
- Revopush/Kimi first attempt failed on a Kimi profile usage/auth quota before final files; single-cell retry completed: 11 segments, 275 arms, 11 packs, 55 touches.
- Connectro/GLM first attempt failed with Qwen API 400 data_inspection_failed; single-cell retry completed after 27.4 minutes: 18 segments, 114 arms, 18 packs, 90 touches.
- Final roll-up after retries: 15/15 completed and JSON validation ok.
## What To Open
- Human index: index.html
- Markdown index: index.md
- Machine roll-up: report-data.json
- Scorecard data: model-scorecard.json
- Full outputs: runs/<client>/<model>/market_scan.md and market_scan.json
- Raw transcripts: runs/<client>/<model>/raw-transcript.log