Anton Gladkov · 2026-08-09 · source ↗ · whole document (21)
Market Scan Model Benchmark - Human Scorecard Generated: 2026-08-08T22:35:20.254
## Operational Chronicle
- Full matrix was launched as a local sidecar with max parallelism; production stayed untouched.
- Primary full matrix initially completed 13/15.
- Revopush/Kimi first attempt failed on a Kimi profile usage/auth quota before final files; single-cell retry completed: 11 segments, 275 arms, 11 packs, 55 touches.
- Connectro/GLM first attempt failed with Qwen API 400 data_inspection_failed; single-cell retry completed after 27.4 minutes: 18 segments, 114 arms, 18 packs, 90 touches.
- Final roll-up after retries: 15/15 completed and JSON validation ok.
## What To Open
- Human index: index.html
- Markdown index: index.md
- Machine roll-up: report-data.json
- Scorecard data: model-scorecard.json
- Full outputs: runs/<client>/<model>/market_scan.md and market_scan.json
- Raw transcripts: runs/<client>/<model>/raw-transcript.log
## Caveats
This benchmark answers writer interchangeability. It does not prove production routing, paid/provider sourcing totals, lead purchase quality, deliverability, or Kimi judge acceptance. A production split-model Market Scan would need a small governed implementation: preserve Qwen as baseline, add explicit stage routing, keep deterministic validation to JSON/transport/math only, then compare with Kimi judge and a bounded live free-measurement smoke.
Я понимаю что нихуя не понятно из текста выше, но это потому что вы не утрудились скормить это своей GPT или Claude и поговорить об этом =)
GPT 5.5 xhigh находит 26 гипотез там где GPT 5.6 Sol Max находит 28, что говорит о том что GPT 5.5 xhigh невероятная модель всё еще на сегодняшний день, однако SOL MAX достаточно заебистый чтобы найти ~2300 связок названий компаний и должностей там где GPT удовлетворяется на ~450.
А маркет скан нет смысла делать снова и снова - ты его делаешь раз в жизни чтобы запланировать работу на весь следующий год.
То бишь GPT 5.5 Xhigh дает тебе по тем же гипотезам Total Addressible Market capacity ~ в 4-5раз ниже чем GPT 5.6 Sol Max.