Superintelligence CouncilСовет гения · sim.im

Anton Gladkov · 2026-08-09 · 01-market-scan · source ↗

Market Scan Model Benchmark - Human Scorecard Generated: 2026-08-08T22:35:20.254

Market Scan Model Benchmark - Human Scorecard

Generated: 2026-08-08T22:35:20.254Z

Run directory: /Users/antongladkov/.codex/worktrees/SLSBMB-Sender/market-scan-model-benchmark/var/market-scan-model-benchmark/20260808T201848Z-market-scan-model-benchmark

## Scope

This is an offline sidecar benchmark for model interchangeability on Market Scan writer work. It used production PostgreSQL client/project/input rows exported read-only, plus the current writer-visible Market Scan doctrine from this repo. It did not mutate production, queues, providers, or customer-facing state. It also did not run the production Kimi judge or paid/provider sourcing measurement.

## Bottom Line

- Keep Qwen3.8 max as the production-safe baseline until a separate code change implements stage-specific model routing. The current production path is intentionally fail-closed there.
- Use GPT 5.6 Sol Max for research expansion and maximum sourcing-map breadth. It found the widest segment/arms universes on every client, but it is slower and can over-produce arms that need pruning/measurement.
- If Qwen is unavailable and we need to write letters, use GPT 5.5 xhigh first. It completed every cell fastest, kept copy controlled, and produced balanced maps without the Sol-scale explosion.
- Use Kimi k3 max as the second writer / voice challenger when capacity is healthy. It produced strong copy and wide maps, but the first Revopush attempt hit a Kimi profile quota/auth limit before the successful retry.
- Do not make GLM5.2 via Qwen CLI the first fallback. It completed, but the route/attestation is less clean and the first Connectro attempt hit Qwen API data inspection before a long successful retry.

## Recommended Model Routing

| Market Scan phase | Primary | Fallback | Why |
|---|---|---|---|
| Research map / adjacent discovery | GPT 5.6 Sol Max | GPT 5.5 xhigh | Sol produced the broadest segment maps on all three clients; GPT 5.5 is the faster, more controlled fallback. |
| Sourcing arms before provider measurement | GPT 5.6 Sol Max | Kimi k3 max if capacity healthy; Qwen3.8 max for conservative maps | Sol gives maximal coverage; Kimi is strong on arms/voice; Qwen is narrower but production-shaped. All arms still need measurement/pruning before purchase. |
| Copy packs / letters | Qwen3.8 max in current production path; GPT 5.5 xhigh if Qwen is unavailable | Kimi k3 max | GPT 5.5 was the best operational writer fallback: fast, complete, controlled. Kimi is a strong alternate voice but capacity-sensitive. |
| Whole-run emergency fallback | GPT 5.5 xhigh | GPT 5.6 Sol Max for quality-max, Kimi for second pass | GPT 5.5 had the cleanest end-to-end operational profile. Sol is better for breadth, worse for latency/overproduction. |
| Judge / acceptance | Not evaluated here | Existing Kimi judge remains current contract | This benchmark evaluated writer cells only. |

## Aggregate Matrix

| Model | Completed | Avg min | Segments | Arms | Packs | Touches | Arms / segment | Note |
|---|---:|---:|---:|---:|---:|---:|---:|---|
| Qwen3.8 max | 3/3 | 41.1 | 50 | 575 | 50 | 250 | 11.5 | Production-safe baseline; conservative, doctrine-compatible, not the broadest discovery engine. |
| GPT 5.5 xhigh | 3/3 | 16.4 | 48 | 661 | 48 | 240 | 13.8 | Best first fallback when Qwen is unavailable: fastest complete cells, balanced maps, controlled copy. |
| GPT 5.6 Sol Max | 3/3 | 41.8 | 73 | 4407 | 73 | 365 | 60.4 | Best research and breadth generator; use before pruning/measurement, not as default final writer without review. |
| Kimi k3 max | 3/3 | 29 | 39 | 1275 | 39 | 195 | 32.7 | Strong voice and sourcing challenger; use as second writer/reviewer when capacity profile is healthy. |
| GLM5.2 max via Qwen CLI | 3/3 | 31.1 | 39 | 236 | 39 | 195 | 6.1 | Useful compact independent check, but not first fallback due route attestation, provider-screen, and latency caveats. |

## Cell Matrix

| Client | Model | Status | Min | Segments | Arms | Packs | Touches | Self-fit research/sourcing/copy |
|---|---|---:|---:|---:|---:|---:|---:|---|
| connectro | Qwen3.8 max | completed | 42.1 | 22 | 175 | 22 | 110 | good/acceptable/good |
| connectro | GPT 5.5 xhigh | completed | 16.7 | 27 | 463 | 27 | 135 | good/acceptable/acceptable |
| connectro | GPT 5.6 Sol Max | completed | 36.8 | 29 | 2301 | 29 | 145 | good/good/good |
| connectro | Kimi k3 max | completed | 39.3 | 21 | 886 | 21 | 105 | good/good/good |
| connectro | GLM5.2 max via Qwen CLI | completed | 27.4 | 18 | 114 | 18 | 90 | good/acceptable/good |
| masha | Qwen3.8 max | completed | 38.7 | 11 | 82 | 11 | 55 | good/acceptable/good |
| masha | GPT 5.5 xhigh | completed | 15.2 | 9 | 54 | 9 | 45 | good/acceptable/good |
| masha | GPT 5.6 Sol Max | completed | 40.2 | 16 | 757 | 16 | 80 | good/acceptable/good |
| masha | Kimi k3 max | completed | 36.7 | 7 | 114 | 7 | 35 | good/acceptable/good |
| masha | GLM5.2 max via Qwen CLI | completed | 32.3 | 8 | 59 | 8 | 40 | good/acceptable/good |
| revopush | Qwen3.8 max | completed | 42.5 | 17 | 318 | 17 | 85 | good/acceptable/good |
| revopush | GPT 5.5 xhigh | completed | 17.4 | 12 | 144 | 12 | 60 | good/acceptable/good |
| revopush | GPT 5.6 Sol Max | completed | 48.4 | 28 | 1349 | 28 | 140 | good/good/good |
| revopush | Kimi k3 max | completed | 10.9 | 11 | 275 | 11 | 55 | good/acceptable/good |
| revopush | GLM5.2 max via Qwen CLI | completed | 33.7 | 13 | 63 | 13 | 65 | good/acceptable/good |

## Per-Client Reading Notes

### masha
- Qwen3.8 max: 11 seg, 82 arms, 11 packs, 38.7 min; subjects: Distribution after the model works / EMEA supply, partnerships
- GPT 5.5 xhigh: 9 seg, 54 arms, 9 packs, 15.2 min; subjects: When AI starts needing a commercial spine / EMEA partnerships that carry revenue
- GPT 5.6 Sol Max: 16 seg, 757 arms, 16 packs, 40.2 min; subjects: EMEA partnerships after product-market fit / The partner layer between supply and demand
- Kimi k3 max: 7 seg, 114 arms, 7 packs, 36.7 min; subjects: After the founder-led deals / Your EMEA partner coverage
- GLM5.2 max via Qwen CLI: 8 seg, 59 arms, 8 packs, 32.3 min; subjects: Partnerships at {{prospect.companyName}}: who should own them / EMEA partner growth at {{prospect.companyName}}

### connectro
- Qwen3.8 max: 22 seg, 175 arms, 22 packs, 42.1 min; subjects: The license is approved. The bank still says no / Where the subscription money lands
- GPT 5.5 xhigh: 27 seg, 463 arms, 27 packs, 16.7 min; subjects: Named account rails for licensed gaming volume / A named account route for adult dating revenue
- GPT 5.6 Sol Max: 29 seg, 2301 arms, 29 packs, 36.8 min; subjects: When a licensed operator gets a blanket no / The fiat account after crypto registration
- Kimi k3 max: 21 seg, 886 arms, 21 packs, 39.3 min; subjects: When the bank letter arrives / The 0.4% behind your 1.5%
- GLM5.2 max via Qwen CLI: 18 seg, 114 arms, 18 packs, 27.4 min; subjects: Licensed, audited, and still declined / The checkbox declines the business, not the documents

### revopush
- Qwen3.8 max: 17 seg, 318 arms, 17 packs, 42.5 min; subjects: The fix is live, but not on their phones / Hotfixes for markets that never close
- GPT 5.5 xhigh: 12 seg, 144 arms, 12 packs, 17.4 min; subjects: Smaller OTA patches for finance apps / OTA patches for crypto app hotfixes
- GPT 5.6 Sol Max: 28 seg, 1349 arms, 28 packs, 48.4 min; subjects: When the banking fix is still downloading / A hotfix during a moving market
- Kimi k3 max: 11 seg, 275 arms, 11 packs, 10.9 min; subjects: The one-line compliance fix that weighs 25 MB / The day the market moves is the day your app gets judged
- GLM5.2 max via Qwen CLI: 13 seg, 63 arms, 13 packs, 33.7 min; subjects: OTA egress math for a banking app / The last mile of an incident fix

## Model Notes

### GPT 5.6 Sol Max
Best pure quality/breadth candidate for research and sourcing-map generation. It produced 73 total segments and 4,407 arms across the three clients, far above the rest. That is useful for discovery and adjacent-market capture, but it should feed a pruning/measurement stage rather than go straight to purchase or final campaign without review.
### GPT 5.5 xhigh
Best operational fallback. It finished all three cells in 15-17 minutes each, with balanced segment counts and controlled copy. If Qwen is unavailable and Anton asks specifically who should write letters, this is the first answer.

### Qwen3.8 max
Best current production baseline because it is already the intended runtime lane and its outputs are conservative and doctrine-shaped. It was not the widest or fastest in this sidecar run, but it is the safest comparison anchor.

### Kimi k3 max
Strong second writer and style challenger. Good breadth on Connectro and Revopush and good copy self-audits. Operational caveat: one first-attempt quota/auth failure on Revopush before successful retry, so capacity preflight matters.

### GLM5.2 max via Qwen CLI
Completed all three validated outputs after retry, but should stay a tertiary fallback / compact independent check. It was narrower, route attestation was less clean, and Connectro first hit Qwen API data inspection before the successful retry.

## Operational Chronicle

- Full matrix was launched as a local sidecar with max parallelism; production stayed untouched.
- Primary full matrix initially completed 13/15.
- Revopush/Kimi first attempt failed on a Kimi profile usage/auth quota before final files; single-cell retry completed: 11 segments, 275 arms, 11 packs, 55 touches.
- Connectro/GLM first attempt failed with Qwen API 400 data_inspection_failed; single-cell retry completed after 27.4 minutes: 18 segments, 114 arms, 18 packs, 90 touches.
- Final roll-up after retries: 15/15 completed and JSON validation ok.

## What To Open

- Human index: index.html
- Markdown index: index.md
- Machine roll-up: report-data.json
- Scorecard data: model-scorecard.json
- Full outputs: runs/<client>/<model>/market_scan.md and market_scan.json
- Raw transcripts: runs/<client>/<model>/raw-transcript.log

## Caveats

This benchmark answers writer interchangeability. It does not prove production routing, paid/provider sourcing totals, lead purchase quality, deliverability, or Kimi judge acceptance. A production split-model Market Scan would need a small governed implementation: preserve Qwen as baseline, add explicit stage routing, keep deterministic validation to JSON/transport/math only, then compare with Kimi judge and a bounded live free-measurement smoke.
Я понимаю что нихуя не понятно из текста выше, но это потому что вы не утрудились скормить это своей GPT или Claude и поговорить об этом =)

GPT 5.5 xhigh находит 26 гипотез там где GPT 5.6 Sol Max находит 28, что говорит о том что GPT 5.5 xhigh невероятная модель всё еще на сегодняшний день, однако SOL MAX достаточно заебистый чтобы найти ~2300 связок названий компаний и должностей там где GPT удовлетворяется на ~450.

А маркет скан нет смысла делать снова и снова - ты его делаешь раз в жизни чтобы запланировать работу на весь следующий год.

То бишь GPT 5.5 Xhigh дает тебе по тем же гипотезам Total Addressible Market capacity ~ в 4-5раз ниже чем GPT 5.6 Sol Max.

И все это стоит просто лишних 20-30 минут заебистого рассуждения модели.

Вы не в моей нише, поэтому вы не понимаете насколько это важно, но это блядь важно просто шо пиздец)

И другие цифры там есть веселые которые можно трактовать занятным образом, оставляю вас гадать на них. Когда я релизну свой продукт я думаю все в итоге объяснится его же юзерами которые будут тут рассказывать свои наблюдения)
Резюмируя: в мелочах кажется что модели недалеко друг от друга ушли, но правда в том что GPT 5.6 Sol Max ушел от GPT 5.5 xhigh настолько далеко, что он выдает в бизнес выражении контента за лишние 15 минут рассуждения как разница между 1 годом и 10 годами скурпулезной работы отдела продаж по рисечу.

Это реально кроме шуток ускорение отдела продаж в 8 раз на десятилетнем отрезке, за лишние 20 минут копеечного палева токенов.

Это сравнение для самых маленьких. Если там копаться в деталях, мы легко найдем ещё бОльшее ускорение, а так же вещи которые предыдущими моделями просто в принице были недостижимы, типа определенных игрищ текстовыми конструкциями, но сегодня нет сил расписывать это)