Superintelligence CouncilСовет гения · sim.im

Anton Gladkov · 2026-08-09 · source ↗ · whole document (21)

Market Scan Model Benchmark - Human Scorecard Generated: 2026-08-08T22:35:20.254

Market Scan Model Benchmark - Human Scorecard

Generated: 2026-08-08T22:35:20.254Z

Run directory: /Users/antongladkov/.codex/worktrees/SLSBMB-Sender/market-scan-model-benchmark/var/market-scan-model-benchmark/20260808T201848Z-market-scan-model-benchmark

## Scope

This is an offline sidecar benchmark for model interchangeability on Market Scan writer work. It used production PostgreSQL client/project/input rows exported read-only, plus the current writer-visible Market Scan doctrine from this repo. It did not mutate production, queues, providers, or customer-facing state. It also did not run the production Kimi judge or paid/provider sourcing measurement.

## Bottom Line

- Keep Qwen3.8 max as the production-safe baseline until a separate code change implements stage-specific model routing. The current production path is intentionally fail-closed there.
- Use GPT 5.6 Sol Max for research expansion and maximum sourcing-map breadth. It found the widest segment/arms universes on every client, but it is slower and can over-produce arms that need pruning/measurement.
- If Qwen is unavailable and we need to write letters, use GPT 5.5 xhigh first. It completed every cell fastest, kept copy controlled, and produced balanced maps without the Sol-scale explosion.
- Use Kimi k3 max as the second writer / voice challenger when capacity is healthy. It produced strong copy and wide maps, but the first Revopush attempt hit a Kimi profile quota/auth limit before the successful retry.
- Do not make GLM5.2 via Qwen CLI the first fallback. It completed, but the route/attestation is less clean and the first Connectro attempt hit Qwen API data inspection before a long successful retry.