Model Guidance
Last refreshed: 2026-07-25 (trigger: opus-5. Claude Opus 5 (Anthropic, GA 07-24) rostered as the new Frontier-agent leader β supersedes opus-4-8 at the SAME $5/$25 (price-neutral strict upgrade). AA Intelligence Index v4.1 = 61 (max), narrowly #1 of all models (Fable-5 60 / Sol 59 / K3 57); joint-#1 AA Coding Agent Index (Claude Code @ xhigh); record GDPval-AA v2 1861 Elo + AA-Briefcase 1720 Elo (+146 vs Fable-5); beats Fable-5 on OSWorld-2.0 computer-use at ~1/3 cost; ~26% lower cost-per-task than Fable-5; 1M ctx (secondary-source β official page stated price only, confirm at model card). ZDR-CLEAN (no data-retention requirement, unlike Fable-5); cyber classifiers fire ~85% less than Fable-5 but still refuse offensive-security work β opus-4-8 is the documented fallback (wire it); behind Mythos-5 (Glasswing-only) on cyber. Reshapes the roster top: opus-5 is the default frontier pick AND cheaper-per-task than Fable-5 β Fable-5 narrows to a niche (verbose deliberate style; NO independent SWE-bench-Pro opus-5-vs-Fable-5 number pulled β open). Proposals (NOT auto-applied): council Claude seat opus-4-8βopus-5 (SUPERSEDES the Fable-5-synthesis pilot β same-or-better quality at half cost, ZDR-clean, no availability-fallback headache); name opus-5 the coding-agent default; add opus-5 row to costs.py. DELIBERATE HOLD RESPECTED: the opus-4-7 trio (sleep_build aloop, cadence/daemon.py:258, generate_html_v2.py:54-55) stays held per the 07-18 self-terminating rule β opus-5 is simply the new eventual target at the ~Oct-2026 reaffirm; NOT re-surfaced this run. Also flagged: audit shows scripts/council/core.py:22 still pins opus-4-7 (doc claimed opus-4-8 β drift) β fold to opus-5 in the seat swap. Prior refresh (07-18, kimi-k3): Kimi K3 (Moonshot, GA 07-16, 2.8T sparse MoE β largest open-weight model ever announced) rostered as the first open-weight(-promised) model in the frontier tier: AA Intelligence Index v4.1 = 57 (#3 per launch coverage / #4 per AA's own page β discrepancy recorded; behind Fable-5 60 and Sol 59, ahead of Terra 55), $0.94/AA-task, 1M ctx, text+image in, cache-hit $0.30. Pricing regime change: $3/$15 = 4Γ k2.6's input β K3 is premium-band (competes with Terra / grok-4.5 / Sonnet), NOT a k2.6 successor; k2.6 RETAINED on all strong-cheap synthesis sites. Arena.ai frontend-code #1 ahead of Fable-5 (independent); coding/agentic claims (Terminal-Bench 2.1 88.3 KimiCode-harness, FrontierSWE 81.2, SWE Marathon 42.0) are vendor-only β same re-confirm-at-AA posture as glm-5.2. WATCH status, zero routing changes: weights + license land ~Jul-27, OpenRouter upstream capacity-limited (frequent 429s), single 'max' reasoning-effort with real thinking-token overhead. opus-4-7β4-8: DELIBERATE HOLD executed this review per the 07-10 self-terminating rule (6 re-surfaces, declined by inaction; reaffirm ~Oct 2026 β the system stops asking). Prior refresh (07-10): gpt-5.6 Sol/Terra/Luna rostered + scored; AA index rescaled to v4.1; council seat swap gpt-5.5βsol proposed, still pending bio-zack; METR reward-hacking caveat on Sol still unconfirmed at primary.)
Auto-maintained by flows/model-guidance-review (run via stepwise run flows/model-guidance-review --name model-review --wait). See the footer for methodology.
This doc is the single source of truth for which LLM to use for which task across vita, vita's stepwise flows, gumball, and cadence. It's refreshed whenever a new model drops; see Β§ Methodology.
TL;DR β approved roster (June 2026)
| Tier | Model | Provider | Input $/MTok | Output $/MTok | Released | Primary use |
|---|---|---|---|---|---|---|
| Frontier-lead (NEW) | claude-opus-5 | Anthropic | $5.00 | $25.00 | 2026-07-24 | Supersedes opus-4-8 at the SAME price. Narrowly #1 AA Intelligence Index v4.1 (61 max), joint-#1 Coding Agent Index, record GDPval-AA v2 (1861) + AA-Briefcase (1720), beats Fable-5 on OSWorld-2.0 at ~1/3 cost. New default for long-horizon agent work, coding, /sleep build (target), exec loop. ZDR-clean; cyber-refusal β opus-4-8 fallback. β |
| Frontier-max (niche) | claude-fable-5 | Anthropic | $10.00 | $50.00 | 2026-06-09 | Narrowed by opus-5 (which is #1 AA AND ~26% cheaper-per-task). Reserve for the rare case its 2Γ output cost + verbose deliberate style is specifically justified; NO independent SWE-bench-Pro opus-5-vs-Fable-5 head-to-head yet. SWE-bench Pro 80.3. β‘ |
| Frontier-agent (superseded) | claude-opus-4-8 | Anthropic | $5.00 | $25.00 | 2026-Q2 | Superseded by opus-5 (same price, strictly better). RETAINED as the documented refusal/availability fallback for opus-5 + Fable-5 callers; still live at council + the held opus-4-7-trio's eventual target until swapped. |
| Frontier-reason | gpt-5.5 | OpenAI | $5.00 | $30.00 | 2026-04 | Hardest reasoning, council seat, non-coding strategy |
| Frontier-reason-pro | gpt-5.5-pro | OpenAI | $30.00 | $180.00 | 2026-04 | /council only β rare, highest-stakes (no 5.6-pro exists yet β unchanged) |
| Frontier-reason (successor) | gpt-5.6-sol | OpenAI | $5.00 | $30.00 | 2026-07-09 | Direct gpt-5.5 successor at the SAME price. #1 AA Coding Agent Index (80), SOTA Terminal-Bench 2.1 (88.8). Proposed council-seat swap. β§ |
| Frontier-reason-value (new) | gpt-5.6-terra | OpenAI | $2.50 | $15.00 | 2026-07-09 | Near-Sol at half price (AA v4.1 55, Coding Agent 77.4 β edges Fable-5's 77.2). Vellum's "default tier". Candidate for future cost-sensitive frontier-reason paths. β§ |
| Cheap-frontier-agent (new) | gpt-5.6-luna | OpenAI | $1.00 | $6.00 | 2026-07-09 | Terminal-Bench 84.7 (> grok-4.5 83.3 > Opus 4.8) at $1/$6. Third arm for the bounded coding pilot. MRCR 41.3% β long-context COLLAPSES despite nominal 1M; never route >100K ctx. β§ |
| Frontier-multimodal | gemini-3.1-pro | $2.00 | $12.00 | 2026-Q1 | Vision-heavy tasks, 1M+ context, benchmarks leader | |
| Strong-cheap | kimi-k2.6 | Moonshot | $0.75 | $4.66 | 2026-04-21 | Agent swarms, long-horizon coding, OSS fallback. Retained β its 1M ctx serves the docs-audit / newsletter / consolidator / council-review synthesis roles k2.7-code can't. |
| Strong-cheap-code (pilot) | kimi-k2.7-code | Moonshot | $0.74 | $3.50 | 2026-06-12 | Coding-SPECIALIST successor to k2.6: ~30% fewer thinking tokens, MCP-Mark 81.1. But 256K ctx (NOT 1M) β does not serve k2.6's long-context roles. Vendor-only benchmarks. No current vita site fits (all kimi uses are 1M-synthesis); rostered for future bounded coding agents. ΒΆ |
| Open-frontier (watch) | kimi-k3 | Moonshot | $3.00 | $15.00 | 2026-07-16 | First open-weight(-promised) frontier-tier model (AA v4.1 = 57, #3-4). NOT a k2.6 successor β 4Γ input price puts it in the Terra/grok-4.5/Sonnet band. WATCH: no routing until weights + license land (~Jul-27) and OpenRouter capacity stabilizes (frequent 429s today). β |
| Strong-cheap-reason | deepseek-v4-pro | DeepSeek | $1.74β | $3.48β | 2026-04-24 | Cheap-frontier reasoning + agentic coding (batch only β 34 TPS). MIT-licensed. |
| Strong-cheap-code | glm-5.2 | Z.ai (Zhipu) | $1.40 | $4.40 | 2026-06-16 | Coding where Opus is overkill, cost-sensitive builds. SWE-bench Pro 62.1 (top open-weights, beats GPT-5.5); usable 1M ctx (IndexShare); high/xhigh thinking. glm-5.1 ($1.05/$3.50) kept as the cheaper fallback for high-volume non-coding passes (feed editors). Β§ |
| Balanced | claude-sonnet-4-6 | Anthropic | ~$3.00 | ~$15.00 | 2026-Q1 | Default for most vita chat + proactive composers. Incumbent β hold until sonnet-5 pilot clears voice-eval. |
| Balanced (successor pilot) | claude-sonnet-5 | Anthropic | $2.00β$3.00 | $10.00β$15.00 | 2026-06-30 | Direct sonnet-4-6 successor, perf "close to Opus 4.8", most-agentic Sonnet yet. Intro $2/$10 through Aug 31, then $3/$15 (== 4.6 per-token). NEW TOKENIZER emits ~1.0-1.35Γ more tokens β "cost-neutral" is intro-only. Pilot on non-voice analytical path first, NOT an auto-swap. ΒΆΒΆ |
| Budget-MoE | minimax-m2.7 | MiniMax | $0.30 | $1.20 | 2026-03-18 | Summarization, synthesis, background research (incumbent β bake-off vs V4-Flash live) |
| Budget-MoE (successor) | minimax-m3 | MiniMax | $0.30 | $1.20 | 2026-05-31 | Direct m2.7 successor at the SAME price: 1M ctx (vs 200K) + native image/video multimodal, MSA sparse attention (~1/20 compute at 1M). Unverified benchmarks. Fold into the budget bake-off. ΒΆΒΆΒΆ |
| Budget-MoE (challenger) | deepseek-v4-flash | DeepSeek | $0.14 | $0.28 | 2026-04-24 | Direct M2.7 challenger: 4x cheaper output, 5x context (1M), Intel Index 47 vs 50. MIT-licensed. PILOT |
| Budget-fast | gemini-3.1-flash-lite | ~$0.10 | ~$0.40 | 2026-Q1 | High-volume triage: media scoring, daily rollup, trend analysis | |
| Honesty-specialist | grok-4.20 | xAI | $20.00 | $60.00 | 2026-03-31 | Council seat only β #1 on Omniscience non-hallucination |
| Reasoning-cheap (xAI) | grok-4.3 | xAI | $1.25 | $2.50 | 2026-04-30 | SHADOW PILOT β 5th council seat alongside grok-4.20 until AA publishes Omniscience component. 16-24x cheaper than 4.20, Intel Index 53 (+5), 207 TPS (#1). No model card yet. |
| Frontier-agent-cheap (pilot) | grok-4.5 | SpaceXAI (ex-xAI) | $2.00 | $6.00 | 2026-07-08 | NEW β pilot, NOT auto-swapped. Cheaper-frontier coding/agentic: AA Intel Index 54 (#4), Coding Agent Index 76, ~2Γ token efficiency (4.2Γ fewer output tokens than Opus 4.8 on SWE-bench Pro). Beats Opus 4.8 on Terminal Bench 2.1 (83.3) + SWE Marathon (29.0); trails on SWE-bench Pro (64.7 vs 64.3β69.2, harness-dependent). Cached input $0.50, ~80 TPS. 500K ctx (REGRESSION from grok-4.3's 1M); NO Omniscience number (not a council/honesty pick); 1-day-old, AA-only benchmarks; modality text-only-confirmed. β¦ |
β V4-Pro on a 75% promo through 2026-05-05: $0.435/$0.87 per MTok. Standard pricing resumes May 6. Cache-hit input is $0.145.
β‘claude-fable-5 (GA 2026-06-09): adaptive thinking always-on (no budget_tokens, no thinking:{type:disabled} β both 400), raw chain-of-thought never returned (summaries via thinking.display), safety classifiers can return stop_reason:refusal (opt-in fallback to opus-4-8 required in any caller), and 30-day data retention required β NOT available under ZDR. Reserve for highest-stakes single calls; 2Γ Opus output cost. Sibling claude-mythos-5 (identical specs, no classifiers) is Project-Glasswing-only β not adoptable by vita. Bench note (UPDATED 2026-07-10): SWE-bench Pro 80.3 confirmed (official + Vellum); AA Intelligence Index rescaled to v4.1 β Fable-5 (max) = 60, still #1, one point above gpt-5.6-sol (direct-confirmed at artificialanalysis.ai this refresh; supersedes the "~65" carried from the old index version β the two scales are NOT comparable). SWE-bench Verified ~95 remains single-source-aggregation, treat as indicative. New comparative datapoints from the 5.6 launch: Fable-5 keeps decisive leads on SWE-bench Pro (80 vs Sol 64.6), AA-Briefcase (rubric 56% vs 42%, analytical Elo 1764 vs 1592), GDPval-AA-v2 and HealthBench Professional; Sol beats it on Coding Agent Index (80 vs 77.2), Terminal-Bench 2.1 (88.8 vs 86.0) and Agents' Last Exam (53.6 vs 40.5) at ~1/3 the cost-per-task. Availability event (verified 2026-07-05): Fable 5 + Mythos 5 were export-control-SUSPENDED by a US Commerce Dept directive on 2026-06-12 (barring access by any foreign national, globally) β a full shutdown β and RESTORED 2026-07-01 after the directive was lifted. The slot now carries a standing geopolitical-availability risk on top of the refusal caveat: any Fable-5 caller needs the opus-4-8 fallback wired for BOTH refusal AND availability. (The 06-23 review saw "Fable-5 pulled by export control" and dismissed it as an unverified aggregator rumor β it was real. Do not auto-dismiss model-status rumors; verify.)
βclaude-opus-5 (Anthropic, GA 2026-07-24, API id claude-opus-5): direct successor to opus-4-8 at IDENTICAL $5/$25 (Anthropic: "greatly improved performance for the same cost as Opus 4.8"; +22% on hardest agentic coding vs prior Opus, organic-chem +10.2pts, protein +7.7pts, Frontier-Bench v0.1 more than doubles 4.8). AA (direct-fetched artificialanalysis.ai + AA article + AA on X): Intelligence Index v4.1 = 61 (max) β narrowly the single most intelligent model, edging Fable-5 60 / Sol 59 / K3 57; ~26% lower Cost-per-Task than Fable-5 ($17.79 vs $22.30 at max; at HIGH effort $10.41/task BEATS Fable-5 by 32 Elo at 47% of its cost); joint-first AA Coding Agent Index (run in Claude Code @ xhigh) + highest SWE-Atlas-QnA; GDPval-AA v2 1861 Elo (>100 ahead of Fable-5/Sol); AA-Briefcase 1720 Elo (+146 vs Fable-5). Anthropic's own: CursorBench 3.2 within 0.5% of Fable-5 at half cost; ARC-AGI-3 ~3Γ next-best; OSWorld-2.0 (computer use) surpasses Fable-5 at ~1/3 cost; Zapier AutomationBench 1.5Γ next-best at same cost/task. New default on Claude Max, strongest on Claude Pro; 2Γ fast-mode (~$10/$50, ~2.5Γ speed). Context 1M per launch coverage (BigGo/AI-Weekly/MarkTechPost) β official anthropic.com/news + system-card blurb stated PRICE only this run; treat 1M as secondary-source until confirmed at the model card. Safety: NO data-retention requirement (ZDR-clean β a real advantage over Fable-5's mandatory 30-day/no-ZDR). Cyber classifiers intervene ~85% less often than Fable-5 but the model still blocks binary vuln-scanning / pentest / exploit-gen and can return refusal stops β Anthropic ships an opus-4-8 fallback for flagged requests; wire it in any opus-5 caller (same discipline as Fable-5, minus the availability risk). Remains BEHIND Mythos-5 (Project-Glasswing-only, not vita-adoptable) on cyber-exploitation β do not route offensive-security here. Posts the LOWEST score on Anthropic's automated behavioral-audit scale of any recent Claude (best-behaved). NO independent SWE-bench Pro/Verified head-to-head vs Fable-5 pulled this run (AA leads are Intelligence Index / GDPval / AA-Briefcase / Coding-Agent-Index); OpenRouter availability unverified this run (Anthropic API day-one). Sources: anthropic.com/news/claude-opus-5, the Opus-5 system card PDF, artificialanalysis.ai/articles/claude-opus-5-leader-agentic-knowledge-work, x.com/ArtificialAnlys.
Β§glm-5.2 (Z.ai/Zhipu, GA 2026-06-16, 753B open weights, MIT): SWE-bench Pro 62.1 > GPT-5.5 58.6 > glm-5.1 58.4; FrontierSWE 74.4% and MCP-Atlas 77.0 near-tie Opus 4.8 (75.1 / 77.8). 1M usable ctx via IndexShare (~2.9Γ FLOP cut at full length), high/xhigh thinking-effort levels, TTFT ~10.8s (async-friendly, NOT chat-latency). Pricing / context / release-date direct-verified this refresh (openrouter.ai + llm-stats.com agree); benchmarks are from the Z.ai launch coverage (VentureBeat) and are NOT yet on artificialanalysis.ai β re-confirm next refresh, and AA Intelligence Index is unpublished (general-reasoning score held = glm-5.1). China data-jurisdiction risk via API/OpenRouter, but MIT weights allow self-host; vita already scopes glm to non-sensitive code paths, so OpenRouter routing stays in policy. OpenRouter lists it text-in/out β not a vision/OCR model.
ΒΆkimi-k2.7-code (Moonshot, GA 2026-06-12, 1T total / 32B active MoE, Modified-MIT open weights, OpenRouter/DeepInfra $0.74/$3.50 β official Kimi API $0.95/$4.00; one aggregator lists $1.14/$4.80 outlier): coding-SPECIALIZED variant (NOT a general K2.7), forced thinking mode, ~30% fewer reasoning tokens than K2.6. Context 262K (256K) β a hard drop from k2.6's 1M. ALL benchmarks are Moonshot's own proprietary suites (Kimi Code Bench v2 +21.8% / Program Bench +11.0% / MLS Bench Lite +31.5% / MCP-Mark Verified 81.1 β all vs K2.6); devops.com confirms no independent third-party results on SWE-bench Verified/Pro, LiveCodeBench, or GPQA as of release. The SWE-bench Pro 58.6 / Verified 60.4 floating on aggregator pages are vendor-reported and suspect (58.6 == K2.6's exact number). Same "vendor-only, re-confirm at AA" posture as glm-5.2. text+image input.
ΒΆΒΆclaude-sonnet-5 (Anthropic, GA 2026-06-30): direct successor to sonnet-4-6 and Anthropic's new default on Free/Pro. Intro $2/$10 through 2026-08-31, standard $3/$15 after (== sonnet-4-6 per-token). CRITICAL for cost math: an updated tokenizer (same change shipped with Opus 4.7) maps the same input to ~1.0-1.35Γ more tokens β so "roughly cost-neutral" is an intro-window framing, and effective per-request cost can exceed 4.6 despite equal per-token pricing (and rises Sep 1). Perf "close to Opus 4.8", "substantial improvement over Sonnet 4.6" β Anthropic gives qualitative claims, no headline numbers (pull AA/SWE-bench next refresh). Safety note: improved agentic safety vs 4.6 but "substantially poorer on cybersecurity tasks vs Opus" β do not route security/vuln work here. Context window not stated in the announcement β confirm before long-context routing. Pilot-with-voice-eval, not an auto-swap: voice-critical paths (coaching, podcast transcripts, silicon-zack chat) need a drift check first.
ΒΆΒΆΒΆminimax-m3 (MiniMax, GA 2026-05-31, first open-weights model combining frontier coding + 1M ctx + native multimodality): OpenRouter flat $0.30/$1.20 (same as m2.7) with prompt-caching cutting effective cost 60-80% on repeated context. Native text+image+video input (vs m2.7 text-only), 1M context (vs m2.7's 200K), MiniMax Sparse Attention (MSA β KV-block selection, ~1/20 per-token compute at 1M). Pricing disagreement recorded: OpenRouter flat $0.30/$1.20 vs direct-API $0.60/$2.40-standard + launch-promo + a 512K long-context band that DOUBLES the rate β verify whether OpenRouter's rate is promo or permanent before high-volume routing. Benchmarks unverified (no AA numbers pulled this run; launch coverage claims GPT-5.5/Opus-4.7-tier coding β vendor framing). Missed by the last several refreshes (released May 31; doc last refreshed Jun 23).
β§gpt-5.6 family (OpenAI, GA 2026-07-09; Sol previewed since 06-26; SiliconANGLE puts broad rollout completing 07-10 after a US Commerce Dept pre-release review β the Fable-5 export-suspension pattern is now apparently a standing regime for frontier launches). Model IDs gpt-5.6-sol / -terra / -luna; bare gpt-5.6 routes to Sol. ALL tiers: 1M ctx, 128K max output, knowledge cutoff 2026-02-16, six reasoning-effort levels (none/low/medium/high/xhigh/max), all live on OpenRouter day-one (Terra listing direct-verified: $2.50/$15, 1M ctx, text+image). NEW Responses-API features: programmatic tool calling (model-written JavaScript orchestrates tool calls in an isolated no-network V8 β 38β63.5% token reductions claimed), native subagent spawning, explicit prompt-cache breakpoints (30-min minimum cache life; cache writes now billed 1.25Γ uncached input, reads keep the 90% discount: Sol $0.50 / Terra $0.25 / Luna $0.10 cached reads). OpenAI claims 54% token-efficiency gain vs prior gen. Benchmarks (AA direct-fetched + Vellum): AA Intelligence Index v4.1 Sol 59 / Terra 55 / Luna 51 (Fable-5 = 60, #1); Coding Agent Index Sol 80 #1 / Terra 77.4 / Luna 74.6 (Fable-5 77.2); Terminal-Bench 2.1 Sol 88.8 SOTA (91.9 "Ultra") / Terra 87.4 / Luna 84.7 (Fable-5 86.0); Agents' Last Exam Sol 53.6 / Terra 50.4 / Luna 50.3 (Fable-5 40.5); MRCR Sol 91.5 / Terra 89.6 / Luna 41.3 β long-context disqualified; AA cost-per-task Sol $1.04 (~1/3 Fable-5) / Terra $0.55 / Luna $0.21. SWE-bench Pro: Sol 64.6 vs Fable-5's 80 β OpenAI disputes the benchmark's validity and omitted SWE-bench Verified/GPQA/AIME/MMLU/FrontierMath where Claude leads (selective reporting β weigh vendor-favorable numbers accordingly). METR caveat: multiple launch-coverage articles report METR found Sol's detected reward-hacking rate the HIGHEST of any public model it has evaluated β primary METR post unfetched this run (403s), re-confirm before load-bearing use. Omniscience: minor accuracy gain over gpt-5.5 but hallucination rate HIGHER β NOT a council-honesty candidate (grok-4.20 unaffected). No gpt-5.5 deprecation announced (explicitly checked), but gpt-5.5 is now price-dominated: Sol matches its $5/$30 with better agentic numbers; Terra β beats it at half price. Willison's hands-on: Sol competent but did not outperform Fable-5 on his complex coding tasks.
β¦grok-4.5 (SpaceXAI, GA 2026-07-08 devs / 07-09 public): the provider rebranded xAI β SpaceXAI at this launch (affects all xAI-provider rows + the access matrix β confirm the XAI_API_KEY/base-URL path is unchanged before relying on it). Pricing $2/$6, cached input $0.50; ~80 TPS. Built on the 1.5T-param "V9" foundation, trained alongside Cursor (available in Cursor all-plans + Grok Build + xAI console + OpenRouter). Benchmarks are AA-verified (Intelligence Index 54 #4, Coding Agent Index 76, GDPval-AA v2 1543 Elo #4) plus launch-coverage SWE numbers β RECORD BOTH sides of two disagreements: (1) DeepSWE provider-harness 62.0% vs neutral-harness 53%; (2) SWE-bench Pro lists Opus 4.8 at 69.2% (max) in xAI's chart vs this doc's roster 64.3% Pro β different harness/effort, don't overwrite. Context window SHRANK to 500K (grok-4.3 = 1M, grok-4.20 = 2M) β grok-4.5 is NOT a long-context pick. NO AA-Omniscience component published β not a council honesty seat (grok-4.20 keeps that slot). Modalities: OpenRouter lists text in/out; image/video NOT confirmed this run β do not route OCR/PDF here until confirmed. Not in the EU at launch (xAI expects mid-July). A 2026-06-29 "private beta only, no public access" report was superseded by the 07-08 public GA β pre-GA status flips fast (Fable-5 lesson), verify at GA.
βkimi-k3 (Moonshot, GA 2026-07-16 API-first; 2.8T-param sparse MoE, largest open-weight model ever announced; weights PROMISED by 2026-07-27, license unspecified at launch β do NOT assume K2.x's Modified-MIT carries over): new Kimi Delta Attention (KDA) + Attention Residuals (AttnRes) architecture; ONE reasoning-effort level ('max') β no cheap-effort tiers; automatic context caching (no cache IDs/TTL), cache-hit $0.30/MTok, Moonshot claims >90% hit rates in coding workloads. AA-verified (direct-fetched 07-18): Intelligence Index v4.1 = 57, #4 of 187 per AA's page β launch coverage says #3; on known v4.1 values (Fable-5 60 / Sol 59 / Terra 55) a 57 slots #3, so AA's #4 implies an unlisted β₯57 entry (likely a gpt-5.5 high/xhigh effort-tier row β unresolved, both recorded); cost-per-task $0.94 (Sol $1.04, Opus 4.8 $1.80); 62 TPS (#91 β slow-ish); AA logged it verbose (130M output tokens across its eval) despite the vendor's "21% fewer output tokens than K2.6" claim. Arena.ai frontend-code #1, ahead of Fable-5 (independent). Vendor-only coding suite (KimiCode harness @ max effort, no third-party reproduction): Terminal-Bench 2.1 88.3 (Sol's 88.8 SOTA stands β not harness-comparable), DeepSWE 67.5, ProgramBench 77.8, FrontierSWE 81.2, SWE Marathon 42.0 (would crush grok-4.5's 29.0 IF independently reproduced). NO SWE-bench Pro/Verified anywhere. Moonshot's own framing: "mostly beats Opus 4.8 max and GPT-5.5 high; trails Fable-5 and Sol on aggregate." Willison hands-on: pelican test burned 13,241 reasoning tokens for 3,417 output tokens (~$0.25/query) β max-only effort has real small-task overhead; his cited "+732 Elo vs K2.6" is implausible as written (fetch artifact), the 1547 Elo itself is indicative (cf. grok-4.5's GDPval-AA 1543). Council seat evaluated β DECLINED (no Omniscience number, capacity-limited upstream, open-weights/price is not a council criterion). OpenRouter live day-one but single-provider with a standing capacity warning ("may return frequent 429 errors") until weights land. Sources fetched this run: openrouter.ai/moonshotai/kimi-k3, artificialanalysis.ai/models/kimi-k3, simonwillison.net (07-16), mlq.ai.
By-domain scoring (1β10, vita-calibrated)
Scores blend published benchmarks with fit for vita workloads. See Β§ Methodology.
| Domain | claude-opus-4-8 | gpt-5.5 | gemini-3.1-pro | kimi-k2.6 | v4-pro | glm-5.2 | sonnet-4-6 | minimax-m2.7 | v4-flash | grok-4.20 | grok-4.3 | fable-5 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Agentic coding (SWE-bench Pro / Verified) | 10 (64.3% Pro) | 9 (58.6%) | 8 (54.2%) | 9 (58.6%) | 9 (80.6% Verified, no Pro) | 9 (62.1% Pro β top open-weights) | 7 | 8 (56.2%) | 8 (~79% Verified) | 6 | ? (no SWE-bench published) | 10 (80.3% Pro β new leader) |
| Long-horizon task completion | 10 | 9 | 8 | 10 (300 subagents, 4k steps) | 8 | 8 (FrontierSWE/SWE-Marathon wins, high/xhigh thinking) | 7 | 7 | 7 | 6 | 7 (16-agent Heavy carryover, unverified) | 10 (built for minutes-long autonomous runs) |
| General reasoning (AA Intel Index) | 9 | 10 (57) | 10 (57) | 9 (54) | 9 (52) | 7 | 7 | 7 (50) | 7 (47) | 6 (48) | 8 (53) | 10 (60 on v4.1, #1 β confirmed 07-10) |
| Vision / screenshots / PDFs | 9 (2576px OCR) | 8 | 10 | 7 | 5 (text only) | 6 | 8 | 6 | 5 (text only) | 7 | 8 (text+image, video native β first frontier with native video) | 10 (high-res; GDP.pdf 29.8 top) |
| Long context (>200k) | 8 (1M beta) | 8 (400k) | 10 (1M+ stable) | 9 (1M) | 9 (1M, 83.5% MRCR) | 9 (1M usable, IndexShare) | 7 | 7 (200k) | 9 (1M) | 9 (2M) | 8 (1M API-exposed; 2M claimed; tiered pricing past 200k) | 9 (1M default) |
| Cost-efficiency @ quality | 5 | 5 | 8 | 9 | 9 (10 during May-5 promo) | 9 ($1.40/$4.40; glm-5.1 keeps the 10) | 8 | 9 | 10 | 3 | 9 | 3 ($10/$50 β priciest GA frontier) |
| Non-hallucination / honesty | 8 | 8 | 8 | 7 | 7 | 7 | 8 | 7 | 7 | 10 (78% Omniscience) | ? (component score not yet published β pilot decision blocker) | 8 (no Omniscience published) |
| Writing / voice fidelity | 10 | 9 | 8 | 7 | 6 | 7 | 9 | 7 | 6 | 7 | 6 (verbose by default β 88M tokens during AA eval) | 10 (clearer/warmer per launch notes) |
| Tool use stability | 10 | 9 | 8 | 9 | 8 | 9 (MCP-Atlas 77.0, near Opus 4.8) | 9 | 7 | 7 | 7 | 7 (assumed carryover, unverified) | 10 (SOTA agentic) |
| Streaming speed (TPS) | 7 | 7 | 9 | 9 | 4 (~34 TPS) | 7 (TTFT ~10.8s, thinking; async-friendly) | 9 | 10 | 8 | 7 | 10 (206.9 TPS, #1; TTFT 12.65s due to extended thinking) | 6 (long deliberate turns; minutes-long single requests) |
Blank cells = insufficient public data, scored conservatively by inference.
New 2026-07-05 entrants not yet in the grid (deliberate):
claude-sonnet-5,minimax-m3, andkimi-k2.7-codeare in the roster above but are NOT scored in this 1-10 table this refresh β none has verified third-party benchmarks yet (Sonnet 5: qualitative Anthropic claims only; M3: no AA numbers pulled; K2.7-code: Moonshot-proprietary suites only, no independent SWE-bench). Scoring them now would be fabrication. They're tracked as pilots/successors in the roster + routing sections; pull AA Intelligence Index + SWE-bench for all three next refresh, then add columns. (Same discipline the doc applied to glm-5.2's held general-reasoning score.)β AA Intelligence Index rescale (2026-07-10): artificialanalysis.ai moved to Index v4.1 with the gpt-5.6 launch. The parenthetical AA numbers in the grid above (gpt-5.5 = 57, kimi = 54, etc.) are OLD-index values and are NOT comparable to v4.1 values (Fable-5 60 / Sol 59 / Terra 55 / Luna 51). Do not mix scales when scoring; re-pull v4.1 values for the incumbent columns at the next full grid rebuild.
gpt-5.6 family domain scores (new 2026-07-10 β SCORED: AA-verified third-party benchmarks, direct-fetched): Sol β agentic coding 9 (Coding Agent Index 80 #1, Terminal-Bench 2.1 88.8 SOTA; but SWE-bench Pro 64.6, ~15pts behind Fable-5); long-horizon 9 (Agents' Last Exam 53.6 top, OSWorld 62.6 with 85% fewer output tokens than Opus 4.8, native subagent spawning); general reasoning 10 (v4.1: 59, one below Fable-5); vision 8? (image input confirmed,
detail: originaloption new β no OCR benchmarks pulled); long context 9 (1M, MRCR 91.5); cost-eff 7 ($1.04/AA-task β 1/3 Fable-5, but $5/$30 sticker); honesty 7? (Omniscience: accuracy up vs 5.5 but hallucination rate UP β plus the METR reward-hacking caveat); writing/voice 8 (AA-Briefcase 2nd to Fable-5); tool use 9 (programmatic tool calling is genuinely new β unproven in vita's long-session shape); streaming ? (no TPS pulled). Terra β same shape one notch down (v4.1 55, Coding Agent 77.4, Terminal-Bench 87.4, MRCR 89.6) at half price β best value in the family. Luna β coding 8 (Terminal-Bench 84.7 > grok-4.5), reasoning 8 (v4.1 51), long context 3 (MRCR 41.3 β collapsed; hard disqualifier), cost-eff 8 ($0.21/AA-task β but 3-5Γ m2.7's sticker, ~20Γ v4-flash output; NOT a budget-band replacement). Fold into full grid columns at the next rebuild alongside the v4.1 re-pull.grok-4.5 domain scores (new 2026-07-09 β SCORED, unlike the 07-05 entrants, because it has AA-verified third-party benchmarks): Agentic coding 9 (SWE-bench Pro 64.7 β Opus 4.8; Terminal Bench 2.1 83.3 > Opus; Coding Agent Index 76). Long-horizon 9 (SWE Marathon 29.0 > Opus 4.8's 26.0; Cursor-trained agentic). General reasoning 9 (AA Intel Index 54, #4). Vision 7? (modality unconfirmed β OpenRouter lists text-only this run; do NOT route OCR/PDF here until confirmed). Long context 7 (500K usable β a REGRESSION vs the 1M tier). Cost-efficiency @ quality 9 (headline strength β $2/$6 + ~2Γ token efficiency β effective cost-per-task well under Opus/GPT-5.5). Non-hallucination ? (no Omniscience published β conservative, don't assume). Writing/voice 7 (not a voice model). Tool use 8 (native FC/JSON, Cursor-trained; not yet vita-stress-tested for long-session partial-failure shape). Streaming 7 (~80 TPS β mid; grok-4.3 is 207). Fold into a full grid column next refresh once numbers settle.
kimi-k3 domain scores (new 2026-07-18 β PARTIALLY scored: AA-verified reasoning/price only; coding deliberately unscored): general reasoning 9 (v4.1 = 57, #3-4 β above Terra 55, below Sol 59; first open-weight model at this tier); agentic coding ? (ALL SWE-style numbers are KimiCode-harness vendor claims β no independent SWE-bench Pro/Verified; scoring now would be fabrication, same discipline as the 07-05 entrants); long-horizon ? (SWE Marathon 42.0 vendor claim would be #1 by a wide margin β unverified); frontend/web code 10 (Arena.ai #1 ahead of Fable-5 β independent, and the one lane where K3 has a verified lead); long context 9? (1M nominal, no MRCR-style depth test pulled); vision 7? (text+image in per AA; zero OCR benchmarks); cost-efficiency 7 ($0.94/AA-task is frontier-cheap, but $15/MTok output + max-only reasoning effort = real overhead on small tasks β ~$0.25 for a trivial SVG); honesty ? (no Omniscience); writing/voice ? (untested); streaming 5 (62 TPS #91, plus upstream 429s). Full grid column only after weights land + at least one harness-comparable coding number exists.
claude-opus-5 domain scores (new 2026-07-25 β SCORED: AA-verified third-party benchmarks, direct-fetched): general reasoning 10 (AA Intelligence Index v4.1 = 61, narrowly #1 of all models β one above Fable-5's 60); agentic coding 10 (joint-first AA Coding Agent Index @ xhigh in Claude Code, highest SWE-Atlas-QnA β BUT no independent SWE-bench Pro number pulled; the "10" rests on AA Coding-Agent + GDPval, flag if a Pro number lands lower); long-horizon 10 (record GDPval-AA v2 1861 + AA-Briefcase 1720; OSWorld-2.0 beats Fable-5 at ~1/3 cost β the successor to opus-4-8's 10); vision/OCR 9? (image input confirmed, no OCR benchmark pulled β held at opus-4-8's 9); long context 8 (1M nominal but secondary-source + no MRCR pulled β held at opus-4-8's 8 pending model-card confirmation); cost-efficiency 7 (top intelligence at Opus $5/$25 AND ~26% cheaper-per-task than Fable-5 β the new value-frontier; up from opus-4-8's 5, but still Opus-tier sticker so not a cheap-band pick); non-hallucination 8? (no Omniscience pulled; lowest behavioral-audit score of any recent Claude = best-behaved, but that is NOT an Omniscience number β held at 8); writing/voice 10 (Opus lineage, "thoughtful/proactive" per launch); tool use 10 (SOTA agentic, joint-#1 coding agent); streaming 7? (no TPS pulled; 2.5Γ fast-mode available). Fold into a full grid column next refresh alongside an independent SWE-bench-Pro number + AA Omniscience + TPS.
Current vita routing (.aloop/config.json)
Audit as of 2026-04-23. Each mode pins a model for a task class; unpinned modes fall to the default (currently Claude via scripts/inference/claude_backend.py).
| Mode | Current model | Task type | Status |
|---|---|---|---|
silicon-zack / silicon-zack-chat |
silicon-zack (Modal) |
Identity/voice fine-tune | β correct β custom weights |
epd47_generate, media_deep_dive |
anthropic/claude-sonnet-4-6 |
Creative + analytical writing | β bumped from sonnet-4 (2026-04-23) |
epd47_explore, research_analyze, research_extract, research_synthesize, background, deep_analysis, synthesis, speculative, research |
deepseek/deepseek-v4-flash |
Research/synthesis | β swapped from M2.7 (V4-Flash bake-off winner) |
daily_rollup, media_score, media_summarize, monitoring, trend_analysis, view_rebuild, micro_consolidation, rss_process |
google/gemini-3.1-flash-lite-preview |
High-volume triage | β correct β cheap, fast, good-enough |
gemini-flash-test |
google/gemini-3.1-flash-lite-preview |
Test mode | β correct |
sleep_build |
anthropic/claude-opus-4-7 |
Autonomous sleep build phase | π DELIBERATE HOLD β auto-marked 2026-07-18 per the 07-10 self-terminating rule. The opus-4-7β4-8 bump was recommended at 6 consecutive reviews (06-17 β 07-10) and never applied; silence = declined by inaction. Same pin also live at scripts/cadence/daemon.py:258 + scripts/reports/generate_html_v2.py:54-55. No longer re-surfaced in review emails. Reaffirm quarterly (~Oct 2026); one tap overturns. (2026-07-25: the eventual target is now opus-5, not opus-4-8 β still same $5/$25; hold stands, NOT re-surfaced.) |
(unpinned: executive_reflect, sleep, board_meeting, coaching, doctor, voice, telegram, etc.) |
fall-through Claude | Real-time + high-stakes | β correct β Claude is the right default |
council seats (flows/council/setup_models.py) |
opus-4-8 / gpt-5.5 / gemini-3.1-pro / grok-4.20 | Multi-model deliberation | πΆ PROPOSED 07-25: Claude seat + synthesis β opus-5 (same price, AA #1, supersedes Fable-5-synth pilot). grok-4.3 shadow-seat pilot concluded. gpt-5.5βsol still open (gated on METR). |
simple-council, council-review, deep-council, research/codebase-research, vision-strategy flows |
opus-4-8 / gpt-5.5 / gemini-3.1-pro / grok-4.20 | Peer-model fan-out | πΆ sync Claude seat β opus-5 with the above |
scripts/council/core.py COUNCIL_DEFAULT_MODELS |
opus-4-7 / gpt-5.5 / gemini-3.1-pro / grok-4.20 | Programmatic council | β οΈ DRIFT: core.py:22 pins opus-4-7, not opus-4-8 (audit 07-25). Reconcile straight to opus-5. |
Audit snapshot: data/model-guidance/routing-audit-2026-04-23.json (23 pinned modes, 0 stale; 208 hardcoded refs scanned, 28 residual flags β all historical cost-tracking tables or containment-test pins, not live routing).
Sonnet 4.6 swap decisions (2026-04-23)
Sonnet 4.6 was being used heavily in both agentic and non-agentic paths. Audited every use site and swapped the non-voice-critical ones to cheaper Chinese models of equal-or-better benchmark performance on those specific workloads. Voice-critical paths stayed on Sonnet.
| Site | From | To | Rationale |
|---|---|---|---|
flows/code-review/synthesize |
sonnet-4-6 | z-ai/glm-5.1 | Code reasoning + structured output; GLM-5.1 beats Sonnet on SWE-bench Pro (58.4 vs ~55), ~1/3 cost |
flows/council-plan/council |
sonnet-4-6 | z-ai/glm-5.1 | Architectural review of implementation proposals; pure code-reasoning |
flows/docs-audit/review |
sonnet-4-6 | moonshotai/kimi-k2.6 | 1M context lets it read more docs at once |
flows/council-review/synthesize |
sonnet-4-6 | moonshotai/kimi-k2.6 | Long-horizon synthesis across multiple model outputs |
flows/podcast-deep/quality_gate.py |
sonnet-4-6 | minimax/minimax-m2.7 | Structured yes/no judging with rubric; M2.7 adequate, cheap |
flows/vdb/generate_newsletter.py |
sonnet-4-6 | moonshotai/kimi-k2.6 | Long-form data-driven newsletter |
scripts/memory/consolidator.py |
sonnet-4-6 | moonshotai/kimi-k2.6 | Weeklyβmonthly rollup, long context |
scripts/feed/narrative.py |
sonnet-4-6 | z-ai/glm-5.1 | Structured editor pass over timeline entries |
scripts/feed/privacy.py |
sonnet-4-6 | z-ai/glm-5.1 | Structured editor pass for privacy scrubbing |
api/src/services/improve_cycle.py |
sonnet-4-6 | z-ai/glm-5.1 | Improvement proposal generation = code/system reasoning |
gumball/flows/field-guide/segments/signal-check/FLOW.yaml |
sonnet-4-6 | moonshotai/kimi-k2.6 | PILOT β 1 segment only; expand to other segments if quality judgment is positive |
Kept on Sonnet 4.6 (voice-critical):
- flows/podcast-deep/scratchpad β creative brainstorm for the podcast host voice
- flows/podcast-quick/transcript β actual transcript output
- scripts/coaching/generator.py β coaching messages to bio-zack
- scripts/podcast/generate_transcript.py β podcast scripting output
- 6 of 7 gumball/flows/field-guide/segments/* + field-guide-review.flow.yaml β hold until signal-check pilot judged
Estimated monthly savings: ~$20-40 at current volumes. Real value is concentrating Sonnet "quality budget" on voice surfaces where drift is hardest to measure and easiest to miss.
Observability: token usage by model is already logged to data/tokens/YYYY-MM.json and data/inference/usage.jsonl. Watch for quality regressions via (a) bio-zack reports, (b) flows/podcast-deep/quality_gate.py failure rate (now self-judging), (c) /improve proposal quality trend.
Residual refs (intentional, do not touch)
scripts/inference/costs.py,cost_report.py,usage.pyβ keep historical model IDs (claude-opus-4-5-20251101,gpt-5.4, etc.) for retrospective cost computation on token data. New IDs added alongside, not replacing.scripts/tokens/codex_extractor.pyβ parses old codex logs that referencegpt-5.4; intentional.scripts/telegram_bot.pyβminimax-m2.5-faston SambaNova (may not have M2.7-fast hosting yet; bump when available).flows/containment-*/FLOW.yamlβkimi-k2-0905,kimi-k2.5pinned for reproducibility in security boundary tests.
By-task recommendations
Use this as a decision tree when writing a new script, a new stepwise step, or a new .aloop mode.
Real-time chat (telegram, voice, silicon-zack sessions)
β claude-sonnet-4-6 default; escalate to claude-opus-4-8 if the turn involves coding or a multi-step plan. Do not route to OpenAI from the chat path β voice fidelity is tuned to Claude. Sonnet 5 pilot (2026-07-05): do NOT swap chat/voice/composer paths to sonnet-5 yet β pilot it first on a non-voice analytical path (media_deep_dive), read quality, and account for the ~1.0-1.35Γ tokenizer inflation in cost before touching any voice-critical surface.
Proactive composers (scripts/proactive/builders.py, commitment_followup, morning_kickoff)
β claude-sonnet-4-6. Pricing matters here β these fire dozens of times/day. Avoid Opus unless a specific composer has shown quality issues.
Executive loop / sleep reflect / identity work
β claude-sonnet-4-6 for reflection passes (cheaper, runs often), claude-opus-4-8 for the sleep build step and any autonomous-edit agent (VITA #1 issue β persistence quality matters most here).
Coding agents (sleep build, stepwise agent executors, /improve develop)
β claude-opus-5 (2026-07-24) is the new default benchmark leader β same $5/$25 as opus-4-8, joint-#1 AA Coding Agent Index, ~26% cheaper-per-task than Fable-5, ZDR-clean (wire the opus-4-8 cyber-refusal fallback). It supersedes claude-opus-4-8, which stays available as the fallback and at any not-yet-swapped pin. Use kimi-k2.6 when cost matters and the task is bounded. deepseek-v4-pro is a strong batch-mode alternative β slow output (34 TPS) so async only. glm-5.2 is now the strongest open-weight coding option (SWE-bench Pro 62.1, top open-weights, MIT) for non-sensitive code β usable 1M ctx, high/xhigh thinking, but async-friendly only (TTFT ~10.8s); glm-5.1 stays as the ~30%-cheaper fallback for high-volume non-coding structured passes (feed editors). For the hardest, highest-value builds where correctness dominates cost, claude-fable-5 now leads (SWE-bench Pro 80.3 vs 4.8's ~69, and its lead widens as tasks get harder) β but at 2Γ Opus output cost and with mandatory refusal/fallback handling, it's a deliberate per-task escalation, not a default. Pilot candidate for sleep_build once the opus-4-7β4-8 bump settles. kimi-k2.7-code (2026-07-05): coding-specialized k2.6 successor, ~30% fewer thinking tokens and cheaper output ($0.74/$3.50) β but 256K ctx and vendor-only benchmarks. RETAIN k2.6 on all its current sites (they're 1M-context synthesis roles k2.7-code can't serve); rostered for future bounded coding agents where β€256K suffices β no current vita site is a clean swap target. grok-4.5 (2026-07-09): cheaper-frontier coding/agentic model whose value prop is cost-per-task β $2/$6 with ~2Γ token efficiency (4.2Γ fewer output tokens than Opus 4.8 on SWE-bench Pro), AA Coding Agent Index 76, and it beats Opus 4.8 on Terminal Bench 2.1 + SWE Marathon while trailing on SWE-bench Pro. Pilot only, NOT auto-swapped: run it as a challenger on ONE bounded coding path (candidate: flows/code-review/synthesize, currently deepseek-v4-pro) and measure effective cost-per-task WITH the ~2Γ token efficiency applied β per-token rates alone understate it. Keep it OFF any >500K-context role (500K ctx regression) and OFF council (no Omniscience number). Proprietary (not open-weight), so it competes with opus-4-8/gpt-5.5, not glm/kimi. gpt-5.6-luna (2026-07-10): add as a THIRD arm to that same bounded pilot β $1/$6, Terminal-Bench 2.1 84.7 (above grok-4.5's 83.3 and Opus 4.8), Coding Agent Index 74.6, AA-verified. One pilot, three challengers (v4-pro incumbent vs grok-4.5 vs luna), one cost-per-task readout β beats running two separate pilots. Luna's hard constraint: MRCR 41.3% β the task must fit comfortably under ~100K context. kimi-k3 (2026-07-18): WATCH β no coding routing yet. Frontier-tier open-weight (AA v4.1 57) whose only independently-verified coding signal is Arena.ai frontend-code #1, ahead of Fable-5 β so the natural future pilot is a bounded FRONTEND-generation path (the HTML-report implementer in scripts/reports/generate_html_v2.py, or /frontend-slides), NOT the SWE-style bounded-coding pilot lane (where all its numbers are vendor-harness). Three gates before any pilot: (1) weights + license land (~Jul-27), (2) OpenRouter capacity stabilizes (frequent 429s today), (3) at least one harness-comparable independent coding number. At $3/$15 it competes with sonnet/terra/grok-4.5 β not with the glm/kimi cheap band, and it does NOT touch k2.6's slots.
Research, synthesis, summarization (research_*, synthesis, deep_analysis modes)
β minimax-m2.7 is the incumbent (Intel Index 50, $0.30/$1.20). deepseek-v4-flash is the active challenger ($0.14/$0.28, Intel Index 47, 1M context) β bake-off in progress on flows/podcast-deep/quality_gate.py and scripts/media/synthesis.py. If quality holds, fan out to the full M2.7 tier. minimax-m3 (2026-07-05): the direct m2.7 successor at the SAME $0.30/$1.20 but with 1M ctx + native multimodal β add it as a third bake-off arm and pilot it on the equal-price m2.7 hardcoded sites (scripts/speculative/{engine,storage,budget}.py); it's a capability upgrade at zero price delta over m2.7 (though still pricier than v4-flash). Verify the OpenRouter rate is permanent (not promo) first. gpt-5.6-luna is NOT a budget-band candidate (2026-07-10): at $1/$6 it's 3β5Γ m2.7's sticker and ~20Γ v4-flash's output rate β its niche is cheap frontier-agentic coding (see Coding agents), not high-volume synthesis; and its MRCR 41.3% disqualifies the long-context synthesis roles this band leans on. kimi-k3 likewise is NOT a candidate here (2026-07-18): $3/$15 is 10Γ m2.7's sticker β its lane is frontier-value, and k2.6 keeps ALL the 1M-ctx strong-cheap synthesis slots.
High-volume triage (media scoring, daily rollup, trend analysis)
β gemini-3.1-flash-lite stays. Nothing touches its $/M.
Vision / OCR / screenshot analysis (emails with receipts, Libre CGM screenshots, health screenshots)
β claude-opus-4-8 for text-in-image OR gemini-3.1-pro for PDFs + diagrams. Flash-lite is not good enough here.
Council / delphi / tribunal (high-stakes decisions)
β Four-seat fan-out: claude-opus-4-8, gpt-5.5, gemini-3.1-pro, grok-4.20 (for the non-hallucination pressure). Synthesize with claude-opus-4-8. Reserve gpt-5.5-pro for rare, bio-zack-tagged critical calls. PROPOSED 2026-07-25 (not yet applied): bump the Claude deliberation seat AND the synthesis model opus-4-8 β claude-opus-5 β same $5/$25, #1 on AA Intelligence Index (61 v4.1, above Fable-5's 60), joint-#1 Coding Agent Index, record GDPval/AA-Briefcase. This SUPERSEDES the 06-17 Fable-5-synthesis pilot proposal: opus-5 gets you Fable-5-or-better synthesis quality at HALF the cost, is ZDR-clean (no 30-day-retention block), and carries no export-control availability risk β so it's a cleaner synthesis upgrade than Fable-5 on every axis except raw verbose deliberation. Wire the opus-4-8 cyber-refusal fallback (health/doctor-adjacent councils were already OFF Fable-5 for classifier risk; opus-5's classifiers fire ~85% less but keep the fallback). Sites: flows/council/setup_models.py, scripts/council/core.py (audit-flagged: core.py:22 still pins opus-4-7, not opus-4-8 as this doc claimed β reconcile to opus-5), + the 5 downstream flows (simple-council, council-review, deep-council, research/codebase-research, vision-strategy). One change at a time: land the Claudeβopus-5 seat swap independently of the still-open gpt-5.5βgpt-5.6-sol OpenAI-seat proposal below. PROPOSED 2026-07-10 (not yet applied): bump the OpenAI seat gpt-5.5 β gpt-5.6-sol β same $5/$30, successor generation, stronger agentic/reasoning numbers (v4.1 59 vs Fable-5's 60), GDPval β Fable-5. Sites: flows/council/setup_models.py, scripts/council/core.py, + the 5 downstream flows (simple-council, council-review, deep-council, research/codebase-research, vision-strategy). Two reasons this is a proposal and not an auto-swap: (a) the METR reward-hacking report is unconfirmed at primary, (b) Omniscience shows hallucination rate UP vs 5.5 β acceptable for the reasoning seat since grok-4.20 holds the honesty seat, but worth bio-zack's eyes. Keep gpt-5.5-pro as-is (no 5.6-pro exists). Grok-4.3 shadow-seat pilot concluded β see open questions for retirement decision status. grok-4.5 evaluated for a seat (2026-07-09) β DECLINED: it publishes no AA-Omniscience component, and grok-4.20's seat rationale is entirely its 78% Omniscience honesty lead β grok-4.5 adds cheaper capability, not a distinct honesty perspective, and would be redundant with the xAI slot. Keep the 4 seats; re-evaluate only if AA publishes an Omniscience number for it. kimi-k3 evaluated for a seat (2026-07-18) β DECLINED: same test β no Omniscience component published, upstream capacity-limited (frequent 429s β unfit for a reliability-sensitive fan-out), and its distinct angle (open weights, price) is not a council criterion. Re-evaluate only after weights land AND an honesty number exists. Fable-5 synthesis pilot (proposed 2026-06-17): the synthesis step is a single, low-volume, highest-stakes call β the cleanest place to spend Fable-5's 2Γ cost. Recommend piloting claude-fable-5 as the synthesis model only (keep the 4 deliberation seats), once council code carries refusalβopus-4-8 fallback. Health/doctor-adjacent councils stay OFF Fable-5 β bio classifier false-positive risk.
Embeddings / classifiers / small utility models
β Out of scope for this doc β tracked in scripts/vdb/ and scripts/inference/. Revisit when a frontier-embedding model ships.
Provider access matrix
| Provider | Env var | SDK | OpenRouter proxy? | Notes |
|---|---|---|---|---|
| Anthropic | ANTHROPIC_API_KEY |
anthropic python |
yes | Preferred direct; Claude daemon uses this. claude-opus-5 live day-one on Anthropic API ($5/$25, ZDR-clean, 2Γ fast-mode); OpenRouter availability unverified this run. opus-4-8 retained as the cyber-refusal fallback for opus-5/Fable-5 callers. |
| OpenAI | OPENAI_API_KEY |
openai |
yes | Direct for GPT-5.5-pro (response-API feature parity). gpt-5.6 family (gpt-5.6-sol/-terra/-luna; bare gpt-5.6βSol) live day-one direct + OpenRouter; programmatic tool calling is Responses-API-only β direct SDK required for that feature. NEW cache billing on 5.6+: writes 1.25Γ input, reads β90%, 30-min min life |
GEMINI_API_KEY |
google-genai |
yes | Flash-lite used heavily; direct preferred for cost tracking | |
| xAI / SpaceXAI | XAI_API_KEY |
OpenAI-compatible | yes | Provider rebranded xAI β SpaceXAI at the grok-4.5 launch (07-08) β confirm base-URL/key path unchanged. OpenRouter fine β low volume (council only). x-ai/grok-4.5 ($2/$6, 500K ctx, cached $0.50) live on OpenRouter; NOT in EU until ~mid-July. grok-4.3 has tiered pricing past 200k tokens; verify the higher rate before routing long-context jobs. |
| MiniMax | β | β | via OpenRouter only | minimax/minimax-m2.7 (incumbent); minimax/minimax-m3 (successor, same price, 1M ctx + multimodal β pilot) |
| Z.ai | β | β | via OpenRouter only | z-ai/glm-5.2 (current tier); z-ai/glm-5.1 retained as cheaper fallback. China data-jurisdiction risk on API path β non-sensitive code only, or self-host the MIT weights |
| Moonshot | β | β | via OpenRouter only | moonshotai/kimi-k2.6 (1M ctx, general β retained); moonshotai/kimi-k2.7-code (256K, coding-specialist β pilot); moonshotai/kimi-k3 ($3/$15, cache-hit $0.30, 1M ctx, text+image β WATCH: single-provider, upstream capacity-limited w/ frequent 429s, weights promised 07-27) |
| DeepSeek | DEEPSEEK_API_KEY |
OpenAI-compatible | yes | V4-Pro + V4-Flash live since 2026-04-24. OpenRouter ok; direct preferred for V4-Flash if volume warrants (saves 5% markup) |
OpenRouter key lives at OPENROUTER_API_KEY. All non-direct providers route through scripts/inference/factory.py.
Change log
Each refresh appends a row. Oldest-first β tail is freshest.
| Date | Change | Triggered by |
|---|---|---|
| 2026-04-23 | Initial doc. Opus 4.7 adopted for frontier-agent. Kimi K2.6 added to strong-cheap tier. M2.7 flagged as bump from M2.5. Deprecated Sonnet-4 references in .aloop/config.json. |
Manual build by bio-zack request |
| 2026-04-23 | Sonnet 4.6 β cheaper-Chinese swap pass (bio-zack directive): 5 stepwise flows + 5 scripts + 1 API service + 1 gumball pilot segment moved off Sonnet 4.6 onto GLM-5.1 (code-reasoning), Kimi K2.6 (long-horizon synthesis + 1M context), or MiniMax M2.7 (cheap structured judging). Voice-critical paths (coaching/generator, podcast transcripts, podcast-deep scratchpad, 6/7 gumball segments + field-guide-review) kept on Sonnet. See "Sonnet 4.6 swap decisions" section above for per-site rationale. | bio-zack "just do it all and ill watch" |
| 2026-04-23 | Applied all 6 recommendations: (1) .aloop research/background/deep_analysis/synthesis/speculative/research bumped M2.5βM2.7; (2) .aloop epd47_generate + media_deep_dive bumped sonnet-4βsonnet-4-6; (3) added .aloop mode sleep_build pinned to opus-4-7; (4) council roster bumped to opus-4-7/gpt-5.5/gemini-3.1-pro/grok-4.20 in flows/council/setup_models.py + scripts/council/core.py + 5 downstream flows (simple-council, council-review, deep-council, research-v2, vision-strategy); (5) swept 19 script + flow files onto current IDs (memory/consolidator, coaching/generator, cadence/daemon, feed/narrative, feed/privacy, podcast/generate_transcript, improve_cycle, vdb/generate_newsletter, speculative/{budget,storage,engine}, sentinel/poll_agent, media/{deep_analysis,synthesis}, podcast-panel-reactive); (6) DeepSeek V4 tracked in Open Questions. Also: tightened snapshot regex to capture full model strings (fewer false positives); added cost-table entries for opus-4-7. Aloop stale: 11β0. Hardcoded-stale true positives eliminated; 28 residual flags are intentional historical cost-tracking + containment-test pins. |
bio-zack "implement all suggested feedback" |
| 2026-04-23 | Automated audit: confirmed 11 stale aloop modes (8Γ m2.5βm2.7, 2Γ sonnet-4β4.6, 1Γ sleep unpin). 7 stale hardcoded sonnet-4 refs in scripts. Council seats (opus-4.6+gpt-5.4) and 5 Kimi K2 flow refs flagged for bump. 30+ audit false positives identified (pattern matcher too greedy on model-family substrings). No new models (research skipped). Full report: data/reports/2026-04-23-model-guidance-review.md. |
flows/model-guidance-review automated run |
| 2026-04-30 | DeepSeek V4 Pro + Flash adoption pass. Both shipped 2026-04-24 (MIT-licensed, 1M context, OpenRouter-available). V4-Pro added to Strong-cheap-reason tier (Intel Index 52, SWE-bench Verified 80.6%, Codeforces 3206 β frontier-tier reasoning at ~1/6th Opus 4.7 cost). V4-Flash added as Budget-MoE challenger to M2.7 (Intel Index 47, $0.14/$0.28 β exactly matches doc's prior projection). Domain scoring table extended with both. 75% promo on V4-Pro through 2026-05-05 β pilot recommendation: 2-flow V4-Flash bake-off vs M2.7 (podcast-deep/quality_gate.py, media/synthesis.py); add V4-Pro as 5th council seat through promo expiry; pilot V4-Pro on flows/code-review/synthesize (replace GLM-5.1). Audit: 0 stale aloop modes, 28 hardcoded refs all intentional residuals. Full report: data/reports/2026-04-30-model-guidance-review.md. |
flows/model-guidance-review automated run, trigger: new-model-detected |
| 2026-04-30 | Grok 4.3 review (later same day). xAI Grok 4.3 went GA today (beta since Apr 17), priced $1.25/$2.50 per MTok β 16x/24x cheaper than grok-4.20 β with Intel Index 53 (+5) and #1 output speed (206.9 TPS). xAI did not publish a model card, so SWE-bench / GPQA / AA-Omniscience are unavailable β and grok-4.20's roster slot is justified entirely by its 78% Omniscience honesty lead. Decision: add grok-4.3 to the roster as a "Reasoning-cheap (xAI)" tier with SHADOW-PILOT status; recommend a 5th council seat through ~2026-05-14 to gather side-by-side honesty signal. Hold on grok-4.20 retirement until AA publishes the Omniscience component for 4.3. Domain scoring table extended with conservative scores + explicit "?" markers where benchmarks are missing. Audit re-run found 0 stale aloop modes; 28 hardcoded refs remain (all intentional cost-tracking + containment-test residuals). Full report: data/reports/2026-04-30-model-guidance-review-grok-4-3.md. |
flows/model-guidance-review automated run, trigger: new-model-detected (x-ai/grok-4.3) |
| 2026-06-18 | GLM-5.2 lands β Strong-cheap-code tier bumped glm-5.1βglm-5.2. Z.ai/Zhipu GLM-5.2 (GA 2026-06-16, 753B open weights, MIT, $1.40/$4.40, 1M usable ctx via IndexShare, high/xhigh thinking) is the top open-weights coding model: SWE-bench Pro 62.1 (> GPT-5.5 58.6 > glm-5.1 58.4), FrontierSWE 74.4% / MCP-Atlas 77.0 near-tie Opus 4.8. Updated roster + domain scoring (coding 62.1% Pro; long-horizon 7β8; long-ctx 7β9; tool-use 8β9; cost-eff 10β9; streaming 9β7 for ~10.8s TTFT thinking latency). Routing proposals (NOT auto-applied): (1) bump 2 code-reasoning sites glm-5.1βglm-5.2 (flows/council-plan/FLOW.yaml:19, api/src/services/improve_cycle.py:68); (2) HOLD 2 high-volume feed-editor sites on glm-5.1 (scripts/feed/narrative.py:144, scripts/feed/privacy.py:170) β coding gains don't apply, 5.2 is ~30% pricier; (3) re-surface the still-unapplied 2026-06-17 opus-4-7β4-8 bump (sleep_build aloop + scripts/cadence/daemon.py:258; also scripts/reports/generate_html_v2.py:54-55); (4) close cost-tracking gap β add opus-4-8/fable-5/glm-5.2 rows to scripts/inference/costs.py + cost_report.py. Audit: 0 stale aloop modes; 26 hardcoded "stale" hits all documented intentional residuals (cost-history, containment-test pins, SambaNova-gated m2.5-fast TODO, reflect.md placeholder). Benchmarks from Z.ai launch coverage β NOT yet on artificialanalysis.ai; AA Intel Index unpublished; re-confirm next refresh. Full report: data/reports/2026-06-18-model-guidance-review.md. |
flows/model-guidance-review automated run, trigger: new-model-detected (glm-5.2) |
| 2026-06-17 | Claude Fable 5 lands β new Frontier-max tier. Added claude-fable-5 (Anthropic, GA 2026-06-09, $10/$50, 1M ctx, SWE-bench Pro 80.3 β new top reasoner) to roster + domain scoring table as a tier ABOVE opus-4-8; reserved for hardest builds + council synthesis, NOT a default (2Γ Opus output cost). Flagged Fable-5's integration requirements (adaptive-thinking-always-on, refusal/fallback opt-in, 30-day-retention/no-ZDR) and its sibling claude-mythos-5 (Glasswing-only, not adoptable). Routing proposals (NOT auto-applied): (1) pilot Fable-5 as council synthesis model; (2) bump live opus-4-7βopus-4-8 on aloop sleep_build + scripts/cadence/daemon.py:258; (3) add opus-4-8 + fable-5 pricing rows to scripts/inference/costs.py + cost_report.py. Audit: 0 stale aloop modes; hardcoded "stale" hits are documented intentional residuals (cost-history, containment pins, SambaNova-gated m2.5-fast TODO, reflect.md placeholder). Full report: data/reports/2026-06-17-model-guidance-review.md. |
flows/model-guidance-review automated run, trigger: new-model-detected (claude-fable-5) |
| 2026-06-23 | gpt-5.5-cyber evaluated β NOT adoptable; no roster change. OpenAI's cyber-specialist (full version ~2026-06-22; CyberGym 85.6 vs GPT-5.5 81.8, ExploitGym 39.5, SEC-bench Pro 69.8) is gated to verified "trusted defenders" under Daybreak β not generally available, not on public API/OpenRouter, no published pricing. Excluded for access, not capability (handled like claude-mythos-5: noted here, NOT given a roster/domain-table row; its evals are cyber-domain only). OpenAI points most orgs to GA "GPT-5.5 + Trusted Access for Cyber + Codex Security". Re-surfaced carried-forward proposals (NOT auto-applied): (1) the now-thrice-recommended opus-4-7βopus-4-8 bump (aloop sleep_build + scripts/cadence/daemon.py:258 + scripts/reports/generate_html_v2.py:54-55) β drop-in, price-neutral, still unapplied; (2) glm-5.1βglm-5.2 on the 2 code-reasoning sites (flows/council-plan/FLOW.yaml:19, api/src/services/improve_cycle.py:68), HOLD the 2 feed-editor sites on glm-5.1; (3) close the 3-review-old cost-tracking gap (add opus-4-8/fable-5/glm-5.2 rows to scripts/inference/costs.py + cost_report.py). Council roster unchanged (opus-4-8/gpt-5.5/gemini-3.1-pro/grok-4.20). Audit: 0 aloop stale (audit roster lacks opus-4-8/glm-5.2 so can't flag those pins β known blind spot); 26 hardcoded "stale" hits all documented residuals or carried-forward bump targets. Est. cost delta β +$1β3/mo. Broad sweep surfaced UNVERIFIED aggregator claims (GPT-5.6/Gemini-3.2/GLM-6; "Fable-5 pulled by export control") β not recorded. Full report: data/reports/2026-06-23-model-guidance-review.md. |
flows/model-guidance-review automated run, trigger: new-model-detected (gpt-5.5-cyber) |
| 2026-07-05 | Three new models rostered (challengers/pilots) + Fable-5 availability correction. Triggered on kimi-k2.7; broadened after research surfaced two higher-leverage models. (1) claude-sonnet-5 (GA 06-30, $2/$10 intro β $3/$15, perf near Opus 4.8) added as Balanced successor pilot β NOT auto-swapped: a new tokenizer inflates token counts ~1.0-1.35Γ (so "cost-neutral" is intro-only) and voice-critical paths need a drift check; pilot on media_deep_dive first. (2) minimax-m3 (GA 05-31, missed by recent refreshes) added as Budget-MoE successor β same $0.30/$1.20 as m2.7 but 1M ctx + native multimodal; fold into the budget bake-off / pilot on the speculative/* m2.7 sites. (3) kimi-k2.7-code (GA 06-12, the trigger) added as Strong-cheap-code pilot β ~30% fewer thinking tokens but 256K ctx and vendor-only benchmarks; RETAIN k2.6 (its 1M ctx serves the synthesis roles k2.7-code can't) β no current site is a clean swap. New entrants deliberately NOT scored in the 1-10 grid (no verified third-party benchmarks β would be fabrication). Fable-5 correction: the export-control suspension the 06-23 review dismissed as a rumor was REAL (suspended Jun 12-30, restored Jul 1) β slot now carries a geopolitical-availability caveat + needs opus-4-8 fallback for availability, not just refusal. Carried-forward (still unapplied): opus-4-7β4-8 bump now recommended 4Γ (sleep_build aloop + cadence/daemon.py:258 + generate_html_v2.py:54-55) β apply or mark deliberate hold; glm-5.1β5.2 on the 2 code-reasoning sites; the now-4-review-old cost-tracking gap (add opus-4-8/fable-5/glm-5.2/sonnet-5/m3/kimi-k2.7-code rows to costs.py+cost_report.py). Audit: 0 aloop stale by the audit's own roster (known blind spot β it lacks opus-4-8/glm-5.2/sonnet-5/m3/k2.7-code, so it can't flag the sleep_buildβopus-4-7 pin); 27 hardcoded "stale" hits are documented residuals or the carried-forward bump targets. ByteDance Seed 2.1 (Jun 24) + GPT-5.6/Gemini-3.5-Pro (unconfirmed) logged to Open Questions. Full report: data/reports/2026-07-05-model-guidance-review.md. |
flows/model-guidance-review automated run, trigger: new-model-detected (kimi-k2.7) |
| 2026-07-09 | grok-4.5 lands β new Frontier-agent-cheap pilot tier. SpaceXAI (rebranded from xAI) shipped grok-4.5 (GA 07-08 devs / 07-09 public, $2/$6, cached $0.50, 500K ctx, ~80 TPS, 1.5T "V9" foundation, Cursor-trained). AA-verified: Intelligence Index 54 (#4), Coding Agent Index 76, GDPval-AA v2 1543 Elo. Beats Opus 4.8 on Terminal Bench 2.1 (83.3) + SWE Marathon (29.0); trails on SWE-bench Pro (64.7). Headline = ~2Γ token efficiency (4.2Γ fewer output tokens than Opus 4.8 on SWE-bench Pro) β cost-per-task well under Opus/GPT-5.5. Added to roster + a scored domain mini-table (SCORED unlike the 07-05 entrants β it has verified benchmarks). NOT auto-swapped anywhere. Two disagreements recorded (DeepSWE provider vs neutral harness 62.0/53; Opus-4.8 SWE-bench-Pro 69.2 vs doc's 64.3). Caveats: 500K ctx REGRESSION (grok-4.3 = 1M), no Omniscience β council UNCHANGED (grok-4.20 keeps the honesty seat), modality text-only-confirmed, not-in-EU-til-mid-July, provider rebrand xAIβSpaceXAI (access matrix updated). Routing proposals (NOT auto-applied): (1) pilot grok-4.5 as a challenger on ONE bounded coding path (flows/code-review/synthesize) measuring cost-per-TASK; (2) opus-4-7β4-8 bump now 5Γ-recommended, DECISION FORCED β apply (sleep_build aloop + cadence/daemon.py:258 + generate_html_v2.py:54-55, audit-confirmed still opus-4-7) or mark a deliberate hold; (3) glm-5.1β5.2 on the 2 code-reasoning sites; (4) close the now-5-review-old cost-tracking gap (add opus-4-8/fable-5/glm-5.2/sonnet-5/m3/k2.7-code/grok-4.5 rows to costs.py+cost_report.py); (5) swap the reflect.md:348 placeholder 'kimi-k2.7,qwen-3.5'β'example-model-x' (real-model collision). Audit: aloop_stale empty (known blind spot β roster lacks opus-4-8/glm-5.2/grok-4.5); 210 hardcoded refs, "stale" hits all documented residuals/carried-forward targets; deep-council/FLOW.yaml:28 opus-4-8 flag is a FALSE POSITIVE (correct current ref). Full report: data/reports/2026-07-09-model-guidance-review.md. |
flows/model-guidance-review automated run, trigger: new-model-detected (grok-4.5) |
| 2026-07-10 | GPT-5.6 family lands (Sol/Terra/Luna) β three SCORED roster rows + council-seat proposal; GPT-Live = watch-only. OpenAI GA'd gpt-5.6 07-09 (preview since 06-26; post-US-Commerce-review β the Fable-5 pattern is now a standing pre-release regime): Sol $5/$30, Terra $2.50/$15, Luna $1/$6; all 1M ctx / 128K out / six effort levels / OpenRouter day-one; new Responses-API programmatic tool calling + subagent spawning + explicit cache breakpoints (writes 1.25Γ input). AA-verified (direct-fetched): Index v4.1 RESCALE β Fable-5 60 #1 / Sol 59 / Terra 55 / Luna 51 (old-index grid numbers NOT comparable; Fable-5 footnote corrected from "~65 single-source" β 60 v4.1 confirmed, partially closing that open question). Sol: #1 Coding Agent Index 80, SOTA Terminal-Bench 2.1 88.8, ALE 53.6 β but SWE-bench Pro 64.6 vs Fable-5's 80 (OpenAI disputes the bench + omits every bench Claude leads) and METR reportedly measured Sol's reward-hacking rate as the highest of any public model it has evaluated (primary unfetched β confirm). Omniscience: hallucination UP vs 5.5 β no council-honesty implications. Luna: Terminal-Bench 84.7 > grok-4.5, but MRCR 41.3 long-context collapse β never >100K ctx. gpt-live-1/-mini (07-08, full-duplex voice, delegates to frontier text model mid-conversation): ChatGPT-only, no API/pricing/OpenRouter β watch item, not rostered (gpt-5.5-cyber handling). No gpt-5.5 deprecation announced, but it's now price-dominated by its successors. Proposals (NOT auto-applied): (1) council OpenAI seat gpt-5.5βgpt-5.6-sol (same $5/$30); (2) opus-4-7β4-8 bump β 6th review, FINAL re-surface: unapplied by next review β auto-marked deliberate-hold; (3) Luna as third arm of the grok-4.5 bounded coding pilot (one pilot, three challengers); (4) fix the audit's frozen snapshot_routing.py:27 CURRENT_ROSTER (root cause of the perpetual "0 aloop stale" lying gauge); (5) cost-tracking gap now 10 models / 6 reviews (adds sol/terra/luna); (6) reflect.md:348 placeholder swap (carry-forward). Audit: aloop_stale [] (blind spot, see #4); 30 hardcoded flags all documented residuals; deep-council:28 opus-4-8 false-positive again. Full report: data/reports/2026-07-10-model-guidance-review.md. |
flows/model-guidance-review automated run, trigger: new-model-detected (gpt-5.6, gpt-live) |
| 2026-07-18 | Kimi K3 lands β rostered as first open-weight frontier-tier model (WATCH, zero routing changes); opus-4-7β4-8 DELIBERATE HOLD executed. Moonshot K3 (GA 07-16, 2.8T sparse MoE, KDA+AttnRes, $3/$15, cache-hit $0.30, 1M ctx, text+image, single 'max' effort) added to roster + partial domain mini-scores: AA v4.1 = 57 (#3 per coverage / #4 per AA page β discrepancy recorded), $0.94/AA-task, 62 TPS, Arena.ai frontend-code #1 ahead of Fable-5 (independent). Coding suite (Terminal-Bench 2.1 88.3 KimiCode-harness, FrontierSWE 81.2, SWE Marathon 42.0) vendor-only β coding domain deliberately UNSCORED. Pricing regime change: 4Γ k2.6 input β NOT a k2.6 successor β k2.6 retained on every strong-cheap site; K3 competes vs Terra/grok-4.5/Sonnet. Council seat DECLINED (no Omniscience, capacity-limited). Future pilot lane = bounded frontend-generation (generate_html_v2.py / /frontend-slides), gated on weights+license (~07-27), capacity (429s today), and one independent coding number. Executed the 07-10 self-terminating rule: the opus-4-7 pins (aloop sleep_build, cadence/daemon.py:258, generate_html_v2.py:54-55) auto-marked DELIBERATE HOLD β 6 re-surfaces, declined by inaction, reaffirm ~Oct 2026, stops being re-surfaced. Carried forward (4): glm-5.1β5.2 on 2 code-reasoning sites; snapshot_routing.py:27 frozen-roster fix (now missing 11 models incl. K3 β "0 aloop stale" still lies); cost-tracking gap (7 reviews, 11 models, K3 needs a cache-hit column); reflect.md:351 placeholder swap (elevated β this run WAS kimi-triggered). Audit: aloop_stale [] (blind spot unchanged); 30 hardcoded flags all documented residuals/carry-forwards; deep-council:28 false-positive again. Full report: data/reports/2026-07-18-model-guidance-review.md. |
flows/model-guidance-review automated run, trigger: new-model-detected (kimi-k3) |
| 2026-07-25 | Claude Opus 5 lands β new Frontier-agent LEADER; supersedes opus-4-8 at the same price. Anthropic Opus 5 (GA 07-24, claude-opus-5, $5/$25 β IDENTICAL to opus-4-8) added to roster as the new frontier-lead, with opus-4-8 reframed as superseded-but-retained (the documented cyber-refusal/availability fallback) and Fable-5 narrowed to a niche. AA-verified (direct-fetched): Intelligence Index v4.1 = 61 (max), narrowly #1 of all models (Fable-5 60 / Sol 59 / K3 57); joint-#1 Coding Agent Index (Claude Code @ xhigh); record GDPval-AA v2 1861 + AA-Briefcase 1720 (+146 vs Fable-5); ~26% lower cost-per-task than Fable-5 (at HIGH effort BEATS Fable-5 by 32 Elo at 47% of its cost). Anthropic: OSWorld-2.0 beats Fable-5 at ~1/3 cost, ARC-AGI-3 ~3Γ next-best, +22% hardest agentic coding vs prior Opus. ZDR-clean (no data-retention req, unlike Fable-5); cyber classifiers fire ~85% less than Fable-5 but still refuse offensive-security β opus-4-8 fallback; behind Mythos-5 (Glasswing-only) on cyber. 1M ctx secondary-source (official page stated price only). Added a SCORED domain mini-block (reasoning 10 / coding 10 / long-horizon 10 / cost-eff 7 β AA-verified). Proposals (NOT auto-applied): (1) council Claude deliberation seat + synthesis opus-4-8βopus-5 (setup_models.py + core.py + 5 flows) β SUPERSEDES the 06-17 Fable-5-synthesis pilot (opus-5 = better synthesis at half cost, ZDR-clean, no availability risk); (2) name opus-5 the coding-agent default (doc done; pins follow); (3) add opus-5 $5/$25 row to costs.py+cost_report.py (now 12 models / 8 reviews overdue); (4) fix frozen snapshot_routing.py:27 CURRENT_ROSTER (still opus-4-7, now missing 12 models incl opus-5 β the "0 aloop stale" gauge lies by omission). NEW audit finding: scripts/council/core.py:22 pins opus-4-7, NOT opus-4-8 as this doc claimed β doc/code drift, reconcile to opus-5. DELIBERATE HOLD RESPECTED: the opus-4-7 trio (sleep_build, cadence/daemon.py:258, generate_html_v2.py:54-55) stays held per the 07-18 self-terminating rule β opus-5 is just the new eventual target at the ~Oct-2026 reaffirm; NOT re-surfaced. Est. cost delta β $0 to β$5/mo (all concrete swaps price-neutral; Fable-5 narrowing is a slight savings). No SWE-bench Pro opus-5-vs-Fable-5 head-to-head + no AA Omniscience/TPS pulled β open for next refresh. Full report: data/reports/2026-07-25-model-guidance-review.md. |
flows/model-guidance-review automated run, trigger: new-model-detected (opus-5) |
Methodology β how to refresh
When to refresh:
- A new frontier model drops (GPT, Claude, Gemini majors)
- A new cheap-strong challenger lands (DeepSeek, Kimi, GLM, MiniMax, Qwen majors)
- Existing model prices change meaningfully (>20%)
- A vita workload shows quality issues traceable to model choice
- Quarterly, even with no new models, to catch drift
How the automated flow works (flows/model-guidance-review):
- Snapshot current routing. Reads
.aloop/config.json,flows/council/setup_models.py, and greps vita/gumball/cadence for hardcoded model IDs. Writesdata/model-guidance/routing-audit-<date>.json. - Research new models. Agent with Camofox + WebSearch pulls provider announcements, pricing pages (
artificialanalysis.ai,pricepertoken.com,llm-stats.com, vendor docs), and top independent benchmarks (SWE-bench Pro, AA Intelligence Index, Terminal-Bench 2, Omniscience). - Score by domain. Agent updates the domain table above using the published numbers + vita-calibrated weights (tool use stability and voice fidelity matter more than raw reasoning for our workloads).
- Generate diff. Agent proposes concrete changes to
.aloop/config.jsonand councilsetup_models.py. Proposals are diffs, not auto-applied. - Email bio-zack. HTML summary with: new models, recommendation summary, diff, and a link to this doc for deep-dive.
- Persist. Updates this doc (
docs/pages/model-guidance.md), appends to change log, and writes a signed report todata/reports/YYYY-MM-DD-model-guidance-review.md.
Manual refresh:
# from vita root
stepwise run flows/model-guidance-review --name model-review --wait
Sleep-triggered refresh:
The /sleep reflect step watches for new-model mentions in data/discoveries/*.json (from web_discover), in recent bio-zack telegram/voice mentions, and in data/reports/*. If any named model is not covered in this doc's roster, the refresh job is enqueued.
Calibration notes:
- vita-specific weights diverge from public benchmarks in three places: tool use stability (we run long agent sessions, partial-failure shape matters), voice fidelity (identity coherence is a first-class requirement), and streaming speed (telegram sessions are latency-sensitive).
- Non-hallucination matters disproportionately for the coaching + doctor agents β grok-4.20's Omniscience lead is why it gets a council seat.
- OpenRouter adds ~5% markup + latency; direct SDK preferred for high-volume modes.
- When in doubt: cheaper model first, escalate only on measured quality regression. VITA #1 issue means bad outputs persist β but so does over-spending.
Open questions
- Should we add Qwen 3.5 to the roster? Not yet evaluated; OpenRouter availability + benchmarks need a pass.
- Cerebras / Groq hosted inference: not yet part of vita's routing. Worth a pass if latency becomes a pain point on silicon-zack-chat.
- DeepSeek V4 post-promo decision (May 6+): at standard $1.74/$3.48, is V4-Pro worth a permanent council seat vs grok-4.20 ($20/$60)? Cheaper, but grok has the unique honesty/Omniscience profile. Decide based on observed contributions during the May-5 promo window.
- V4-Flash bake-off result (~2026-05-07): if quality holds across
podcast-deep/quality_gate.pyandmedia/synthesis.py, fan out to the full 9-mode M2.7 tier. Otherwise document the regression and stay on M2.7. - GPT-5.5-pro policy: currently "bio-zack tag only" β should we automate escalation on council when a question scores above some stakes threshold?
- gpt-5.5-cyber GA watch (new 2026-06-23): evaluated and excluded β gated to verified "trusted defenders" (Daybreak), not on public API/OpenRouter, no published pricing. Strong cyber-domain evals (CyberGym 85.6) but not vita-routable. Re-evaluate ONLY if OpenAI opens general access with pricing. Until then, any autonomous vuln-analysis/patching need falls to the existing coding tier (opus-4-8 / fable-5 / glm-5.2); orgs with vetted access use the GA "GPT-5.5 + Trusted Access for Cyber + Codex Security" path. Don't add a roster/domain-table row (handled like the non-adoptable
claude-mythos-5). - opus-4-7β4-8 bump β DELIBERATE HOLD EXECUTED 2026-07-18 (rule fired; no longer re-surfaced): still unapplied at this review, so per the self-terminating rule below the row is now marked "DELIBERATE HOLD β bio-zack declined by inaction, reaffirm quarterly (~Oct 2026)". The pins stay: aloop
sleep_build,scripts/cadence/daemon.py:258,scripts/reports/generate_html_v2.py:54-55β all opus-4-7, price-neutral to flip whenever he wants. Original context: recommended at 06-17, 06-18, 06-23, 07-05, 07-09, 07-10 and still unapplied β audit-confirmed still live at aloopsleep_build,scripts/cadence/daemon.py:258,scripts/reports/generate_html_v2.py:54-55(the 07-10 audit showsdaemon.py:258= opus-4-7,sleep_build= opus-4-7). Price-neutral ($5/$25 both). The 07-09 review forced a binary; no decision landed. Self-terminating rule, effective now: if still unapplied at the next review, that review auto-marks this row "DELIBERATE HOLD β bio-zack declined by inaction, reaffirm quarterly" and STOPS re-surfacing it. Silence has been the answer five times; the recite loop itself is now the bigger failure mode than the stale pin. One tap to apply, or do nothing and the system stops asking. - grok-4.5 coding pilot design (new 2026-07-09): grok-4.5's value is cost-per-TASK (not per-token) β a fair bake-off must measure effective cost WITH the ~2Γ token-efficiency applied, not the $2/$6 sticker. Candidate bounded path:
flows/code-review/synthesize(currently deepseek-v4-pro) as a challenger arm, or a bounded sleep-build sub-agent. Do NOT fan out until one clean side-by-side. Keep OFF council (no Omniscience number) and OFF >500K-context roles (500K ctx). Also unresolved: modality (OpenRouter lists text-only β confirm before any OCR/PDF routing) and whether the SpaceXAI rebrand changed theXAI_API_KEY/base-URL path. grok-4.5 does NOT replace grok-4.3 (4.3 keeps the 1M-ctx xAI reasoning-cheap slot). - Sonnet 5 pilot design (new 2026-07-05): sonnet-4-6 is vita's default chat/composer/voice model, so a wholesale swap is the highest-blast-radius change available β do NOT auto-apply. Recommended: pilot sonnet-5 on ONE non-voice analytical path (
media_deep_dive, which has a quality signal), track effective cost WITH the ~1.0-1.35Γ tokenizer inflation applied (not just per-token rates), and hold all voice-critical + composer paths on 4-6 until the pilot reads clean. Re-decide at standard pricing (Sep 1) since the intro discount evaporates then. Open sub-question: does the tokenizer change alter silicon-zack voice fine-tune assumptions? - MiniMax M3 vs the m2.7/v4-flash bake-off (new 2026-07-05): M3 is same price as m2.7 with strictly better specs (1M ctx, native multimodal) β arguably it should just replace m2.7 as the incumbent, making the live bake-off a 3-way (m3 / m2.7 / v4-flash). But benchmarks are unverified and the OpenRouter $0.30/$1.20 may be a launch promo (direct API is $0.60/$2.40 + a 512K-doubling band). Pilot on the
speculative/*m2.7 sites first; verify pricing permanence before any high-volume fan-out. - kimi-k2.7-code has no clean vita swap target (new 2026-07-05): the trigger model is coding-specialized with 256K ctx, but EVERY current kimi-k2.6 site (
docs-audit/review,council-review/synthesize,vdb/generate_newsletter.py,memory/consolidator.py, gumball signal-check) was chosen FOR the 1M window. So k2.7-code can't slot in anywhere today without a context regression. Rostered for future bounded coding agents; k2.6 retained. Re-evaluate if a general (non-Code) K2.7 ships with 1M ctx. (2026-07-18: kimi-k3 IS the general 1M-ctx successor β but at $3/$15 it's a different band entirely; the k2.6 slots stay put. The "cheap 1M Moonshot successor" this bullet waits for still doesn't exist.) - Fable-5 availability fallback (new 2026-07-05): the Jun 12-30 export-control suspension proved the Frontier-max slot can vanish by policy, not just refusal. Before ANY Fable-5 adoption (council-synthesis pilot, sleep_build), the caller must carry an opus-4-8 fallback that triggers on BOTH
stop_reason:refusalAND provider-unavailable/403. Also: mythos-5 remains Glasswing-only. Don't route to Fable-5 without the dual fallback. - ByteDance Seed 2.1 specs pull (new 2026-07-05): confirmed on the llm-stats tracker (GA 2026-06-24, Pro + Turbo variants) but no price/context/benchmark pulled this run β not on vita's roster. Pull specs next refresh and decide whether it's roster-relevant (likely a budget/mid-tier challenger). Do not adopt on zero data.
- Gemini 3.5 Pro (still unconfirmed; GPT-5.6 half RESOLVED 2026-07-10): the 07-05 row held both as unverified rumors β correct call twice over: GPT-5.6 turned out real (GA 07-09, now rostered β though at 1M ctx, not the rumored 1.5M) and the quarantine discipline cost nothing. Gemini 3.5 Pro remains UNCONFIRMED (not checked this run β trigger-scoped research). If real, it's a direct gemini-3.1-pro (Frontier-multimodal seat) bump β verify at next refresh.
- gpt-5.6-sol council-seat swap (new 2026-07-10, bio-zack's call): proposed gpt-5.5 β gpt-5.6-sol on the OpenAI council seat (same $5/$30; sites in the Council section). Gated on bio-zack's OK because of two open flags: (a) METR reward-hacking β launch coverage reports METR measured Sol's detected reward-hacking rate as the highest of any public model it has evaluated; the primary METR post 403'd this run β CONFIRM at metr.org before or shortly after the swap; (b) Omniscience hallucination rate UP vs 5.5 (fine for the reasoning seat while grok-4.20 holds honesty, but worth eyes). gpt-5.5-pro unchanged (no 5.6-pro). Terra note: if the seat swap lands and reads clean, Terra at $2.50/$15 is the natural candidate for any future cost-sensitive frontier-reason path β but one change at a time.
- GPT-Live API watch (new 2026-07-10): gpt-live-1/-mini (full-duplex voice, delegates to frontier text models mid-conversation, replaced Advanced Voice Mode in ChatGPT 07-08) is ChatGPT-only β API "planned", no pricing, no model IDs, not on OpenRouter. NOT rostered (gpt-5.5-cyber/mythos-5 handling: excluded for access, not capability). Vita relevance is real β this is the architecture a silicon-zack voice surface wants β so re-evaluate the moment OpenAI ships API access with pricing. Known launch wrinkles: Hindi translation demo quality issues; Willison's preview hit a laughing-interruption bug; the 75.5 "pleasantness" score is OpenAI-internal.
- Audit CURRENT_ROSTER is frozen β the "0 aloop stale" gauge lies by omission (root-caused 2026-07-10):
flows/model-guidance-review/snapshot_routing.py:27still listsclaude-opus-4-7as CURRENT and lacks opus-4-8, fable-5, glm-5.2, sonnet-5, minimax-m3, kimi-k2.7-code, grok-4.3, grok-4.5, the deepseek v4 family, and now the gpt-5.6 trio. Consequence:aloop_stalerenders empty every review while thesleep_buildβopus-4-7 pin β the exact thing the audit exists to catch β sails through unflagged, anddeep-council/FLOW.yaml:28(opus-4-8, CORRECT) false-positives every run. Four consecutive reviews have footnoted this as a "known blind spot"; it's a 10-minute fix to the set literal. Instrumentation-lying class: a passing check that cannot fail is not a check. - reflect.md placeholder collision (new 2026-07-05):
flows/sleep/prompts/reflect.md:348uses the literal examplenew_models = 'kimi-k2.7,qwen-3.5'as a placeholder β and "kimi-k2.7" is now a real (if -Code-suffixed) model. Low risk, but the placeholder string doubling as a real model name is exactly the kind of thing that can spuriously trigger new-model detection. Consider swapping the placeholder to an obviously-fake token (example-model-x). - Grok 4.3 vs 4.20 retirement decision (~2026-05-14): swap requires (a) AA publishes Omniscience component score β₯75% AND (b) shadow-seat outputs track 4.20 on at least 5 council runs. If both conditions hit, retire 4.20 (saves ~$15-20/mo on council fan-out). If either misses, hold and re-evaluate at next refresh.
- GPT-5.5 reasoning-effort tiers: AA leaderboard now lists
gpt-5.5 (high)= 59 andgpt-5.5 (xhigh)= 60 vs the standardgpt-5.5= 57 currently used in vita. Need to investigate whether OpenAI exposes a request-timereasoning_effortknob and whether council seats should pin to a specific tier. - grok-4.3 context window discrepancy: AA + OpenRouter list 1M; review articles say 2M carryover from 4.20. Treat 1M as the safe operating limit until xAI publishes a model card. Tiered pricing past 200k tokens β verify higher-rate before routing long-context jobs.
- Fable-5 council-synthesis pilot (bio-zack's call): swap the council synthesis model opus-4-8 β claude-fable-5? Highest-leverage adoption (one low-volume call on the highest-stakes step), but requires council code to carry a refusalβopus-4-8 fallback first. Decision gated on that integration + one side-by-side synthesis comparison.
- Fable-5 for
sleep_build: Fable-5 is purpose-built for long-horizon autonomous work (sleep_build's exact profile, VITA #1 persistence surface). Tempting, but 2Γ output cost on a nightly job + adaptive-thinking-always-on + refusal/fallback integration. Recommend: land the conservative opus-4-7β4-8 bump first, observe persistence quality, then consider a Fable-5 pilot. Don't leapfrog. - Fable-5 refusal risk on health/bio surfaces: do NOT route
doctor, health-coaching, or CGM/labs-analysis paths to Fable-5 without fallback β bio-adjacent content can trip its bio safety classifier and returnstop_reason:refusal. This is a sharp, model-specific caveat (Mythos 5 lacks the classifier but is Glasswing-only). Any Fable-5 caller needs the opt-in fallback wired before it sees personal-health content. - Fable-5 benchmark confirmation (PARTIALLY CLOSED 2026-07-10): AA Intelligence Index confirmed direct at artificialanalysis.ai β 60 on Index v4.1, #1 (the "~65" was an old-index single-source figure; scale changed, both were "right" on their own scale). SWE-bench Pro 80 re-confirmed incidentally via the gpt-5.6 comparison coverage. Still open: SWE-bench Verified ~95 remains single-source β low priority now that the Pro number and AA index are both grounded.
- Cost-tracking gap (now 6 reviews old, 2026-07-10 β 10 models missing):
scripts/inference/costs.py(audit-confirmed still only opus-4-5/opus-4-6/sonnet-4-5) +cost_report.pyhave no rows for opus-4-8 ($5/$25), fable-5 ($10/$50), glm-5.2 ($1.40/$4.40), sonnet-5 ($3/$15), minimax-m3 ($0.30/$1.20), kimi-k2.7-code ($0.74/$3.50), grok-4.5 ($2/$6), or now gpt-5.6-sol ($5/$30), gpt-5.6-terra ($2.50/$15), gpt-5.6-luna ($1/$6). New-tier spend currently mis-costs (falls back to nearest-tier estimate). Add all alongside the historical entries (don't replace). Note for the 5.6 rows: cache writes bill at 1.25Γ input β the cost model needs a cache-write column for 5.6+ models or it will under-report. (2026-07-18: add kimi-k3 $3/$15 + cache-hit $0.30 β now 11 models / 7 reviews; Moonshot's cache-hit discount, like the 5.6 cache-write multiplier, needs the cache-aware column.) - GLM-5.2 AA Intelligence Index unpublished: SWE-bench Pro 62.1 + the FrontierSWE/MCP-Atlas figures are from Z.ai launch coverage (VentureBeat); artificialanalysis.ai had no GLM-5.2 entry at refresh time and llm-stats listed no benchmark numbers. Re-confirm the coding scores + pull an AA Intel Index next refresh before treating as load-bearing. General-reasoning domain score held at 7 (= glm-5.1) pending AA data.
- GLM-5.2 feed-editor hold revisit:
scripts/feed/narrative.py:144+scripts/feed/privacy.py:170stay on glm-5.1 (cheaper, non-coding structured-editor passes; glm-5.2's gains are coding-specific). Revisit only on a measured quality regression β cheaper-first per methodology. - GLM-5.2 modality: OpenRouter lists glm-5.2 as text-in/out; vision domain score held at 6 (unchanged from 5.1) and it is NOT recommended for any OCR/screenshot/PDF path. Confirm modality before routing image work.
flows/code-review/synthesizedoc-drift: the 2026-04-23 "Sonnet 4.6 swap decisions" table still lists this site as glm-5.1, but it was moved todeepseek/deepseek-v4-proby the 2026-04-30 V4-Pro pilot (flows/code-review/FLOW.yaml:136). Historical table left as-dated; live routing is v4-pro. Not a glm-5.2 bump target.- kimi-k3 weights + license watch (new 2026-07-18): weights PROMISED by 07-27; license unspecified at launch (K2.x was Modified-MIT β don't assume it carries). Watch for: the actual license, independent SWE-bench Pro/Verified or common-harness Terminal-Bench numbers, an AA Omniscience component, OpenRouter capacity stabilizing (frequent 429s at launch), and second providers post-weights (which usually fixes both capacity and price). Re-run this review at the weights drop β the /sleep new-model watch should catch the coverage; if not, manual
stepwise run flows/model-guidance-review. - kimi-k3 AA-rank discrepancy (new 2026-07-18): launch coverage (Bloomberg et al) says K3 debuted #3 on the AA leaderboard, behind only Fable-5 and Sol; AA's own model page says #4 of 187. On the known v4.1 values (Fable-5 60 / Sol 59 / Terra 55) a 57 slots #3 β AA's #4 implies an unlisted β₯57 entry, most plausibly a gpt-5.5 high/xhigh effort-tier row (see the GPT-5.5 reasoning-effort question above β the two puzzles may be one puzzle). Also: Willison's cited "+732 Elo vs K2.6" is implausible as written (fetch artifact); the 1547 Elo itself is indicative (cf. grok-4.5's GDPval-AA 1543). Resolve both at next refresh.
- kimi-k3 frontend-generation pilot (new 2026-07-18, triple-gated): the one lane with an INDEPENDENT K3 #1 (Arena.ai frontend-code, ahead of Fable-5) maps directly onto vita's HTML-report pipeline β candidate sites: the implement stage of
scripts/reports/generate_html_v2.py(currently opus-4.7 under the deliberate hold, ~$3-7/report) or/frontend-slides. Run K3 as a challenger arm on a small report batch, judge with the pipeline's existing critique stage, measure cost-per-report. Gates: weights/license landed + capacity stable + one harness-comparable coding number. Do NOT fold into the grok-4.5/luna SWE-style bounded-coding pilot β different lane, different evidence base. - opus-5 council seat swap (new 2026-07-25, bio-zack's call): bump the Claude deliberation seat + synthesis model opus-4-8 β opus-5. Price-neutral, #1 AA, and it SUPERSEDES the still-open 06-17 Fable-5-synthesis pilot (opus-5 = better/equal synthesis at half Fable-5's cost, ZDR-clean, no export-control availability risk). Only real gate: wire the opus-4-8 cyber-refusal fallback into council code first (health/
doctorcouncils were already OFF Fable-5). Reconcile thecore.py:22opus-4-7 drift in the same edit. Independent of the gpt-5.5βsol OpenAI-seat question β do one seat at a time. - opus-5 vs Fable-5 β the Frontier-max slot may be collapsing (new 2026-07-25): opus-5 is #1 on AA Intelligence Index AND ~26% cheaper-per-task than Fable-5 AND ZDR-clean β so the "hardest builds" default arguably moves from Fable-5 to opus-5. What KEEPS Fable-5 alive: (a) no independent SWE-bench-Pro opus-5-vs-Fable-5 head-to-head was pulled (Fable-5's 80.3 Pro lead over opus-4-8's ~69 is real; whether opus-5 closes it is unknown), (b) Fable-5's verbose deliberate style may still win the rare max-correctness single call. Pull an independent SWE-bench-Pro opus-5 number next refresh; if it matches/beats Fable-5, retire the Frontier-max tier to a footnote. Until then Fable-5 stays as a narrow niche, not a default.
- opus-5 spec confirmations pending (new 2026-07-25): (a) context window β 1M is launch-coverage only; anthropic.com/news + the system-card blurb stated PRICE only this run. Confirm at the model card before any >200K-ctx routing. (b) OpenRouter availability β unverified this run (Anthropic API is day-one). (c) AA Omniscience + TPS/streaming not pulled β the domain mini-block holds honesty at 8 and streaming at 7? pending. (d) opus-5's ZDR-clean status makes it the preferred frontier model for any sensitive/health path where Fable-5's no-ZDR + refusal risk disqualified it β note when routing
doctor/CGM/labs analysis that needs frontier reasoning. - opus-5 for
sleep_build(new 2026-07-25 β HELD, not a re-surface): the held opus-4-7 trio's eventual target is now opus-5 (not opus-4-8), same $5/$25. Per the 07-18 self-terminating rule this is NOT being re-surfaced as an ask β it reaffirms ~Oct 2026, one tap to apply. Recorded here only so the target is correct when the hold next comes up. opus-5 (agentic-coding + computer-use leader, ZDR-clean) is a materially better sleep_build fit than either opus-4-7 or a Fable-5 pilot.