Current study & catalog
Exact endpoints, visible evidence.
The 16-endpoint real-call study appears first. A separate QwenCloud addendum records two real Qwen 3.8 Max arms and two successful Epicure calls, while keeping the mutable alias outside every scoring pool. Content-addressed provider snapshots support discovery; they do not establish availability, season eligibility, or model quality.
Current real-call endpoints
| Model | Canonical response identity | Provider route | Matched pairs | Evidence status |
|---|---|---|---|---|
| GPT-5.6 Sol (OR pro) | openai/gpt-5.6-sol-pro-20260709 | OpenRouter openai/flex | 13 / 20 | real calls unranked |
| Claude Fable 5 | anthropic/claude-5-fable-20260609 | OpenRouter anthropic | 17 / 20 | real calls unranked |
| Claude Opus 5 | anthropic/claude-opus-5-20260723 | OpenRouter anthropic | 19 / 20 | real calls unranked |
| Claude Sonnet 5 | anthropic/claude-sonnet-5-20260630 | OpenRouter anthropic | 18 / 20 | real calls unranked |
| Gemini 3.1 Pro | google/gemini-3.1-pro-preview-20260219 | OpenRouter google-ai-studio/flex | 14 / 20 | real calls unranked |
| Gemini 3.6 Flash | google/gemini-3.6-flash-20260721 | OpenRouter google-ai-studio/flex | 17 / 20 | real calls unranked |
| Grok 4.5 | x-ai/grok-4.5-20260708 | OpenRouter xai/zdr | 8 / 24 | real calls unranked |
| Kimi K3 | k3 | Kimi direct API kimi-code-direct | 16 / 20 | real calls unranked |
| GLM 5.2 | z-ai/glm-5.2-20260616 | OpenRouter deepinfra/fp4 | 10 / 20 | real calls unranked |
| DeepSeek V4 Pro | deepseek/deepseek-v4-pro-20260423 | OpenRouter cloudflare | 10 / 20 | real calls unranked |
| DeepSeek V4 Flash | deepseek/deepseek-v4-flash-20260731 | OpenRouter deepinfra/fp4 | 18 / 20 | real calls unranked |
| MiniMax M3 | minimax/minimax-m3-20260531 | OpenRouter morph | 9 / 24 | real calls unranked |
| Nemotron 3 Ultra | nvidia/nemotron-3-ultra-550b-a55b-20260604 | OpenRouter together | 14 / 20 | real calls unranked |
| Mistral Medium 3.5 | mistralai/mistral-medium-3.5-20260430 | OpenRouter mistral | 8 / 28 | real calls unranked |
| Command A Plus | command-a-plus-05-2026 | Cohere direct API cohere-direct | 12 / 12 | real calls unranked |
| Command A Reasoning | command-a-reasoning-08-2025 | Cohere direct API cohere-direct | 8 / 12 | real calls unranked |
These rows report gross provider-and-Epicure execution, not culinary quality. The task validity hold removes 32 pairs from the original collection, leaving 179; source-verified coverage recovery adds 7, while one predecessor pair remains on a formal policy hold, yielding 186 retained uplift pairs. No endpoint becomes rankable until admissible independent judgments are collected and a release snapshot is approved. A frozen coverage repair requires 25 new real arms. Its first execution recorded 9 arms and then stopped fail-closed with 9 of 13 endpoint-task cells unresolved. No complete-repair or ranking claim is made.
protocol comparison sha256: 2f99cec4e2e79029a81b8a601f76cf5bba5b237809e9d6d5981d7adc9153adb8Direct-Qwen evidence and candidates
A separate engineering smoke exercised the catalog-observed Qwen 3.8 Max mutable alias. Its successor run delivered two real arms, six provider responses, and two successful live Epicure calls. QwenCloud returned usage but no rate or charged amount, so cost is unknown rather than free; USD 2 remains held for the successor and USD 4 cumulatively across both attempts. The alias is not a frozen model release and has zero quality judgments. It enters neither Season 0 scoring pool, any quality figure, nor a ranking. The dated 3.7 releases and Qwen 3.5 397B remain unobserved candidates.
| Model | Identity | Compatibility | Observed arms | Ranking |
|---|
Discovered model catalog
0 visibleVerifying the frozen catalog snapshot and its acquisition receipt.