Development evidence · quality release held
One protocol for open battles. Another for controlled evaluation.
FlavourBench compares culinary model responses and measures the effect of offering Epicure access. The current 16-endpoint study used real provider calls and the deployed Epicure MCP. It defines a blinded review workload, but it does not yet contain an admissible human quality judgment and therefore does not qualify as a public model ranking.
Current study
Nine frozen real-call collections ran 16 exact endpoints over 28 human-authored development tasks on 2–3 August 2026. One of the four quarantined tasks also remains under a source-rights hold. The collections scheduled 320 matched Epicure-off/on pairs and completed 211, yielding 524 normalized real answers. The records preserve 1,883 provider generation identifiers and 488 live Epicure calls, of which 380 succeeded. No synthetic task, response, tool trace, or ballot enters this study.
The 211 gross completed pairs remain in operational denominators. Quarantining four held tasks removes 32 pairs, leaving a 179-pair base uplift pool; source-verified coverage recovery originally added eight pairs. A later pair-policy review placed one of those recovery pairs on hold, yielding 186 review pairs and 372 arms. The gross arena pool contained 1,024 same-task pairings from 218 Epicure-on answers. Quarantine left 876 pairings from 185 answers, and coverage recovery added 39 pairings and seven answer identities. The corrected workload therefore contains 915 pairings over 188 compared answers, plus four answers with no same-task peer: 192 Epicure-on answer identities in total. These reuse answers and are not independent observations. No preference model is fitted until eligible blinded judgments are released.
Post-freeze QwenCloud route check
On 8 August, a separate engineering smoke exercised the catalog-observed qwen3.8-max mutable alias. The successor run delivered 2 real response arms, 6 provider responses, and 2/2 successful live Epicure calls. Both final responses ended with stop.
QwenCloud returned token usage but no rate or charged amount. The USD 2 successor ceiling and USD 4 cumulative ceiling therefore remain conservative exposure, not provider cost. Recorded zero cost means unknown, not free. The alias is not a frozen release, has zero quality judgments, and enters neither the 186-pair uplift pool, the 915-comparison arena pool, any quality figure, nor any ranking.
Historical audit
A separate retrospective audit covers 120 public questions, 12 historical endpoints, and 2,880 attempted response arms. Four automated judge endpoints evaluated available pairs in both presentation orders.
After finish-reason correction, only 393 of 1,486 judgable comparisons reached the automated-panel consensus rule. That evidence is retained as a reliability and selective-survival audit. It is not pooled with the current study and is not a human culinary-quality leaderboard.
Why no ranking is published
The current study has zero admitted quality judgments. Its 915 retained candidate pairings reuse 188 compared answers, with four further unpaired answers, so neither the pairing count nor graph connectivity constitutes an effective judged sample. The post-quarantine audit had 94 empty model-pair-by-family cells; the corrected workload has 73 empty cells out of 480. A frozen coverage-repair schedule requires 25 new real arms across 13 endpoint-task cells. Its first governed execution stopped fail-closed after recording nine real arms from six source artifacts: four cells completed their required conditions, two finalized with one condition missing, one reservation produced no source artifact, and six cells were never started. Nine cells remain unresolved. Source reconstruction admits only qualifying recorded outputs and does not turn the incomplete repair into full family coverage. No family-specific or overall ranking follows from this partial run. The historical audit likewise has no independent human criterion cohort.
The public development questions did not complete qualified independent item review. Four current-study tasks are under an explicit validity hold for degraded context, self-resolution, construct ambiguity or drift, or missing visual information. They remain in operational denominators but cannot enter an official preference fit unchanged.
Reasoning configuration and missing sensitivity
Every exact-frontier development record explicitly requested low reasoning effort in both the intermediate tool-use phase and the final-answer phase. This is a material execution choice, not a provider default, and it limits comparisons among frontier endpoints.
Three governed attempts to qualify matched default- and high-effort sensitivity arms recovered no usable pair. The final route audit records four safe provider rejections, zero accepted generations, and no default- or high-effort arms. The current evidence therefore supports no claim that endpoint ordering is stable across reasoning effort.
Response-envelope route qualification
A new v4 gate qualified the response-envelope path using one real DeepSeek V4 Flash matched pair on the fixed DeepInfra route. A source-reconstructing verifier re-opened and re-hashed the immutable journal, provider generations, routing parameters, reasoning settings, cost records, and complete MCP trace. Both arms were usable; seven accepted provider generations and five successful Epicure calls were reconciled at an actual cost of USD 0.002899.
This is a route-smoke result only. It contains zero culinary judgments, contributes no ranking or uplift observation, and does not establish default- or high-effort sensitivity. The identifiers are closed against replay. The gate authorized materialization of the separate coverage-repair workload; it did not certify that workload as complete.
Epicure reconstruction boundary
The deployed Epicure runtime is bound to its application, tool-schema, and 1,790-item representation hashes. The version 3 packet verifies all 11 exact runtime-data files at Git commit 14ddf04aba81a76b75efa6554041f6bff48992c6, with bundle digest 98d0403115bf8eb4fb71dbb89a53362e9b9acafda7494c464763df7d764174d1. A downloadable reconstruction kit now fetches those immutable public inputs and the 41 hash-locked dependency wheels, then installs offline and runs fixed identity and tool probes. The clean execution receipt was produced by the study operator; it is not an independent reproduction or a public training-lineage release.
A source audit recovered a plausible Cooc training run, seed, job specification, and graph provenance, but rejected it as the exact ancestor because its recorded matrix statistics differ from the deployed payload. Training lineage, redistribution rights, a redistributable payload release, and independent reconstruction remain unresolved; the present Epicure condition is consequently development evidence, not a reproducible public treatment release.
The downloadable exact-source archive preserves all 28 application files and 116,668 source bytes used by the study. 21 files match the frozen public Git reference; the other 7 are preserved in the research archive. The source archive and reconstruction kit contain no data/model payload or credentials.
The packet also binds the dependency lock and SBOM, same-operator rebuild receipt, and rejected lineage candidate. Payload rights and upstream source rights are not attested; the exact immutable OCI image, private wheelhouse, original training input, complete corpus lineage, and independent reproduction receipt remain unavailable. Public-input reconstruction is technically specified and same-operator verified; redistribution, training-lineage recovery, an independently reproduced receipt, a signed immutable runtime release, and official rank eligibility remain blocked.
Two evaluation populations
Public Arena is designed to accept free-form culinary prompts and public votes. It remains disabled while the quality-release gates are open. Once released, its estimates will describe that changing prompt population.
Controlled Evaluation uses a sealed task release, fixed repetitions, fixed endpoint contracts, and a fixed rater plan. It produces a signed private run card for model and evaluation teams.
The service stores the data stratum on every battle and leaderboard snapshot. Public and controlled observations cannot enter the same estimate.
Controlled run contract
- An operator admits a model already present in the frozen season roster.
- Submitted endpoint, model-card, data-policy, analysis-plan, and rater-plan digests are recorded.
- The pre-run card binds the model roster, task schedule, Epicure release, prompts, decoding policy, data stratum, and spending cap.
- A content-addressed result snapshot reports preference intervals and operational measures separately.
- Results stay private unless the lab requests release and the common eligibility gates pass.
The present service does not yet onboard arbitrary customer endpoints. Execution fees do not alter scoring, eligibility, release review, or publication thresholds. The initial service will remain invite-only until endpoint onboarding, the sealed task bank, and the independent rater study pass their gates.
Next study
The frozen prospective campaign begins with 1,052 attributed, public, human-authored culinary questions from 756 source authors. Automated screening marked 418 provisional strict passes and 505 questions eligible for blind family review. These labels set review priority; they are not task validity judgments.
A fixed 180-candidate slate contains 45 scheduled questions per family. The target is 120 admitted tasks, 30 per family, with 60 reserve candidates. Every task requires two qualification-matched blind reviews, a criterion pack, disagreement-only adjudication, objective checks where possible, and campaign-level rights and contamination audits. The campaign currently has 0 human ballots and authorizes no model generation. Public-source contamination risk is bounded and measured; it is never described as contamination-free.
Automated judges must be calibrated against independent culinary ratings before they can support collection at scale. Public, independent-expert, author-evaluator, and automated cohorts remain separate.
Author-evaluator track
Josef Chen may contribute a separately reported author-evaluator cohort. The frozen workload targets 960 primary comparisons and 120 concealed, side-swapped reliability repeats. Task validity is sealed while only the prompt is visible. The answers are released afterward and receive nine anchored scores, side-specific failure tags, confidence, practical-verification status, and a comparative rationale.
The system records server-side review duration, preserves batch blinding, and reports repeat agreement and score drift. It enforces 320 model-arena and 640 uplift judgments rather than treating the tracks as equally sized. These controls strengthen the author-evaluator cohort but do not convert it into independent expert validation.
Open expert workspaceConflict and availability
Josef Chen develops Epicure and FlavourBench, designed this study, and plans to offer FlavourBench evaluations commercially. His dual role as author and evaluator, together with that prospective commercial interest, is material to interpretation of the development evidence.
Research lead: Josef Chen
Research artifacts
The paper, arXiv source, privacy-safe review-workload identities, frozen endpoint evidence, exact Epicure application source, reconstruction kit, and execution receipt are served with their SHA-256 identities. The exact runtime data bytes are hash-verified at the immutable public commit named in the packet. Technical public-input reconstruction has passed; payload rights, a signed immutable image, recovered training lineage, independent reproduction, and official benchmark release remain open.
The August 2026 retrospective measurement audit and exact-frontier development study.
sha256 1bbbfa8cff6f2ef4b1d50e2294d08461123e3237021e00c3e5b5e0e930fbee27Read paperSelf-contained LaTeX source, vector figures, aggregate plot data, and provenance manifests.
sha256 9df3d5ae7e732b288b281d52a8652508ece7098e0283314ec7ef97a0ca431659Download sourceBinds the exact 25-page PDF, arXiv source archive, and full paper artifact inventory distributed by this release.
sha256 f854fe9e19ff8c39e95a7b44a9f857727495aa111f429c2244ea820b5562aa28Verify release hashesLists the manuscript sources, figures, aggregate plot data, provenance records, and readiness artifacts bound into the frozen paper package.
sha256 20eade7d90ed1623969e275711b8c0fc91a61df729feae35a017ba9897d61176Inspect artifact inventoryContent-addressed roster, evidence boundary, retained workload counts, and same-origin artifact map.
sha256 7a2b5bb7dcb6df414d1b07dd73eecd37013dd391947df2442fcb642bc230e594Inspect manifest915 arena comparisons, 192 Epicure-on answer identities (188 compared and 4 unpaired), and 186 uplift pairs. It contains no prompt or answer text, raw tool traces, credentials, ballots, or quality scores and is not the governed raw research dataset.
sha256 a50f7edfe24b995b94771b4df75a1154e3a76e19519edd177a4539d3631b4622Inspect review identities2 real arms, 6 provider responses, and 2/2 successful Epicure calls. Provider cost is unknown, the alias is mutable, and the record changes no season count or quality fit.
sha256 b2f7790b3eb18d1df083397ce02b5296c549e5ed3ddb3d3f32ea776db3ddca04Inspect Qwen addendumRecords that terminalizing the complete successor source required no provider call, Epicure call, source mutation, replay, or new reservation.
sha256 7bb5f1392a2422437edc138b14940cd92736caa6bc6328acbf4b2dd73e8d479aInspect recovery1,052 attributed public human-authored questions yielded 418 provisional strict passes and 505 blind-review-eligible candidates. The frozen slate contains 180 candidates for a 120-task target. Source records remain undistributed while rights and contamination review is incomplete.
sha256 76b248477b3adc81b6eb198666a93538534db8e945567e2a99fc69085f709709Version 3 binds the 28-file exact study application, 11 runtime-data files at an immutable public commit, dependency lock, SBOM, same-operator rebuild receipt, and lineage audit. It does not establish payload rights, an immutable executable image, exact training lineage, or independent reproduction.
sha256 554ddcede820860a3150e621e9834c41bf959dee4878cd13cb01eb2c2512d52dDownload packetDeterministic source archive, immutable 11-file data map, 41-wheel dependency lock, SBOM, tool catalog, parity fixture, and a standard-library reconstruction program. The kit contains no model/data payload, dependency wheels, credentials, or training material.
sha256 bc5b109840c3c7d468fdc050eff7b907b66a0bc3a31899bd4c39ea8bcb1ae4b6Download reconstruction kitA clean CPython 3.12/Linux x86-64 run fetched and hashed all declared public inputs, installed every locked wheel offline, verified wheel RECORD hashes, reproduced the runtime payload identity, and passed fixed tool probes. It is not an independent-operator receipt.
sha256 0d28bb310614e707edf678977b2a6e548bdafb0bdf7adea185a7e46cff0a4519Inspect execution receiptBinds the verified reconstruction while keeping independent reproduction, payload redistribution clearance, training lineage, immutable OCI identity, and benchmark rank eligibility false.
sha256 878d93f8872f3f948b9a125dcb6084f01bbf06cc204e461d409611d96946f384Inspect release boundary28 exact source files (116,668 bytes) under the observed MIT code license. The archive contains neither runtime data/model payloads nor credentials.
sha256 d08fb475e9c325a8c41daf5b789e6b4bca547228139eece4578f9b06c324703cDownload exact sourceThe 16-endpoint main-stratum roster, canonical identities, fixed provider routes, and operational pair counts used by the paper and public models route. Post-freeze QwenCloud evidence is kept in its separate addendum. This record contains zero quality judgments and zero rankable comparisons.
sha256 5d834be66ec8fca611a298c2cbed75315b1f2cfc6bcd135fc465adf93c57a47dInspect frozen roster338 discovery records observed 3 August 2026. This is not a live availability, compatibility, eligibility, or quality claim.
sha256 0a9cbb38cc519cc90c337b12a55d30c5e1f931c0696eb3d1b4791af68500d0baInspect raw snapshotUnauthenticated request boundary, selected response headers, raw byte count, content hash, and claim limitations for the OpenRouter snapshot.
sha256 2d831746e833a7c3c2e04e3d124adbfbdd07319d5cbcc47bed6ca69257404cf6Inspect receiptProtocol, task holds, inference boundaries, governance, privacy, and limitations for the current study.
Read methodsRaw prompts, outputs, and tool traces remain private pending source-rights, provider-output, privacy, and identity-leakage review.
No public archival deposit or DOI is claimed on this site yet.