Repository navigation
Default to Claude Opus 5.5 with refreshed benchmarks - #11
Merged
Merged
Conversation
Pull a fresh Artificial Analysis snapshot (673 models, 2026-09-23) and regenerate the bundled mappings: 69 bindings covering the Opus 5.5 advisor (direct and opencode routes, all efforts), GPT-6 Sol/Luna, Grok 4.7, new effort levels for Fable 5.1 and GPT-6 Astra, gateway DeepSeek V4.1 Flash routes, and Muse Spark 1.2. The generator now fails on missing slugs instead of silently skipping them, and a seed test asserts every bundled binding resolves to a snapshot record.
Change the default advisor from anthropic/claude-fable-5-1#xhigh to anthropic/claude-opus-5-5#xhigh, with one opencode retry when the Anthropic route is unavailable; the serving route is recorded on the metrics row. Generalize the self-consultation guard from hard-coded Fable to the configured advisor model (any route, any effort) with a new skipped_self outcome, keeping legacy skipped_fable history parsing intact. Treat a recorded default effort as unset for benchmark matching, and parameterize the usage report's eligibility proxy by advisor model.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
anthropic/claude-opus-5-5#xhigh(wasclaude-fable-5-1#xhigh), with oneopencoderetry when the Anthropic route is unavailable. The serving route is recorded asadvisorProvideron metrics rows.skipped_selfoutcome; legacyskipped_fablehistory parsing kept.opencode, all efforts), GPT-6 Sol/Luna, Grok 4.7, new Fable 5.1 / GPT-6 Astra effort levels, gateway DeepSeek V4.1 Flash routes, and Muse Spark 1.2. The generator now fails on missing slugs instead of skipping them, and a seed test asserts every bundled binding resolves."default"treated as unset for benchmark matching; usage-report eligibility proxy parameterized (--advisor-model, defaultclaude-opus-5-5).Behavioral notes
planagent onclaude-fable-5-1#xhigh) will now see the advisor tool; Opus 5.5 sessions are now the blocked ones.missing_advisoruntil Artificial Analysis publishes one. Watchskipped_typesaferates for the first few days.deepseek/deepseek-flashandmuse-spark-1.3#highstay unmapped (no AA entry; no effort substitution per policy).Verification
bun test: 361 pass, 0 fail;tsc --noEmitclean;bun run buildclean.benchmarks status --model: 25/27 top-usage modelsmatched(2 intentional exclusions above);anthropic/claude-opus-5-5#xhighmatched with evaluated effortxhigh.Models: anthropic/claude-opus-5-5:xhigh, meta/muse-spark-1.3:max, anthropic/claude-fable-5-1:xhigh