Skip to content

Default to Claude Opus 5.5 with refreshed benchmarks - #11

Merged
judagent[bot] merged 2 commits into
masterfrom
feat/opus-5-5-advisor
Sep 23, 2026
Merged

judagent[bot] merged 2 commits into
masterfrom
feat/opus-5-5-advisor

Conversation

@judagent

@judagent judagent Bot commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Default advisor: anthropic/claude-opus-5-5#xhigh (was claude-fable-5-1#xhigh), with one opencode retry when the Anthropic route is unavailable. The serving route is recorded as advisorProvider on metrics rows.
  • Self-consultation guard generalized from hard-coded Fable to the configured advisor model (any provider route, any effort, normalized across nested/dotted IDs). New skipped_self outcome; legacy skipped_fable history parsing kept.
  • Benchmarks: fresh Artificial Analysis snapshot (673 models, fetched 2026-09-23) and 69 bindings covering Opus 5.5 (direct + opencode, all efforts), GPT-6 Sol/Luna, Grok 4.7, new Fable 5.1 / GPT-6 Astra effort levels, gateway DeepSeek V4.1 Flash routes, and Muse Spark 1.2. The generator now fails on missing slugs instead of skipping them, and a seed test asserts every bundled binding resolves.
  • Requester variant "default" treated as unset for benchmark matching; usage-report eligibility proxy parameterized (--advisor-model, default claude-opus-5-5).

Behavioral notes

  • Fable sessions (including the plan agent on claude-fable-5-1#xhigh) will now see the advisor tool; Opus 5.5 sessions are now the blocked ones.
  • Opus 5.5 has no AA Coding Index yet, so coding comparisons report missing_advisor until Artificial Analysis publishes one. Watch skipped_typesafe rates for the first few days.
  • deepseek/deepseek-flash and muse-spark-1.3#high stay unmapped (no AA entry; no effort substitution per policy).

Verification

  • bun test: 361 pass, 0 fail; tsc --noEmit clean; bun run build clean.
  • Compiled CLI benchmarks status --model: 25/27 top-usage models matched (2 intentional exclusions above); anthropic/claude-opus-5-5#xhigh matched with evaluated effort xhigh.
  • Fallback covered on both discovery namespaces (legacy catalog + V2); failed fallback reports the primary reason.

Models: anthropic/claude-opus-5-5:xhigh, meta/muse-spark-1.3:max, anthropic/claude-fable-5-1:xhigh

judagent Bot added 2 commits September 23, 2026 01:08
Pull a fresh Artificial Analysis snapshot (673 models, 2026-09-23) and
regenerate the bundled mappings: 69 bindings covering the Opus 5.5 advisor
(direct and opencode routes, all efforts), GPT-6 Sol/Luna, Grok 4.7, new
effort levels for Fable 5.1 and GPT-6 Astra, gateway DeepSeek V4.1 Flash
routes, and Muse Spark 1.2. The generator now fails on missing slugs
instead of silently skipping them, and a seed test asserts every bundled
binding resolves to a snapshot record.
Change the default advisor from anthropic/claude-fable-5-1#xhigh to
anthropic/claude-opus-5-5#xhigh, with one opencode retry when the
Anthropic route is unavailable; the serving route is recorded on the
metrics row. Generalize the self-consultation guard from hard-coded
Fable to the configured advisor model (any route, any effort) with a
new skipped_self outcome, keeping legacy skipped_fable history parsing
intact. Treat a recorded default effort as unset for benchmark matching,
and parameterize the usage report's eligibility proxy by advisor model.
@judagent judagent Bot assigned judsd Sep 23, 2026
@judagent
judagent Bot merged commit 60c33ca into master Sep 23, 2026
3 checks passed
@judagent
judagent Bot deleted the feat/opus-5-5-advisor branch September 23, 2026 01:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant