Skip to content

Cross-provider benchmark matching + generator filename fix - #10

Merged
judsd merged 3 commits into
masterfrom
fix/artificial-analysis-followups
Sep 20, 2026
Merged

judsd merged 3 commits into
masterfrom
fix/artificial-analysis-followups

Conversation

@judagent

@judagent judagent Bot commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Follow-ups to PR #9 and #8:

  • Optional cross-provider benchmark matching: benchmarks.matchAnyProvider defaults to false. When no exact provider binding exists, it can resolve another provider's binding for the same model ID and variant. Exact provider bindings still win, variants stay independent, and local overrides take precedence over bundled mappings. Exposed through the plugin option, OCADVISOR_BENCHMARKS_MATCH_ANY_PROVIDER, and ocadvisor benchmarks status --match-any-provider true. Baseline mappings grew from 24 to 28.
  • Material-value advisor screening: replace the broad need question with the evaluated structured question and explicit true/false criteria. Keep skipBelow: 0.20. Direct lookups, deterministic operations, mechanical edits, and unchanged already-answered questions can skip; substantive decisions, diagnosis, correctness concerns, changed evidence, and explicit user requests remain grounds for consultation. Missing task context preserves uncertainty. Model/benchmark evidence remains available and effort selection stays independent.
  • Durable evaluation report: records the exact selected prompt, alternatives, 264 live experimental calls, held-out results, case-level probabilities, limitations, and implementation verification. Six additional opt-in evaluation cases cover the new policy boundaries.
  • Generator path cleanup: the earlier generator commit in this PR inlines the canonical snapshot path. This is behavior-preserving; the previous snapshotPath already resolved to the same filename. The earlier PR description overstated it as a filename bug fix.

Gate evaluation report

Full report at implementation commit b9fd18e

  • 24 synthetic scenarios, DeepSeek/Muse/Astra requester profiles, Fable advisor profile; gate model pinned to jev-1.13.0.
  • 192 comparison calls: selected criteria at 0.20 avoided 60/60 routine calls and preserved 84/84 useful calls. Reserved holdout: 24/24 routine skips, 24/24 useful proceeds.
  • 72 production-shaped confirmation calls: 30/30 routine skips, 42/42 useful proceeds, zero fallbacks.
  • Raising the selected prompt's threshold to 0.40 wrongly skipped a focused correctness review, so the implementation retains 0.20.
  • These are repeated synthetic scenario results, not an estimate of production error rates. No advisor generation was invoked for the evaluations.

Verification at b9fd18e

  • bun test: 338 passed, 0 failed.
  • bun run typecheck, bun run build, and git diff --check: passed.
  • Production need question and criteria exactly match the frozen selected experimental prompt.
  • Updated 16-case live evaluation, before → after: routine skips 0/6 → 6/6, useful consultations preserved 10/10 → 10/10, zero fallbacks in both runs.
  • SDK-boundary regression verifies structured instructions and both criteria are serialized, with exactly the normal need and effort questions.
  • Earlier genuine-snapshot checks verified strict and cross-provider matching for opencode/deepseek-v4.1-flash#max.

Implementation complete; awaiting review/merge. The running server configuration has not been switched to this worktree.

Models: openai/gpt-6-astra:xhigh, anthropic/claude-fable-5-1:xhigh, meta/muse-spark-1.3:max

@judagent

judagent Bot commented Sep 20, 2026

Copy link
Copy Markdown
Contributor Author

Implemented: material-value gate at threshold 0.20

Committed and pushed in b9fd18e.

The experiment report is now versioned in the repository and linked from the README:

Advisor gate policy evaluation — full report and exact prompt

It preserves the alternatives considered, 264 live TypeSafe experiment calls, discovery/holdout results, requester-by-requester confirmation, case-level probabilities, exact selected instructions and criteria, and limitations. Keeping the cutoff at 0.20 matters: a 0.40 cutoff with the new question suppressed a legitimate review in the experiment.

Fresh implementation checks

Check Result
Automated tests 338 pass, 0 fail
Typecheck / build / diff check Passed
Exact equality to tested question and criteria Passed
16-case live evaluation: routine skips Before 0/6 → after 6/6
16-case live evaluation: useful consultations preserved Before 10/10 → after 10/10
Live evaluation fallbacks 0 before and after

The six added regression scenarios cover direct script lookup, mechanical rename, spelling review, unchanged repeat question, explicit user request for the advisor, and consequential one-line authorization change. Benchmark evidence and independent effort selection remain in the same request.

Implementation complete, awaiting review/merge. These results are synthetic scenario evaluations, not production error-rate estimates; the running server has not been switched to the worktree build.

Models: openai/gpt-6-astra:xhigh

@judsd
judsd merged commit 436fd1b into master Sep 20, 2026
4 checks passed
@judsd
judsd deleted the fix/artificial-analysis-followups branch September 20, 2026 01:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant