Repository navigation
Cross-provider benchmark matching + generator filename fix - #10
Conversation
Implemented: material-value gate at threshold 0.20Committed and pushed in b9fd18e. The experiment report is now versioned in the repository and linked from the README: Advisor gate policy evaluation — full report and exact prompt It preserves the alternatives considered, 264 live TypeSafe experiment calls, discovery/holdout results, requester-by-requester confirmation, case-level probabilities, exact selected instructions and criteria, and limitations. Keeping the cutoff at 0.20 matters: a 0.40 cutoff with the new question suppressed a legitimate review in the experiment. Fresh implementation checks
The six added regression scenarios cover direct script lookup, mechanical rename, spelling review, unchanged repeat question, explicit user request for the advisor, and consequential one-line authorization change. Benchmark evidence and independent effort selection remain in the same request. Implementation complete, awaiting review/merge. These results are synthetic scenario evaluations, not production error-rate estimates; the running server has not been switched to the worktree build. Models: openai/gpt-6-astra:xhigh |
Summary
Follow-ups to PR #9 and #8:
benchmarks.matchAnyProviderdefaults tofalse. When no exact provider binding exists, it can resolve another provider's binding for the same model ID and variant. Exact provider bindings still win, variants stay independent, and local overrides take precedence over bundled mappings. Exposed through the plugin option,OCADVISOR_BENCHMARKS_MATCH_ANY_PROVIDER, andocadvisor benchmarks status --match-any-provider true. Baseline mappings grew from 24 to 28.skipBelow: 0.20. Direct lookups, deterministic operations, mechanical edits, and unchanged already-answered questions can skip; substantive decisions, diagnosis, correctness concerns, changed evidence, and explicit user requests remain grounds for consultation. Missing task context preserves uncertainty. Model/benchmark evidence remains available and effort selection stays independent.snapshotPathalready resolved to the same filename. The earlier PR description overstated it as a filename bug fix.Gate evaluation report
Full report at implementation commit b9fd18e
jev-1.13.0.Verification at b9fd18e
bun test: 338 passed, 0 failed.bun run typecheck,bun run build, andgit diff --check: passed.opencode/deepseek-v4.1-flash#max.Implementation complete; awaiting review/merge. The running server configuration has not been switched to this worktree.
Models: openai/gpt-6-astra:xhigh, anthropic/claude-fable-5-1:xhigh, meta/muse-spark-1.3:max