Skip to content

Prefer relevant evidence induction in search - #6

Merged
samth merged 1 commit into
mainfrom
feat/evidence-first-induction
Sep 30, 2026
Merged

samth merged 1 commit into
mainfrom
feat/evidence-first-induction

Conversation

@samth

@samth samth commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Search can spend its budget on data induction before trying induction on a relevant hypothesis. Prefer evidence induction when the head predicate of the hypothesis occurs in the target.

This adds a small InductionPlan.evidenceFirst adapter over the existing Hooks.order interface and applies it to search mode. It stably partitions each candidate batch after the existing ordering. Every alternative remains available, and order within each partition is preserved. Batches without evidence induction take the original path. Committed mode, inference operations, effort accounting, and the trial schedule are unchanged.

The relevance test uses the weak-head-normalized hypothesis type and syntactic constants in the target. The source comment describes its limitations: hidden predicates can be missed and incidental occurrences can be preferred. There are no theorem-specific rules or numeric thresholds.

Validation

  • Focused tests cover the default search integration, composition with a reversed inner ordering, stable ties, unrelated and variable-headed hypotheses, absent major premises, and batches without evidence.
  • Complete package tests passed on Lean 4.33.1 on CPU node kj (lake test, 60 targets, exit 0). GitHub CI passed, including the Lean 4.30.0 / 4.33.1 matrix and website checks.
  • The completed prototype evaluation and raw results checked all 82 current case-study files without changing their proofs, hints, or budgets: all 2,227 measured tactic calls succeeded in both configurations. Summed internal time fell from 1,859.849 s to 1,801.782 s (3.12%); raw heartbeats fell 2.39%. P90 fell 7.99% and P99 fell 4.02%, with essentially unchanged P50.
  • Penn STLC weakening fell from 1,255 to 531 attempts; extended-STLC weakening fell from 1,897 to 769. Both still require the explicit helper hint.

Those timings are single observations of the equivalent ordering prototype on the PR #5 base with premise retrieval off, not a repeated speedup estimate or a new measurement of this clean integration. This PR is independently based on main and includes neither premise selection nor the unsuccessful bounded-contour experiments.

@samth
samth merged commit 12be67f into main Sep 30, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant