Skip to content

LizardCalculator: one lizard process per run, not per file #138

Description

@julianrubisch

Attractor::Lizard.analyze (#137) spawns one lizard process per file. lizard (and more so uvx lizard on a cold cache) starts in a few hundred ms, so a 2,000-file repository spends minutes on process startup alone.

lizard accepts many paths or a directory in one call and prints the file column per row, so the calculator could run it once per calculate and index the rows by path. The obstacle is BaseCalculator#calculate, which yields per change after computing target_paths privately; the subclass sees files one at a time. Options:

  1. Give BaseCalculator a before_calculate(target_paths) hook (or yield the whole list first) so LizardCalculator can prefetch.
  2. Run lizard over Dir.pwd lazily on the first yield, index by normalized path, serve subsequent yields from the index. Simpler, but scans files the churn filter excluded.

Related, surfaced in the same review: Cache keys on file path and commit only, not on calculator type. Two plugins claiming the same extension would share rows, and a cached row from before the details shape change is replayed as is. Worth a cache-key namespace (type or calculator class) when touching this.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions