Dive into a git repo's history: per-commit snapshots, an indexed metrics catalog and an interactive dashboard
Point repo-dive at any git repository and get an explorable catalog of insights derived from its history:
cd /path/to/your/repo
npx repo-diveOne command runs the whole pipeline — scan, index, dashboard — and opens the results in your browser.
- Map: walk the repo's commits (all or sampled) and let pluggable collectors capture raw snapshots per commit — language/LOC breakdowns, author stats, lint diagnostics and more.
- Reduce: index those snapshots into a local metrics store shaped like a data cube — numbers at intersections of open-ended categories (author, language, date, lint rule, …).
- Explore: query the cube to draw charts, export shareable reports and ask AI questions about how the codebase evolved.
Everything is local-first, incremental and resumable: results live in a catalog folder inside the repo being analyzed and are refined over multiple runs.
See docs/specs for the architecture and docs/research/prior-art.md for a survey of existing tools and why none of them fills this niche.
Live dashboards for a few popular repositories, produced by running the tool on their full history:
- curl (C, since 1999)
- effect (TypeScript, since 2020)
- ollama (Go, since 2023)
- prettier (JavaScript, since 2016)
- react (JavaScript, since 2013)
- transformers (Python, since 2018)
- vite (TypeScript, since 2020)
Each one is a single self-contained HTML file exported with repo-dive report and redeployed weekly by a scheduled workflow — see examples for how they are defined.
Run from inside the repository you want to analyze (or pass --repo /path/to/repo).
Node 22.15 or newer is required.
npx repo-dive # the whole pipeline: scan + index + dashboard
npx repo-dive scan # collect snapshots into .repo-dive/
npx repo-dive index # roll up into the metrics cube + dashboard data
npx repo-dive dashboard # serve the interactive dashboard
npx repo-dive status # show catalog coverage
npx repo-dive collectors # list available collectors
npx repo-dive report # export one shareable self-contained HTML file
npx repo-dive mcp # serve the cube to AI agents (Model Context Protocol)
npx repo-dive gc # clean up the catalog interactively
npx repo-dive ignore # keep other tools out of the catalog
npx repo-dive query "SELECT metric, sum(value) FROM facts GROUP BY metric"scan walks the repository's history and runs collectors against every commit (or a sample, per collector), writing raw snapshots into a .repo-dive/ catalog inside the analyzed repo.
It is resumable: re-running skips everything already collected, and bumping a collector's version invalidates only that collector's outputs.
Checkout-based collectors use temporary detached worktrees — the analyzed repo's working tree is never touched.
Collectors so far:
- commit-meta — identities, dates, parents, subject and trailers (incl. AI co-authors)
- churn — lines added/deleted per commit, by file extension
- file-types — file count and bytes per extension at each commit's tree
- directives — eslint-disable comments by rule (block disables tracked as gray areas) and
@ts-ignore/@ts-expect-error/@ts-nocheck - dependencies — resolved package totals from lockfiles, per package manager (pnpm, npm, yarn classic and yarn berry), plus direct/dev/optional dependencies and manifest counts read straight from
package.jsonfiles; version-aware, monorepo-aware and extensible to more managers - todo-comments — TODO/FIXME/HACK/XXX counts
- languages — lines and file count per language across a commit's source files (lockfiles, minified bundles and generated data excluded)
- survival —
git blameline survival by extension, author and age cohort (sampled monthly) - file-survival — file survival by extension, creator and creation cohort, renames followed (sampled monthly)
The catalog hides itself from git, but other tools that walk the repository (prettier, markdownlint, cspell, docker builds) each read one ignore file at its root.
scan warns when the catalog is missing from those; repo-dive ignore adds it to every one that needs it, writing the entry in the shape the file is already written in and skipping the files whose tool learns about the catalog elsewhere.
index normalizes raw snapshots into .repo-dive/index/metrics.sqlite — a facts-by-categories cube, rebuildable at any time — plus dashboard.json.
dashboard then serves a local React app with interactive charts: languages over time, file counts over time, a GitHub-style commit calendar, monthly commits with AI-assisted share, churn, lint-suppression trends, dependency counts over time, code survival by cohort and author, and more.
Everything works with zero config.
To refine it, drop a repo-dive.config.ts at the root of the repository you analyze (.mjs/.js also work):
import { defineConfig } from "repo-dive/config";
export default defineConfig({
contributors: {
aliases: [
// Shorthand: emails only, the first is canonical.
["alice@work.example", "alice@personal.example"],
// Rich form: a display name, a profile link and an explicit kind.
{
displayName: "Bob",
emails: ["bob@work.example", "12345+bob@users.noreply.github.com"],
url: "https://github.com/bob",
},
],
// How many contributors charts keep before folding the rest into "Other" (default 10).
maxInCharts: 10,
},
charts: {
// First day of the week in calendar-shaped charts (default "monday").
weekStartsOn: "monday",
},
catalog: {
// Where snapshots, caches and the cube live (default ".repo-dive").
dir: ".repo-dive",
},
});charts.weekStartsOn sets the first day of the week in calendar-shaped charts such as the commit calendar ("monday" by default, "sunday" also supported).
contributors.aliases merges the multiple identities one person commits under (work + personal email, GitHub noreply, name variants) so attribution, the contributors table and code-survival-by-contributor count them once.
A group can also carry a displayName, a profile url and a kind (human/bot/ai, otherwise auto-derived).
The dashboard badges bots and AI agents, listing them apart from humans.
catalog.dir moves the catalog; point it outside the repository (e.g. "../repo-dive-catalogs/my-repo") to leave the analyzed working tree untouched altogether, ignore files included.
Apart from catalog, which every command needs, the config is read by index.
See docs/specs/07-config.md for details.
The repository doubles as a composite GitHub Action, so any project can produce and refresh its report right in CI instead of on somebody's laptop. Commit a workflow like this to the repository you want to analyze:
name: repo-dive
on:
schedule:
- cron: "27 5 * * 1" # weekly
workflow_dispatch:
permissions:
contents: read
jobs:
report:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
with:
fetch-depth: 0 # the whole history is the whole point
- uses: kachkaev/repo-dive@mainEvery run restores the catalog from the Actions cache, scans only the commits that are new since the previous run and uploads the self-contained report as an artifact, viewable straight from the run page. Long first scans are banked too: when the scan hits its time limit, the progress is cached and the next run resumes where this one stopped. See docs/github-action.md for all inputs, publishing the report to GitHub Pages and analyzing repositories other than the workflow's own.
repo-dive mcp serves the metrics cube over the Model Context Protocol on stdio, so an agent can explore a repository's history by asking SQL questions.
Two tools:
schema— tables, available metrics with row counts, sample category keys per metric and the commit range; worth calling before writing queries.query— one read-only statement (SELECT/WITH/EXPLAIN) against the cube, returning{ columns, rows, truncated }(up to 200 rows).
Run scan and index first: the server exits immediately if there is no cube at .repo-dive/index/metrics.sqlite.
The database is opened read-only, so nothing an agent asks can change the catalog.
For Claude Code, run this inside the repository you want to ask questions about:
claude mcp add repo-dive -- npx -y repo-dive mcpOr commit a project-scoped .mcp.json at the repository root, so everyone on the team gets the same server:
{
"mcpServers": {
"repo-dive": {
"command": "npx",
"args": ["-y", "repo-dive", "mcp"]
}
}
}Then ask things like "which languages grew fastest last year?" or "how has the share of AI-assisted commits changed?".
The same stdio server works with any MCP client — point yours at npx repo-dive mcp, adding --repo /path/to/repo if the client does not start it inside the repository being analyzed.
The project is written in TypeScript with Effect v4 and its built-in CLI toolkit (effect/cli).
pnpm install
pnpm test
pnpm lint
pnpm fixTo see how heavy the published package would be:
pnpm build && pnpm report-package-sizeIt prints the tarball and unpacked sizes with a per-file breakdown, comparing them against the previous measurement.
CI runs the same report on every push and adds it to the job summary, comparing against the latest main.
Thanks to @WillJack20 for suggesting the name repo-dive, formerly repo-insighter.