Skip to content

cleanup invalid indexes before reindex - #66

Merged
Komzpa merged 1 commit into
mainfrom
feat-cleanup-invalid-indexes
Oct 2, 2025
Merged

Komzpa merged 1 commit into
mainfrom
feat-cleanup-invalid-indexes

Conversation

@sleeping-h

@sleeping-h sleeping-h commented Oct 2, 2025 •

Copy link
Copy Markdown
Contributor

Summary by CodeRabbit

  • New Features

    • Automatically detects and drops invalid PostgreSQL indexes during maintenance, improving cluster health with minimal disruption.
  • Bug Fixes

    • Reduces reindex failures by running invalid-index cleanup before btree reindexing.
    • De-duplicates btree bloat report results and standardizes sorting for clearer output.
  • Chores

    • Updates maintenance workflow to include an additional pre-reindex safety step for smoother operations.

@coderabbitai

coderabbitai Bot commented Oct 2, 2025 •

Copy link
Copy Markdown

Walkthrough

Adds a pre-step in the reindex_btree_loop to run a new SQL script that drops invalid indexes before executing the existing reindex script. Updates the btree bloat SQL to select distinct table names and sort by the first column. Introduces a PL/pgSQL script to conditionally drop invalid indexes with non-blocking locks.

Changes

Cohort / File(s) Summary
Build automation
Makefile
In reindex_btree_loop, add execution of psql -qf scripts/drop_invalid_indexes.sql immediately before scripts/reindex-bloated-btrees.sh.
Bloat analysis SQL
scripts/btree_bloat-superuser.sql
Add DISTINCT to table name selection; change ORDER BY from idxname to ORDER BY 1.
Invalid index cleanup script
scripts/drop_invalid_indexes.sql
New PL/pgSQL DO block: iterate invalid indexes, try ACCESS EXCLUSIVE LOCK on parent table NOWAIT; on success, DROP INDEX; on lock failure, log and skip; continue processing.

Sequence Diagram(s)

sequenceDiagram
  autonumber
  participant Make as Make reindex_btree_loop
  participant Psql as psql
  participant SQL as drop_invalid_indexes.sql
  participant ReIdx as reindex-bloated-btrees.sh
  Make->>Psql: Run scripts/drop_invalid_indexes.sql
  Psql->>SQL: Execute DO block
  Note over SQL: For each invalid index:<br/>TRY lock parent table NOWAIT<br/>IF lock ok → DROP INDEX<br/>ELSE → LOG skipped
  Make->>ReIdx: Run scripts/reindex-bloated-btrees.sh
Loading
sequenceDiagram
  autonumber
  participant SQL as drop_invalid_indexes.sql
  participant Catalog as pg_catalog
  loop Each invalid index
    SQL->>Catalog: SELECT invalid indexes (indisvalid = false)
    SQL->>Catalog: LOCK TABLE ... ACCESS EXCLUSIVE NOWAIT
    alt Lock acquired
      SQL->>Catalog: DROP INDEX CONCURRENTLY? (DROP INDEX)
      Note right of SQL: NOTICE: Dropped invalid index
    else Lock not available
      Note right of SQL: NOTICE: Skipped (table busy)
    end
  end
Loading

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Pre-merge checks and finishing touches

✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title Check ✅ Passed The title clearly and concisely describes the primary change—dropping invalid indexes before running the reindex procedure—and aligns directly with the added cleanup script and updated Makefile.
Docstring Coverage ✅ Passed No functions found in the changes. Docstring coverage check skipped.
✨ Finishing touches
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch feat-cleanup-invalid-indexes

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@Komzpa
Komzpa merged commit 978f572 into main Oct 2, 2025
1 of 2 checks passed
@Komzpa
Komzpa deleted the feat-cleanup-invalid-indexes branch October 2, 2025 16:51

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

📜 Review details

Configuration used: CodeRabbit UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between eb020c8 and d673887.

📒 Files selected for processing (3)
  • Makefile (1 hunks)
  • scripts/btree_bloat-superuser.sql (2 hunks)
  • scripts/drop_invalid_indexes.sql (1 hunks)
🧰 Additional context used
📓 Path-based instructions (2)
Makefile

📄 CodeRabbit inference engine (AGENTS.md)

Makefile: If you need intermediate result from other target in Makefile, split it into two and depend on the intermediate result.
Format target comments as self-documented Makefile, on same line: target: dependencies | order_only_deps ## Description
Remember that Makefile uses tabs.
Explain high-level architecture and quirks in Makefile.
Makefile: there are comments on the same line after each target separated by ## - they are used in debug graph visualization, need to be concise and descriptive of what's going on in the code itself.

Files:

  • Makefile
**/*.sql

📄 CodeRabbit inference engine (AGENTS.md)

**/*.sql: Format SQL queries in a way so it's easy to copy them out of the codebase and debug standalone.
SQL is lowercase, PostGIS functions follow their spelling from the manual (st_segmentize -> ST_Segmentize).
SQL files should be idempotent: drop table if exists; add some comments to make people grasp queries faster.
Prefer indexed SQL operators when dealing with jsonb ( tags @> '{"key": "value"}' instead of tags ->> 'key' = 'value' ).

Files:

  • scripts/drop_invalid_indexes.sql
  • scripts/btree_bloat-superuser.sql
🔇 Additional comments (2)
scripts/btree_bloat-superuser.sql (1)

28-105: Output deduping looks good.

Selecting distinct table names and ordering on the first column makes the downstream loop deterministic without altering the underlying bloat detection.

Makefile (1)

30-31: Pre-step integration LGTM.

Running the cleanup script ahead of the reindex loop tightens the maintenance flow without altering existing behavior.

Comment on lines +1 to +31
DO $$
DECLARE
idx RECORD;
BEGIN
FOR idx IN
SELECT n.nspname AS schema_name,
c.relname AS index_name,
t.oid AS table_oid,
t.relname AS table_name
FROM pg_index i
JOIN pg_class c ON c.oid = i.indexrelid
JOIN pg_namespace n ON n.oid = c.relnamespace
JOIN pg_class t ON t.oid = i.indrelid
WHERE NOT i.indisvalid
LOOP
BEGIN
-- try to take an ACCESS EXCLUSIVE lock on the table, without waiting
EXECUTE format('LOCK TABLE %I.%I IN ACCESS EXCLUSIVE MODE NOWAIT',
idx.schema_name, idx.table_name);

-- if lock succeeds, drop index
RAISE NOTICE 'Dropping invalid index: %.%', idx.schema_name, idx.index_name;
EXECUTE format('DROP INDEX IF EXISTS %I.%I;', idx.schema_name, idx.index_name);

EXCEPTION WHEN lock_not_available THEN
RAISE NOTICE 'Skipped index %.% (table %.% is busy)',
idx.schema_name, idx.index_name, idx.schema_name, idx.table_name;
END;
END LOOP;
END;
$$ LANGUAGE plpgsql;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Follow repository SQL casing conventions.

Per the project SQL guidelines, keywords should be lowercase. Please downcase DO, DECLARE, SELECT, LOCK, DROP, etc., in this script for consistency.

🤖 Prompt for AI Agents
In scripts/drop_invalid_indexes.sql around lines 1 to 31, the SQL script uses
uppercase SQL keywords; update the file to follow repository SQL casing
conventions by converting all SQL keywords (e.g., DO, DECLARE, BEGIN, FOR,
SELECT, JOIN, WHERE, LOOP, EXECUTE, FORMAT, RAISE NOTICE, EXCEPTION WHEN, END,
LANGUAGE, etc.) to lowercase while leaving identifiers (schema/table/index
names), string literals, and punctuation unchanged; ensure formatting and
indentation remain the same and verify the script still runs after the casing
changes.

Comment on lines +21 to +27
-- if lock succeeds, drop index
RAISE NOTICE 'Dropping invalid index: %.%', idx.schema_name, idx.index_name;
EXECUTE format('DROP INDEX IF EXISTS %I.%I;', idx.schema_name, idx.index_name);

EXCEPTION WHEN lock_not_available THEN
RAISE NOTICE 'Skipped index %.% (table %.% is busy)',
idx.schema_name, idx.index_name, idx.schema_name, idx.table_name;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

Abort risk when invalid index backs a constraint.

DROP INDEX raises SQLSTATE 2BP01 (dependent_objects_still_exist) for primary-key/unique indexes that back constraints. Because the DO block doesn’t trap it, the first such index aborts the whole script and your Make target fails, leaving the reindex loop untouched. Please skip (or explicitly handle) indexes referenced by pg_constraint before issuing DROP INDEX.

Apply a guard in the index selection, e.g.:

-        WHERE NOT i.indisvalid
+        WHERE NOT i.indisvalid
+          AND NOT EXISTS (
+              SELECT 1
+              FROM pg_constraint c
+              WHERE c.conindid = i.indexrelid
+          )

Committable suggestion skipped: line range outside the PR's diff.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants