fix(indexes): avoid __dbt_tmp/__dbt_backup leakage in CCI names (#578) - #751
fix(indexes): avoid __dbt_tmp/__dbt_backup leakage in CCI names (#578)#751mopthe wants to merge 1 commit into
Conversation
There was a problem hiding this comment.
Pull request overview
This PR fixes SQL Server index-macro behavior to prevent “orphaned” clustered columnstore index (CCI) naming during materializations that involve intermediate __dbt_tmp / __dbt_backup relations, by deriving index names from the final relation name and improving identifier quoting. It also adds unit + integration coverage around determinism/idempotence, INCLUDE columns, and quoted identifiers.
Changes:
- Refactors index macros to use deterministic index naming derived from the final relation identifier (suffix-stripped) and improved quoting.
- Updates the CCI creation macro to avoid
__dbt_tmp/__dbt_backupleaking into the index name. - Adds comprehensive unit tests for macro rendering plus live SQL Server integration tests validating idempotence, INCLUDE semantics, quoted identifiers, and the orphaned-CCI regression.
Reviewed changes
Copilot reviewed 4 out of 4 changed files in this pull request and generated 3 comments.
| File | Description |
|---|---|
dbt/include/sqlserver/macros/adapters/indexes.sql |
Adds quoting/qualification helpers and refactors index/CCI macros to produce deterministic names and safer SQL. |
dbt/include/sqlserver/macros/relations/table/create.sql |
Updates the inline comment to document the new “final relation name” index naming behavior. |
tests/unit/adapters/mssql/test_indexes.py |
Adds Jinja-rendered unit tests for quoting, suffix-stripping, deterministic names, and emitted SQL shapes. |
tests/functional/adapter/mssql/test_index_macros.py |
Adds integration tests against live SQL Server for incremental stability, INCLUDE columns, quoted identifiers, and orphaned-CCI regression. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Migration Note: temporary workaround
SQL to inspect and generate drop commands: -- Identify orphaned CCI indexes
SELECT
name AS index_name,
OBJECT_SCHEMA_NAME(object_id) AS schema_name,
OBJECT_NAME(object_id) AS table_name
FROM sys.indexes
WHERE name LIKE '%__dbt_tmp%cci';
-- Drop each one
DROP INDEX [index_name] ON [schema_name].[table_name]; |
|
I think you have converted the file endings because the file is showing a huge number of changes and I don't think there are any there, it makes the pull request difficult to parse. |
6497d0e to
8a2916f
Compare
|
@Benjamin-Knight Thanks! Should be readable now. |
c845896 to
f6f601b
Compare
|
After a careful analysis, this is a bug with how Commvault works, we deliberatelly chose to make a swap and avoid building this heavy index while the table is being read, this approach doesnt fix the issue because commvault just uses names to check object ownership, can you warrant it indeed fixes it @mopthe ? |
|
Thanks for the Commvault context. Just to clarify, this PR only changes the CCI name by removing the If Commvault tracks ownership by index name, consistent naming should improve compatibility, not break it. Could you share more detail about the specific interaction you're seeing? |
Head branch was pushed to by a user without write access
f6f601b to
ac9c3e7
Compare
|
The index work I did also means that a certain prefix is required for DBT to consider indexes part of its managed set for that code, and that it could drop codes if you enable the flag that DBT does not think its managing. I believe CCIs are exempt from this but best to check how your changes interact with that code. |
|
@Benjamin-Knight You are right that CCI is exempt by type, not name. Reconciliation checks |
|
@mopthe The problem, Commvault will still save the ownership using the parent name, so I don't understand how this fix the issue, meaning, if Commvault already got meta about the temporary table and the CCI, parent which will be always later missing, it doesn't matter how is the CCI called, it will already have the wrong map. Probably, the fix is to change the name of the CCI index too, not only the parent table, so Commvault (and any other external like that) understands its completely missing. But I wouldn't like to implement something without real testing, can you test that or reference good docs about this? |
|
@axellpadilla I understand the issue now. I mistakenly associated this PR with a problem that only partially matches what this PR actually addresses. I was focused on the __dbt_tmp index naming and jumped to the wrong conclusion. Thanks for the additional context. I'm away from the office at the moment, but I'll look into this when I'm back in about a week. |
|
Thanks for the detailed review. I do not plan to continue this PR further from my side. If this direction is not useful for the Commvault scenario, I am fine with abandoning this PR. I think we should split this into two separate issues:
This PR is scoped to Issue A only. Important limitation: |
ac9c3e7 to
ba45ce0
Compare
…msft#578) Add strip_dbt_suffix logic for CCI name generation and keep only dbt-msft#578-focused unit coverage.
ba45ce0 to
b5d2ad9
Compare
Partially addresses #578: this change fixes CCI naming leakage only; it does not change the tmp->rename build/swap flow.