Skip to content

Fix server startup after an interrupted CREATE OR REPLACE on a non-atomic database disk - #123902

Merged
alexey-milovidov merged 1 commit into
masterfrom
fix-tmp-replace-duplicate-uuid-on-load
Oct 9, 2026
Merged

alexey-milovidov merged 1 commit into
masterfrom
fix-tmp-replace-duplicate-uuid-on-load

Conversation

@alexey-milovidov

@alexey-milovidov alexey-milovidov commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

On a database disk without atomic renames (plain_rewritable object storage), the rename of the temporary _tmp_replace_* table of CREATE OR REPLACE copies the metadata file and then removes the source. If the server is killed in between, two metadata files refer to the same table, and the server cannot start anymore:

Mapping for table with UUID=968b64dd-... already exists ... Cannot parse definition from metadata file store/832/.../_tmp_replace_8c0ed4c7a2869076_dngsgbppbfhdnsfe.sql

This became frequent in CI after union system log tables (system.all_*) were enabled by default: they are created with CREATE OR REPLACE at the first flush after each startup, and test_reloading_storage_configuration restarts the server with kill -9 many times.

Now the metadata files of _tmp_replace_* tables are processed after all others, and if such a file has the same UUID as another table in the same database, only this metadata file is removed (the data is addressed by the UUID). For both the interrupted plain rename and the interrupted exchange, keeping the file under the final name gives a consistent state.

CI report: https://s3.amazonaws.com/clickhouse-test-reports/praktika.html?REF=master&sha=446d1bb88275a0bccf37e0187ee94a769e00dce8&name_0=MasterCI&name_1=Integration%20tests%20%28amd_asan_ubsan%2C%20db%20disk%2C%206%2F8%29
CI report: https://s3.amazonaws.com/clickhouse-test-reports/praktika.html?REF=master&sha=a537662755efda42670a7a6afa440a1519f5d736&name_0=MasterCI&name_1=Integration%20tests%20%28amd_asan_ubsan%2C%20db%20disk%2C%206%2F8%29

Related: #117943

Changelog category (leave one):

  • Bug Fix (user-visible misbehavior in an official stable release)

Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md):

Fix server startup failure with Mapping for table with UUID=... already exists after the server was killed during CREATE OR REPLACE TABLE when database metadata is stored on an object storage disk (database_disk with plain_rewritable).

🤖 Generated with Claude Code


Workflow [PR]
Sync PR [sync-upstream/pr/123902]

Version info

  • Merged into: 26.10.1.1906-master (included in 26.10 and later)

…atomic database disk

On a database disk without atomic renames (`plain_rewritable` object storage), the rename of the
temporary `_tmp_replace_*` table of `CREATE OR REPLACE` copies the metadata file and then removes
the source. If the server is killed in between, two metadata files refer to the same table, and
the server fails to start with `Mapping for table with UUID=... already exists`.

This became frequent since union system log tables are created by default with `CREATE OR REPLACE`
at the first flush after startup, and `test_reloading_storage_configuration` restarts the server
with `kill -9`.

Now the metadata files of `_tmp_replace_*` tables are processed after the others, and if such a
file has the same UUID as another table in the database, it is removed (only the metadata file,
as the data is addressed by the UUID).

CI report: https://s3.amazonaws.com/clickhouse-test-reports/praktika.html?REF=master&sha=446d1bb88275a0bccf37e0187ee94a769e00dce8&name_0=MasterCI&name_1=Integration%20tests%20%28amd_asan_ubsan%2C%20db%20disk%2C%206%2F8%29

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@clickhouse-gh

clickhouse-gh Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Workflow [PR], commit [11b751f]

Summary: ✅


AI Review

Summary

This PR defers _tmp_replace_* metadata files until after normal metadata is parsed, then drops a temporary file when it is only a duplicate reference to an already-loaded UUID after an interrupted CREATE OR REPLACE rename on a plain_rewritable metadata disk. That matches the failure mode described in the PR, the added stateless regression test covers the startup-recovery path, and I did not find an unresolved correctness or compatibility issue in the current patch.

Final Verdict
  • Status: ✅ Approve

LLVM Coverage Report

⚠️ No coverage measurement for commit 11b751f: incomplete coverage measurement: 1 of 21 shard profiles are missing: LLVM_COVERAGE_FILE_it_8.profdata.

@clickhouse-gh clickhouse-gh Bot added pr-bugfix Pull request with bugfix, not backported by default comp-database-engines Database engine implementations (e.g., Atomic/Replicated) and database-level behavior. labels Oct 4, 2026
@clickhouse-gh

clickhouse-gh Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Build profile diff (arm_release)

Comparing 11b751f4e with master 7cbb8ad28 (stripped binary size, per-symbol sizes and ThinLTO time; object sizes against the warmup build of 796d29517; compile times per translation unit against the most recent warmup build that recompiled it).

✅ No significant changes.

Binary sizes

programs/clickhouse-stripped: smaller than the master baseline by the known offset between the two builds, so the difference is not shown. A delta that differs from the offset by more than 50% of it is shown, in either direction.

The official master build is compiled with -g and a pull request build is not, and XRay counts debug instructions towards its instrumentation threshold, so master instruments thousands of functions more and its binary is ~0.4% larger no matter what the pull request does.

Compile time of recompiled translation units

7 translation units recompiled, 18 s compile time in total, 7 of them have a recent master baseline.

Job report

@alexey-milovidov alexey-milovidov left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The test is quite ad-hoc. Passable.
Instead of removing a stray file, we can rename it for safety, as we usually do.

However, this is good to merge.

@alexey-milovidov alexey-milovidov self-assigned this Oct 9, 2026
@alexey-milovidov
alexey-milovidov added this pull request to the merge queue Oct 9, 2026
Merged via the queue into master with commit f35daaf Oct 9, 2026
149 checks passed
@alexey-milovidov
alexey-milovidov deleted the fix-tmp-replace-duplicate-uuid-on-load branch October 9, 2026 04:12
@robot-clickhouse-ci-1 robot-clickhouse-ci-1 added the pr-synced-to-cloud The PR is synced to the cloud repo label Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp-database-engines Database engine implementations (e.g., Atomic/Replicated) and database-level behavior. pr-bugfix Pull request with bugfix, not backported by default pr-synced-to-cloud The PR is synced to the cloud repo

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants