Conversation
With use_embedded_backup_restore_cluster set, BACKUP/RESTORE ... ON CLUSTER already makes every ClickHouse node write or read its own shards/N/replicas/M part of one shared destination, but clickhouse-backup metadata (metadata.json, per table .json, the .sql fixes before RESTORE) is node local, so only the node which ran the command was described. The node which runs create, create_remote, restore, restore_remote or delete local is now the initiator: it sends the same command with the new --embedded-on-cluster-worker flag to every other node of the cluster through its system.backup_actions table (INSERT INTO FUNCTION remote(...)) and waits for it, so every node needs a running clickhouse-backup server with api.create_integration_tables: true. A worker skips the BACKUP / RESTORE SQL and only describes, uploads, downloads and fixes its own shards/N/replicas/M prefix; workers fix their .sql before the initiator issues RESTORE ON CLUSTER, .backup and metadata.json stay initiator only. Per table .json moves under shards/N/replicas/M/metadata/ for every node, the backup is tagged embedded,cluster=<name> and the tag selects the layout, older flat embedded backups keep working. delete local is propagated, delete remote is not because the destination is shared. remote()/remoteSecure() follow clickhouse.secure because system.clusters has no secure column. TestFlows gets a second clickhouse_backup2 server bound to clickhouse2 and an "embedded on cluster" suite on sharded_cluster (create_remote, restore_remote, delete local, unavailable worker) with new RQ.SRS-013.ClickHouse.BackupUtility.EmbeddedBackup.OnCluster.* requirements. Fix #928 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
22.3 has no ON CLUSTER clause for BACKUP/RESTORE and 22.8 has no S3 backup engine, BACKUP/RESTORE became production ready in 23.3, the same gate TestEmbeddedS3 already uses. The bucket listing helper switches from format One (23.8+) to LineAsString so 23.3 passes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Coverage Report for CI Build 35762126642Coverage at 67.824% (no base build to compare)Details
Uncovered Changes
Coverage RegressionsNo coverage regressions found. Coverage Stats
💛 - Coveralls |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fix #928
Problem
With
use_embedded_backup_restore_clusterset,BACKUP/RESTORE ... ON CLUSTERalready makes every ClickHouse node write or read its ownshards/N/replicas/Mpart of one shared destination, but clickhouse-backup metadata (metadata.json, per table.json, the.sqlfixes beforeRESTORE) is node local, so only the node which ran the command was described and nothing drove the other nodes.Change
create,create_remote,restore,restore_remoteordelete localis the initiator: it sends the same command with the new--embedded-on-cluster-workerflag to every other node of the cluster through itssystem.backup_actionstable (INSERT INTO FUNCTION remote(...),remoteSecure()whenclickhouse.secure) and waits for it via the same table. Every node needs a runningclickhouse-backup serverwithapi.create_integration_tables: trueand the sameclickhouse.username/password/secure.BACKUP/RESTORESQL and only describes, uploads, downloads and fixes its ownshards/N/replicas/Mprefix. Order:create=BACKUP ON CLUSTER→ workers → own upload → wait;restore= workers (download + fix own.sql) → wait →RESTORE ON CLUSTER..backupandmetadata.jsonstay initiator only; on a sharedembedded_backup_diska worker does not overwrite the initiator'smetadata.json..jsonmoves undershards/N/replicas/M/metadata/for every node; the backup is taggedembedded,cluster=<name>and the tag selects the layout, older flat embedded backups keep working.delete localis propagated asdelete local --embedded-on-cluster-worker;delete remoteis not, the destination is shared and one call removes everything.system.clusters) behave as before.Tests
clickhouse_backup2server bound toclickhouse2, new suiteembedded on clusteronsharded_cluster(create_remote, restore_remote, delete local, unavailable worker) with newRQ.SRS-013.ClickHouse.BackupUtility.EmbeddedBackup.OnCluster.*requirements (requirements.pyregenerated withtfs, hence the large reformat diff).remote()rendering, tag parsing and layout selection.test/integration/utils.go: the_URLembedded variant now expects the prefixed.jsonpath.Verified locally:
go vet,GOFLAGS= make test, testflowsembedded on clusterandapi, integrationTestEmbeddedS3(all three embedded configs).Known limits
tables --local-backup/--remote-backupstill reads the flat.jsonpath (listing only); theBACKUP TABLElist is taken from the initiator as before; a failed worker leaves its local backup dir on its node.🤖 Generated with Claude Code