Skip to content

Orchestrate embedded BACKUP/RESTORE ON CLUSTER across all nodes (#928) - #1571

Merged
Slach merged 3 commits into
masterfrom
issue-928
Sep 23, 2026
Merged

Slach merged 3 commits into
masterfrom
issue-928

Conversation

@Slach

@Slach Slach commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator

Fix #928

Problem

With use_embedded_backup_restore_cluster set, BACKUP/RESTORE ... ON CLUSTER already makes every ClickHouse node write or read its own shards/N/replicas/M part of one shared destination, but clickhouse-backup metadata (metadata.json, per table .json, the .sql fixes before RESTORE) is node local, so only the node which ran the command was described and nothing drove the other nodes.

Change

  • The node which runs create, create_remote, restore, restore_remote or delete local is the initiator: it sends the same command with the new --embedded-on-cluster-worker flag to every other node of the cluster through its system.backup_actions table (INSERT INTO FUNCTION remote(...), remoteSecure() when clickhouse.secure) and waits for it via the same table. Every node needs a running clickhouse-backup server with api.create_integration_tables: true and the same clickhouse.username/password/secure.
  • A worker skips the BACKUP/RESTORE SQL and only describes, uploads, downloads and fixes its own shards/N/replicas/M prefix. Order: create = BACKUP ON CLUSTER → workers → own upload → wait; restore = workers (download + fix own .sql) → wait → RESTORE ON CLUSTER. .backup and metadata.json stay initiator only; on a shared embedded_backup_disk a worker does not overwrite the initiator's metadata.json.
  • Per table .json moves under shards/N/replicas/M/metadata/ for every node; the backup is tagged embedded,cluster=<name> and the tag selects the layout, older flat embedded backups keep working.
  • delete local is propagated as delete local --embedded-on-cluster-worker; delete remote is not, the destination is shared and one call removes everything.
  • Single node clusters (no non-local rows in system.clusters) behave as before.

Tests

  • TestFlows: second clickhouse_backup2 server bound to clickhouse2, new suite embedded on cluster on sharded_cluster (create_remote, restore_remote, delete local, unavailable worker) with new RQ.SRS-013.ClickHouse.BackupUtility.EmbeddedBackup.OnCluster.* requirements (requirements.py regenerated with tfs, hence the large reformat diff).
  • Unit tests for the worker command rendering, remote() rendering, tag parsing and layout selection.
  • test/integration/utils.go: the _URL embedded variant now expects the prefixed .json path.

Verified locally: go vet, GOFLAGS= make test, testflows embedded on cluster and api, integration TestEmbeddedS3 (all three embedded configs).

Known limits

tables --local-backup/--remote-backup still reads the flat .json path (listing only); the BACKUP TABLE list is taken from the initiator as before; a failed worker leaves its local backup dir on its node.

🤖 Generated with Claude Code

Slach and others added 2 commits September 18, 2026 20:59
With use_embedded_backup_restore_cluster set, BACKUP/RESTORE ... ON CLUSTER
already makes every ClickHouse node write or read its own
shards/N/replicas/M part of one shared destination, but clickhouse-backup
metadata (metadata.json, per table .json, the .sql fixes before RESTORE)
is node local, so only the node which ran the command was described.

The node which runs create, create_remote, restore, restore_remote or
delete local is now the initiator: it sends the same command with the new
--embedded-on-cluster-worker flag to every other node of the cluster
through its system.backup_actions table (INSERT INTO FUNCTION remote(...))
and waits for it, so every node needs a running clickhouse-backup server
with api.create_integration_tables: true. A worker skips the BACKUP /
RESTORE SQL and only describes, uploads, downloads and fixes its own
shards/N/replicas/M prefix; workers fix their .sql before the initiator
issues RESTORE ON CLUSTER, .backup and metadata.json stay initiator only.

Per table .json moves under shards/N/replicas/M/metadata/ for every node,
the backup is tagged embedded,cluster=<name> and the tag selects the
layout, older flat embedded backups keep working. delete local is
propagated, delete remote is not because the destination is shared.
remote()/remoteSecure() follow clickhouse.secure because system.clusters
has no secure column.

TestFlows gets a second clickhouse_backup2 server bound to clickhouse2
and an "embedded on cluster" suite on sharded_cluster (create_remote,
restore_remote, delete local, unavailable worker) with new
RQ.SRS-013.ClickHouse.BackupUtility.EmbeddedBackup.OnCluster.*
requirements.

Fix #928

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
22.3 has no ON CLUSTER clause for BACKUP/RESTORE and 22.8 has no S3
backup engine, BACKUP/RESTORE became production ready in 23.3, the same
gate TestEmbeddedS3 already uses. The bucket listing helper switches
from format One (23.8+) to LineAsString so 23.3 passes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@coveralls

coveralls commented Sep 22, 2026 •

Copy link
Copy Markdown

Coverage Report for CI Build 35762126642

Coverage at 67.824% (no base build to compare)

Details

  • Coverage remained the same as the base build.
  • Patch coverage: 91 uncovered changes across 10 files (296 of 387 lines covered, 76.49%).
  • No coverage regressions found.

Uncovered Changes

File Changed Covered %
pkg/backup/embedded_cluster.go 177 147 83.05%
pkg/server/server.go 48 28 58.33%
pkg/backup/create.go 25 16 64.0%
pkg/backup/restore.go 30 22 73.33%
pkg/backup/upload.go 26 19 73.08%
pkg/backup/delete.go 26 20 76.92%
pkg/backup/download.go 18 14 77.78%
pkg/backup/restore_remote.go 13 9 69.23%
pkg/backup/create_remote.go 10 8 80.0%
pkg/backup/backuper.go 2 1 50.0%
Total (12 files) 387 296 76.49%

Coverage Regressions

No coverage regressions found.


Coverage Stats

Coverage Status
Relevant Lines: 27437
Covered Lines: 18609
Line Coverage: 67.82%
Coverage Strength: 35533.34 hits per line

💛 - Coveralls

@Slach
Slach merged commit 3a0692a into master Sep 23, 2026
30 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

implements embedded BACKUP/RESTORE ON CLUSTER approach

2 participants