A runnable, reproducible lab for operating ClickHouse on Kubernetes with the Altinity operator — the lifecycle way. Provisioning, replication, schema change, rolling upgrades, growing a disk without downtime, backups, and cross-region disaster recovery, each with something you can execute rather than a wiki page you hope is still accurate.
The topology it does all of that on is a hard one on purpose: two independent
Kubernetes clusters, one ClickHouse Keeper quorum shared across both, and one
logical ReplicatedMergeTree replication group spanning the two. Every
lifecycle operation is harder in that shape, which is what makes the answers
worth writing down.
Companion to percona-pg-lab
and kafka-lifecycle-lab.
Why lifecycle-centric? An operator does not delete operational work, it relocates it somewhere reviewable, versioned and repeatable. Promoting a DR region becomes a patched profile rather than a wiki page; growing a volume becomes
make expand-storagerather than a maintenance window; proving replication works becomesmake smokerather than a screenshot from last quarter.
Each of these is a thing teams believe about running ClickHouse until the day it matters. The repo makes each one falsifiable in a few minutes on a laptop:
| Belief | Reality |
|---|---|
remote_servers gives me a DR replica |
It gives you cluster(), remote() and Distributed. It moves zero rows. Replication is a table-engine property. |
| Two clusters, two Keepers, replication between them | Two Keeper quorums is two replication groups that look like one. No error, no lag metric — just silent divergence. |
| Growing a disk means a maintenance window | Only with the operator's default provisioner: StatefulSet. Operator-managed PVCs resize online, with ClickHouse running. |
| The DR region takes over automatically | With 2 Keeper members in primary and 1 in DR, losing the primary region leaves 1/3 — no quorum, no leader, no writes. That is Raft being correct. |
The lab builds the design that survives all of them, then breaks it on purpose.
Kubernetes cluster: PRIMARY Kubernetes cluster: DR
(own control plane, own operator) (own control plane, own operator)
┌────────────────────────────────┐ ┌────────────────────────────────┐
│ CHI ch-primary │ │ CHI ch-dr │
│ macros: shard=01 │ │ macros: shard=01 │
│ replica=replica- │ │ replica=replica-dr │
│ primary │ │ │
│ app user profile: default │ │ app user profile: standby │
│ (WRITES) │ │ (readonly=2) │
│ │ │ │
│ ReplicatedMergeTree ◄─────┼─ parts ─┼─────► ReplicatedMergeTree │
│ │ pulled │ │
│ keeper-1 keeper-2 │ │ keeper-3 │
└──────┬──────────────┬──────────┘ └──────┬─────────────────────────┘
│ │ │
└──────────────┴────────────────────────────┘
ONE Raft quorum, 3 members, majority = 2
One logical cluster. Two Kubernetes clusters. Writes in exactly one region at a
time. Full detail in docs/01-architecture.md.
A ClickHouse DR deployment's lifecycle has six phases, and each has a first-class answer here:
| Phase | What it covers | Where |
|---|---|---|
| Provision | Two clusters, two operators, no shared control plane | docs/02, local/, scripts/create-clusters.sh |
| Coordinate | The shared Keeper quorum and its arithmetic | docs/03, manifests/keeper/ |
| Replicate | Macros, znode paths, interserver endpoints | docs/04, manifests/clickhouse/, sql/ |
| Operate | Upgrades, scaling, schema change, the write gate | docs/05 |
| Grow | Online PVC expansion with no ClickHouse restart | docs/08, manifests/storage/ |
| Recover | Failover, failback, replica rebuild, and the region-loss path | docs/06, manifests/dr/ |
| Protect | Full + incremental backups, for the failures replication cannot cover | docs/09, manifests/backup/ |
| Observe | The four queries that tell you if DR is real | docs/07, manifests/observability/ |
.
├── local/ # two kind clusters standing in for two regions
├── manifests/
│ ├── keeper/ # the shared cross-region quorum (templated per member)
│ ├── clickhouse/ # one CHI template, rendered once per region
│ ├── storage/ # expandable StorageClass + pvc-autoresizer values
│ ├── backup/ # clickhouse-backup CronJob to object storage
│ ├── dr/ # emergency single-member Keeper for region loss
│ └── observability/
├── sql/ # the ReplicatedMergeTree schema + validation queries
├── scripts/ # deploy, verify, fail over, fail back, drill
├── docs/ # one guide per lifecycle phase + troubleshooting
└── Makefile # make up | smoke | drill | upgrade | expand-storage | destroy
Roughly 6 GB of RAM for Docker: two clusters, four nodes, three Keepers and two ClickHouse servers. No cloud account.
git clone https://github.com/dockndevai/clickhouse-lab
cd clickhouse-lab
make up # clusters -> operators -> shared Keeper -> both CHIs -> schema
make smoke # write to primary, watch it appear in DR, watch DR reject a writemake up is idempotent. When it finishes you have lab.events as a single
ReplicatedMergeTree with two replicas living in two Kubernetes clusters.
Then break it on purpose:
make drill # kill primary, fail over, write in DR, recover, fail back, verifyThe drill asserts that all 100 rows accepted by DR during the simulated outage are present in the primary after failback. If they are not, it fails loudly — which is the only useful thing a DR test can do.
make up deliberately uses kind's default storage and no backup target, so it
works with nothing but Docker. Two opt-in steps unlock the rest of the
lifecycle:
# storage: kind's local-path cannot expand, so install a CSI driver that can
make expandable-storage
export CH_STORAGE_CLASS=csi-hostpath-sc CH_PIN_LABEL=lab.dockndev.io/storage-node=true
make clickhouse && make schema && make smoke
make expand-storage SIZE=15Gi # online, zero ClickHouse restarts
# backups: MinIO as the object store, then the sidecar picks it up
make backup-storage
make clickhouse
make backup-drill # full -> incr -> drop BOTH -> restore -> resyncmake backup-drill covers the failure replication propagates rather than
protects against: it destroys the database in both regions and rebuilds
from an incremental backup, asserting row counts at every step.
Tear down:
make clean # remove ClickHouse + Keeper, keep the clusters
make destroy # delete both kind clustersPods being Running proves nothing. Three Keeper pods that cannot reach each
other are three separate single-node clusters, each losing an election to
itself forever, each perfectly Ready.
make keeper-health # exactly one leader, >= 2 members alive?
make repl-status # lag, queue depth, row counts per regionkeeper-health probes all three members over the same cross-region endpoints
Raft itself uses, so a green result also proves the network path.
make upgrade VERSION=24.8.14.10459 # rolling upgrade, passive region first
make storage-report # per-PVC usage across both regions
make expand-storage SIZE=75Gi # grow volumes online, no restart
make backup-full # full backup to object storage
make backup-incremental # incremental, diffed against the latest
make backup-drill # full -> incr -> destroy both -> restoreupgrade refuses to touch the primary until DR has upgraded, rejoined the
replication group and drained its queue — a version bump that leaves a replica
read-only should stop, not continue onto the only working one.
expand-storage refuses to patch volumes a StorageClass cannot grow, refuses a
shrink, refuses to cross a PVC's ceiling annotation, and refuses to run against
a StatefulSet-provisioned CHI without ALLOW_RESTART=1 — because an "online"
expansion that quietly restarts ClickHouse is worse than a planned window.
docs/08 covers why that setting decides
everything, and wires up pvc-autoresizer for the automated case.
kind's default
rancher.io/local-pathcannot expand at all, somake expandable-storageinstalls a CSI driver that can. The local run then exercises the whole path — patch, resize, capacity, CHI reconcile, ClickHouse still serving with zero restarts. Seedocs/08.
The manifests are the production shape; only the endpoints are laptop-sized. Three substitutions:
- NodePort → internal LoadBalancer. Every cross-region port here is a NodePort on a shared Docker bridge. In Azure/AWS that becomes an internal LoadBalancer over a VNet peering or VPC peering. Keep the constraint that the raft port and the interserver port cannot be translated — the process advertises the port it listens on.
- Node IPs → private DNS.
scripts/lib.shresolves kind node IPs; pointPRIMARY_NODE_IP/DR_NODE_IPat stable private names instead. - Secrets.
clickhouse-credentialsis created from a literal. Replace it with External Secrets, Key Vault or Secrets Manager.
Latency is the real constraint, not YAML: keep primary↔DR RTT under ~50 ms.
Past ~100 ms, Raft elections get unstable and every metadata commit pays the
round trip. docs/01-architecture.md covers when to
stop using native replication and put Kafka or S3 between the regions instead.
MIT — see LICENSE.