Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

clickhouse-lab

A runnable, reproducible lab for operating ClickHouse on Kubernetes with the Altinity operator — the lifecycle way. Provisioning, replication, schema change, rolling upgrades, growing a disk without downtime, backups, and cross-region disaster recovery, each with something you can execute rather than a wiki page you hope is still accurate.

The topology it does all of that on is a hard one on purpose: two independent Kubernetes clusters, one ClickHouse Keeper quorum shared across both, and one logical ReplicatedMergeTree replication group spanning the two. Every lifecycle operation is harder in that shape, which is what makes the answers worth writing down.

Companion to percona-pg-lab and kafka-lifecycle-lab.

Why lifecycle-centric? An operator does not delete operational work, it relocates it somewhere reviewable, versioned and repeatable. Promoting a DR region becomes a patched profile rather than a wiki page; growing a volume becomes make expand-storage rather than a maintenance window; proving replication works becomes make smoke rather than a screenshot from last quarter.


The beliefs this lab tests

Each of these is a thing teams believe about running ClickHouse until the day it matters. The repo makes each one falsifiable in a few minutes on a laptop:

Belief Reality
remote_servers gives me a DR replica It gives you cluster(), remote() and Distributed. It moves zero rows. Replication is a table-engine property.
Two clusters, two Keepers, replication between them Two Keeper quorums is two replication groups that look like one. No error, no lag metric — just silent divergence.
Growing a disk means a maintenance window Only with the operator's default provisioner: StatefulSet. Operator-managed PVCs resize online, with ClickHouse running.
The DR region takes over automatically With 2 Keeper members in primary and 1 in DR, losing the primary region leaves 1/3 — no quorum, no leader, no writes. That is Raft being correct.

The lab builds the design that survives all of them, then breaks it on purpose.


Architecture

     Kubernetes cluster: PRIMARY                Kubernetes cluster: DR
     (own control plane, own operator)          (own control plane, own operator)
  ┌────────────────────────────────┐         ┌────────────────────────────────┐
  │  CHI ch-primary                │         │  CHI ch-dr                     │
  │  macros: shard=01              │         │  macros: shard=01              │
  │          replica=replica-      │         │          replica=replica-dr    │
  │                  primary       │         │                                │
  │  app user profile: default     │         │  app user profile: standby     │
  │            (WRITES)            │         │            (readonly=2)        │
  │                                │         │                                │
  │      ReplicatedMergeTree ◄─────┼─ parts ─┼─────► ReplicatedMergeTree      │
  │                                │  pulled │                                │
  │  keeper-1        keeper-2      │         │  keeper-3                      │
  └──────┬──────────────┬──────────┘         └──────┬─────────────────────────┘
         │              │                            │
         └──────────────┴────────────────────────────┘
                  ONE Raft quorum, 3 members, majority = 2

One logical cluster. Two Kubernetes clusters. Writes in exactly one region at a time. Full detail in docs/01-architecture.md.


What's inside

A ClickHouse DR deployment's lifecycle has six phases, and each has a first-class answer here:

Phase What it covers Where
Provision Two clusters, two operators, no shared control plane docs/02, local/, scripts/create-clusters.sh
Coordinate The shared Keeper quorum and its arithmetic docs/03, manifests/keeper/
Replicate Macros, znode paths, interserver endpoints docs/04, manifests/clickhouse/, sql/
Operate Upgrades, scaling, schema change, the write gate docs/05
Grow Online PVC expansion with no ClickHouse restart docs/08, manifests/storage/
Recover Failover, failback, replica rebuild, and the region-loss path docs/06, manifests/dr/
Protect Full + incremental backups, for the failures replication cannot cover docs/09, manifests/backup/
Observe The four queries that tell you if DR is real docs/07, manifests/observability/
.
├── local/          # two kind clusters standing in for two regions
├── manifests/
│   ├── keeper/     # the shared cross-region quorum (templated per member)
│   ├── clickhouse/ # one CHI template, rendered once per region
│   ├── storage/    # expandable StorageClass + pvc-autoresizer values
│   ├── backup/     # clickhouse-backup CronJob to object storage
│   ├── dr/         # emergency single-member Keeper for region loss
│   └── observability/
├── sql/            # the ReplicatedMergeTree schema + validation queries
├── scripts/        # deploy, verify, fail over, fail back, drill
├── docs/           # one guide per lifecycle phase + troubleshooting
└── Makefile        # make up | smoke | drill | upgrade | expand-storage | destroy

Prerequisites

  • kubectl
  • kind and a running Docker
  • bash, envsubst (from gettext — brew install gettext on macOS)

Roughly 6 GB of RAM for Docker: two clusters, four nodes, three Keepers and two ClickHouse servers. No cloud account.


Quickstart

git clone https://github.com/dockndevai/clickhouse-lab
cd clickhouse-lab

make up        # clusters -> operators -> shared Keeper -> both CHIs -> schema
make smoke     # write to primary, watch it appear in DR, watch DR reject a write

make up is idempotent. When it finishes you have lab.events as a single ReplicatedMergeTree with two replicas living in two Kubernetes clusters.

Then break it on purpose:

make drill     # kill primary, fail over, write in DR, recover, fail back, verify

The drill asserts that all 100 rows accepted by DR during the simulated outage are present in the primary after failback. If they are not, it fails loudly — which is the only useful thing a DR test can do.

The two optional add-ons

make up deliberately uses kind's default storage and no backup target, so it works with nothing but Docker. Two opt-in steps unlock the rest of the lifecycle:

# storage: kind's local-path cannot expand, so install a CSI driver that can
make expandable-storage
export CH_STORAGE_CLASS=csi-hostpath-sc CH_PIN_LABEL=lab.dockndev.io/storage-node=true
make clickhouse && make schema && make smoke
make expand-storage SIZE=15Gi        # online, zero ClickHouse restarts

# backups: MinIO as the object store, then the sidecar picks it up
make backup-storage
make clickhouse
make backup-drill                    # full -> incr -> drop BOTH -> restore -> resync

make backup-drill covers the failure replication propagates rather than protects against: it destroys the database in both regions and rebuilds from an incremental backup, asserting row counts at every step.

Tear down:

make clean     # remove ClickHouse + Keeper, keep the clusters
make destroy   # delete both kind clusters

Is it actually working?

Pods being Running proves nothing. Three Keeper pods that cannot reach each other are three separate single-node clusters, each losing an election to itself forever, each perfectly Ready.

make keeper-health   # exactly one leader, >= 2 members alive?
make repl-status     # lag, queue depth, row counts per region

keeper-health probes all three members over the same cross-region endpoints Raft itself uses, so a green result also proves the network path.


Day-2 operations

make upgrade VERSION=24.8.14.10459   # rolling upgrade, passive region first
make storage-report                  # per-PVC usage across both regions
make expand-storage SIZE=75Gi        # grow volumes online, no restart
make backup-full                     # full backup to object storage
make backup-incremental              # incremental, diffed against the latest
make backup-drill                    # full -> incr -> destroy both -> restore

upgrade refuses to touch the primary until DR has upgraded, rejoined the replication group and drained its queue — a version bump that leaves a replica read-only should stop, not continue onto the only working one.

expand-storage refuses to patch volumes a StorageClass cannot grow, refuses a shrink, refuses to cross a PVC's ceiling annotation, and refuses to run against a StatefulSet-provisioned CHI without ALLOW_RESTART=1 — because an "online" expansion that quietly restarts ClickHouse is worse than a planned window. docs/08 covers why that setting decides everything, and wires up pvc-autoresizer for the automated case.

kind's default rancher.io/local-path cannot expand at all, so make expandable-storage installs a CSI driver that can. The local run then exercises the whole path — patch, resize, capacity, CHI reconcile, ClickHouse still serving with zero restarts. See docs/08.

Run on a real pair of clusters

The manifests are the production shape; only the endpoints are laptop-sized. Three substitutions:

  1. NodePort → internal LoadBalancer. Every cross-region port here is a NodePort on a shared Docker bridge. In Azure/AWS that becomes an internal LoadBalancer over a VNet peering or VPC peering. Keep the constraint that the raft port and the interserver port cannot be translated — the process advertises the port it listens on.
  2. Node IPs → private DNS. scripts/lib.sh resolves kind node IPs; point PRIMARY_NODE_IP / DR_NODE_IP at stable private names instead.
  3. Secrets. clickhouse-credentials is created from a literal. Replace it with External Secrets, Key Vault or Secrets Manager.

Latency is the real constraint, not YAML: keep primary↔DR RTT under ~50 ms. Past ~100 ms, Raft elections get unstable and every metadata commit pays the round trip. docs/01-architecture.md covers when to stop using native replication and put Kafka or S3 between the regions instead.


License

MIT — see LICENSE.

About

Runnable lab + production reference for operating ClickHouse on Kubernetes with the Altinity operator — the whole lifecycle: provisioning, ReplicatedMergeTree replication, upgrades, online PVC expansion, backups, and cross-region DR on a Keeper quorum shared across two clusters.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages