A single script that collects what is needed to plan a migration of persistent volumes to a new storage backend.
Reads your cluster and writes a report. It runs get and describe — nothing
else. It creates no objects, changes no configuration and deletes nothing.
It does not read Secrets or their contents, ConfigMap data, application data, or anything inside your volumes. Only resource metadata: how many volumes there are, how big, which storage classes, which workloads use them, and which platform components are in play.
Everything it collects is written to a directory you can inspect before sending it anywhere.
oc or kubectl, and a context with read access. No python, no jq, no
internet access, no installation.
Read access to cluster-scoped resources (nodes, storage classes, CRDs) gives the complete picture. It runs with less and reports what it could not see.
curl -fsSLO https://raw.githubusercontent.com/CalzDevOps/tp-assessment/main/assess-cluster.sh
sha256sum assess-cluster.sh
bash assess-cluster.sh
Options:
bash assess-cluster.sh -n ns1,ns2 limit the workload scan to namespaces
bash assess-cluster.sh -o /path where to write the output
bash assess-cluster.sh --probe also test whether a helper pod is admitted
bash assess-cluster.sh --deep also sample file counts inside volumes
--probe is the only part that creates anything: one short-lived pod, in one
namespace, which prints a line and is deleted. It answers a question that
cannot be answered by reading policy — whether a pod with elevated privileges
would be admitted here — because admission evaluates the requesting user and
the service account together. Leave it off if you would rather answer that
from your own policy documentation.
--deep runs find and du inside pods that already mount the volumes. It
reads nothing but sizes and counts. It is off by default because it takes time
on large filesystems.
assessment-<context>-<date>/
summary.txt the readable summary
raw/ full output of every command, one file each
assessment-<context>-<date>.tgz
Review it, then send the .tgz back.
Capacity determines how long the bulk copy takes. The number of files determines how long the final synchronisation takes, and the two are not related: a 500 GB database and a 500 GB document store with twelve million files behave completely differently at cutover.
This is the most common reason a migration plan turns out to be wrong, so if only one thing gets measured by hand, measure this on the largest volumes:
kubectl exec -n <namespace> <pod> -- sh -c "find /data -type f | wc -l; du -sh /data"
Some things do not come from a command and need to be answered by whoever runs the applications:
- Acceptable downtime per application group, not one number for everything
- How much data changes per day on the large volumes
- How long the source storage must be kept intact after each move
- Change freeze periods
- Whether there is a non-production environment with the same storage to rehearse on