A hands-on Proof of Concept demonstrating Kubernetes application failover with persistent state stored on an NFS-backed `ReadWriteMany` volume.
This repository demonstrates how Kubernetes recreates an application on another worker after a node failure while preserving the application's counter on shared NFS storage.
The playground uses four kind nodes:
kind-control-plane
kind-worker
└── Timer application at T0
kind-worker2
└── Timer failover target
kind-worker3
└── OpenEBS NFS server and persistent data
The timer increments a counter every ten seconds and writes both the counter and an event history to a PersistentVolumeClaim.
T0 Timer runs on worker1: counter=41, 42, 43
T1 worker1 is stopped
T2 Kubernetes recreates the timer on worker2
T3 The new pod mounts the same NFS PVC: counter=44, 45, 46
The pod changes, but the data remains available because it is stored on the third worker through NFS.
This Proof of Concept demonstrates:
- A multi-node kind cluster
- Dedicated application and storage workers
- Kubernetes Deployment self-healing
- Node affinity and preferred scheduling
- Node failure detection
- Pod eviction and rescheduling
- An OpenEBS NFSv4 server
- Dynamic volume provisioning with the official Kubernetes NFS CSI driver
- A
ReadWriteManyPersistentVolumeClaim - Persistent application state across pod recreation
- The difference between application failover and storage high availability
The following tools are required:
- Docker Desktop
- kind
- kubectl
- Git
Verify the tools:
docker version
kind version
kubectl version --clientCreate one control-plane node, two application workers and one dedicated NFS worker.
kind create cluster \
--name nfs-failover-playground \
--config kind/cluster.yamlVerify the nodes and their labels.
kubectl get nodes -L app-role,storage-roleExpected topology:
NAME APP-ROLE STORAGE-ROLE
nfs-failover-playground-control-plane
nfs-failover-playground-worker timer
nfs-failover-playground-worker2 timer
nfs-failover-playground-worker3 nfs
The Kubernetes node performs the NFS mount on behalf of the pod. Install the NFS client utilities inside both kind application nodes.
docker exec nfs-failover-playground-worker \
apt-get update
docker exec nfs-failover-playground-worker \
apt-get install -y nfs-common
docker exec nfs-failover-playground-worker2 \
apt-get update
docker exec nfs-failover-playground-worker2 \
apt-get install -y nfs-commonVerify the mount helper.
docker exec nfs-failover-playground-worker \
which mount.nfs
docker exec nfs-failover-playground-worker2 \
which mount.nfsExpected path:
/usr/sbin/mount.nfs
kubectl apply -f manifests/namespace.yamlDeploy the OpenEBS NFS server on the dedicated storage worker.
The previous
manifests/nfs/rbac.yamlmanifest is no longer used. The OpenEBS container only runs the NFS server; dynamic provisioning is handled separately by the official NFS CSI driver.
kubectl apply -f manifests/nfs/server.yamlWait for the server.
kubectl rollout status \
deployment/nfs-server \
-n nfs-failover-playground \
--timeout=180sVerify its placement.
kubectl get pods \
-n nfs-failover-playground \
-l app=nfs-server \
-o wideThe NFS pod must run on:
nfs-failover-playground-worker3
The Deployment is named nfs-server. Commands referring to
deployment/nfs-provisioner belong to the previous architecture and must not
be used.
Install the official driver, pinned to version v4.13.2.
curl -skSL \
https://raw.githubusercontent.com/kubernetes-csi/csi-driver-nfs/v4.13.2/deploy/install-driver.sh \
| bash -s v4.13.2 --Wait for its controller and node components.
kubectl rollout status \
deployment/csi-nfs-controller \
-n kube-system \
--timeout=180s
kubectl rollout status \
daemonset/csi-nfs-node \
-n kube-system \
--timeout=180sVerify the registered CSI driver.
kubectl get csidriver nfs.csi.k8s.iokubectl apply -f manifests/nfs/storage-class.yamlVerify it.
kubectl get storageclass nfs-rwxCreate the ReadWriteMany PVC.
kubectl apply -f manifests/workload/pvc.yamlWait until it is bound.
kubectl wait \
--for=jsonpath='{.status.phase}'=Bound \
pvc/timer-data \
-n nfs-failover-playground \
--timeout=180sDeploy the timer.
kubectl apply -f manifests/workload/timer.yamlWait until the Deployment is available.
kubectl rollout status \
deployment/timer \
-n nfs-failover-playground \
--timeout=180sVerify the initial placement.
kubectl get pods \
-n nfs-failover-playground \
-l app=timer \
-o wideThe scheduler should prefer:
nfs-failover-playground-worker
Follow the logs.
kubectl logs \
-n nfs-failover-playground \
deployment/timer \
-fExample output:
2026-07-26T10:00:00Z node=nfs-failover-playground-worker pod=timer-... counter=1
2026-07-26T10:00:10Z node=nfs-failover-playground-worker pod=timer-... counter=2
2026-07-26T10:00:20Z node=nfs-failover-playground-worker pod=timer-... counter=3
Press Ctrl+C after several increments.
Read the counter directly from the PVC.
kubectl exec \
-n nfs-failover-playground \
deployment/timer \
-- cat /data/counterStop the kind container hosting the timer.
docker stop nfs-failover-playground-workerWatch the node state.
kubectl get nodes -wIn another terminal, watch the pods.
kubectl get pods \
-n nfs-failover-playground \
-o wide \
-wThe stopped node first becomes NotReady. The timer pod is then evicted and recreated.
The workload has short not-ready and unreachable tolerations for a faster laboratory demonstration. The full transition can still take several tens of seconds because Kubernetes must detect the failed node.
Verify the new placement.
kubectl get pods \
-n nfs-failover-playground \
-l app=timer \
-o wideThe new timer pod must run on:
nfs-failover-playground-worker2
Follow the new pod logs.
kubectl logs \
-n nfs-failover-playground \
deployment/timer \
-fExpected behavior:
Before failure on worker1: counter=41
Before failure on worker1: counter=42
Before failure on worker1: counter=43
After failover on worker2: counter=44
After failover on worker2: counter=45
After failover on worker2: counter=46
The pod name and node name change, but the counter does not reset.
Display the persisted counter.
kubectl exec \
-n nfs-failover-playground \
deployment/timer \
-- cat /data/counterDisplay the latest persisted events.
kubectl exec \
-n nfs-failover-playground \
deployment/timer \
-- tail -n 20 /data/events.logThe log must contain entries written by pods running on both application workers.
docker start nfs-failover-playground-workerWait for it to become ready.
kubectl get nodes -wThe existing timer remains on worker2. Kubernetes does not automatically move a healthy running pod back to its preferred node.
To demonstrate preferred placement again, delete the current timer pod after worker1 is ready:
kubectl delete pod \
-n nfs-failover-playground \
-l app=timerThe replacement should return to the preferred worker while continuing from the same counter value.
A short overlap between the old and replacement timer pods is possible during
node recovery, eviction or manual pod deletion. For example, the persisted
history may temporarily contain interleaved entries from worker1 and
worker2:
node=worker2 pod=timer-old counter=18
node=worker1 pod=timer-new counter=19
node=worker2 pod=timer-old counter=20
node=worker1 pod=timer-new counter=21
This is a transient split-brain window. The Deployment uses replicas: 1 and
the Recreate strategy, but these settings express the desired state; they do
not provide distributed fencing. The NFS volume uses ReadWriteMany, so both
pods are technically allowed to mount and write to it while they overlap.
The timer performs a read, increment and write sequence without a distributed lock. Sequential counters during an overlap are therefore not guaranteed: two pods could read the same value and one increment could be duplicated or lost.
This behavior is acceptable for this failover playground because its purpose
is to demonstrate shared NFS persistence, not strict single-writer semantics.
A production workload requiring a single writer should use a distributed
lease or another fencing mechanism. Changing the PVC to ReadWriteOnce or
using a StatefulSet alone would not provide that guarantee.
The final state demonstrates application failover with persistent shared storage:
Timer pod on worker1
counter=41
counter=42
counter=43
worker1 stopped
Timer pod recreated on worker2
counter=44
counter=45
counter=46
Kubernetes maintains the Deployment's desired replica count. NFS maintains the application state independently of the timer pod.
The counter is not stored in the container filesystem.
Timer container
│
▼
timer-data PVC
│
▼
nfs-rwx StorageClass
│
▼
Kubernetes NFS CSI driver
│
▼
OpenEBS NFS server on worker3
│
▼
/var/local/nfs-failover
When worker1 is stopped:
- the original timer pod disappears;
- Kubernetes creates a new timer pod on
worker2; - the new pod mounts the same PVC;
- the counter is read from NFS;
- the next increment continues from the persisted value.
The NFS worker is deliberately kept online during this scenario.
This playground demonstrates application failover, not highly available storage.
The third worker is a single NFS server:
worker1 failure → application can fail over
worker2 failure → application can fail over
worker3 failure → NFS storage becomes unavailable
The NFS data is stored in a hostPath on the third kind node. It survives NFS pod recreation and application failover, but deleting the kind cluster deletes the node and its data.
Production environments should use a highly available storage backend instead of this laboratory architecture.
kubectl get nodes -L app-role,storage-role
kubectl get storageclass
kubectl get pvc \
-n nfs-failover-playground
kubectl get pv
kubectl get pods \
-n nfs-failover-playground \
-o wide
kubectl describe deployment timer \
-n nfs-failover-playground
kubectl describe pvc timer-data \
-n nfs-failover-playground
kubectl logs \
-n nfs-failover-playground \
deployment/nfs-serverInspect the NFS data directly inside the storage node.
docker exec nfs-failover-playground-worker3 \
find /var/local/nfs-failover \
-maxdepth 3 \
-type f \
-printVerify the NFS server and CSI driver.
kubectl get pods \
-n nfs-failover-playground \
-l app=nfs-server
kubectl logs \
-n nfs-failover-playground \
deployment/nfs-server
kubectl get pods \
-n kube-system \
-l app.kubernetes.io/name=csi-driver-nfs
kubectl get csidriver nfs.csi.k8s.ioIf the server logs report that a FILEPERMISSIONS_* variable is unbound,
reapply manifests/nfs/server.yaml. The manifest configures the export as
root:root with mode 0777, allowing the CSI driver to create a separate
subdirectory for each PVC.
Verify that the StorageClass uses nfs.csi.k8s.io.
kubectl get storageclass nfs-rwx -o yamlVerify that the NFS client is installed on both application workers.
docker exec nfs-failover-playground-worker \
which mount.nfs
docker exec nfs-failover-playground-worker2 \
which mount.nfsInspect pod events.
kubectl describe pod \
-n nfs-failover-playground \
-l app=timerKubernetes must first detect that the node is unreachable. The timer uses 15-second NoExecute tolerations, but node detection adds additional time.
Watch both resources:
kubectl get nodes -wkubectl get pods \
-n nfs-failover-playground \
-o wide \
-wDelete the workload.
kubectl delete -f manifests/workload/timer.yaml
kubectl delete -f manifests/workload/pvc.yamlDelete the NFS resources.
kubectl delete -f manifests/nfs/storage-class.yaml
kubectl delete -f manifests/nfs/server.yaml
curl -skSL \
https://raw.githubusercontent.com/kubernetes-csi/csi-driver-nfs/v4.13.2/deploy/uninstall-driver.sh \
| bash -s v4.13.2 --Delete the namespace.
kubectl delete -f manifests/namespace.yamlDelete the kind cluster.
kind delete cluster \
--name nfs-failover-playgroundAfter completing this playground you will understand:
- How a Deployment replaces a pod after a node failure
- How node affinity restricts and prefers workload placement
- How Kubernetes detects unreachable nodes
- How pod tolerations influence eviction timing
- How a
ReadWriteManyPVC provides node-independent application state - How NFS can expose shared storage to multiple Kubernetes workers
- Why a recreated pod can resume from persisted data
- Why application failover does not automatically imply storage high availability
- Why a healthy pod is not moved back to its preferred node automatically
-
Kubernetes Deployments
https://kubernetes.io/docs/concepts/workloads/controllers/deployment/ -
Kubernetes Persistent Volumes
https://kubernetes.io/docs/concepts/storage/persistent-volumes/ -
Kubernetes NFS Volumes
https://kubernetes.io/docs/concepts/storage/volumes/#nfs -
Kubernetes Taints and Tolerations
https://kubernetes.io/docs/concepts/scheduling-eviction/taint-and-toleration/ -
kind Configuration
https://kind.sigs.k8s.io/docs/user/configuration/ -
Kubernetes NFS CSI Driver
https://github.com/kubernetes-csi/csi-driver-nfs
OpenMind Systems Lab is an independent French non-profit association dedicated to research, experimental development and technical benchmarking in Cloud Native technologies.
Our mission is to produce practical, reproducible and educational Open Source Proofs of Concept covering Kubernetes, Platform Engineering, Distributed Messaging, Infrastructure Security and Artificial Intelligence.
GitHub Organization:
https://github.com/openmind-systems-lab
Made with ❤️ by OpenMind Systems Lab

