Skip to content

About

A hands-on Proof of Concept demonstrating Kubernetes application failover with persistent state stored on an NFS-backed `ReadWriteMany` volume.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Kubernetes NFS Failover Playground

A hands-on Proof of Concept demonstrating Kubernetes application failover with persistent state stored on an NFS-backed `ReadWriteMany` volume.

License Open Source Status Kubernetes Kubernetes Association


📖 Overview

This repository demonstrates how Kubernetes recreates an application on another worker after a node failure while preserving the application's counter on shared NFS storage.

The playground uses four kind nodes:

kind-control-plane

kind-worker
└── Timer application at T0

kind-worker2
└── Timer failover target

kind-worker3
└── OpenEBS NFS server and persistent data

The timer increments a counter every ten seconds and writes both the counter and an event history to a PersistentVolumeClaim.

T0  Timer runs on worker1: counter=41, 42, 43
T1  worker1 is stopped
T2  Kubernetes recreates the timer on worker2
T3  The new pod mounts the same NFS PVC: counter=44, 45, 46

The pod changes, but the data remains available because it is stored on the third worker through NFS.


🏗️ Architecture

schema


🎯 Objective

This Proof of Concept demonstrates:

  • A multi-node kind cluster
  • Dedicated application and storage workers
  • Kubernetes Deployment self-healing
  • Node affinity and preferred scheduling
  • Node failure detection
  • Pod eviction and rescheduling
  • An OpenEBS NFSv4 server
  • Dynamic volume provisioning with the official Kubernetes NFS CSI driver
  • A ReadWriteMany PersistentVolumeClaim
  • Persistent application state across pod recreation
  • The difference between application failover and storage high availability

📋 Prerequisites

The following tools are required:

  • Docker Desktop
  • kind
  • kubectl
  • Git

Verify the tools:

docker version
kind version
kubectl version --client

🚀 Deployment

Create the Kind Cluster

Create one control-plane node, two application workers and one dedicated NFS worker.

kind create cluster \
  --name nfs-failover-playground \
  --config kind/cluster.yaml

Verify the nodes and their labels.

kubectl get nodes -L app-role,storage-role

Expected topology:

NAME                                    APP-ROLE   STORAGE-ROLE
nfs-failover-playground-control-plane
nfs-failover-playground-worker          timer
nfs-failover-playground-worker2         timer
nfs-failover-playground-worker3                    nfs

Install the NFS Client on the Application Workers

The Kubernetes node performs the NFS mount on behalf of the pod. Install the NFS client utilities inside both kind application nodes.

docker exec nfs-failover-playground-worker \
  apt-get update

docker exec nfs-failover-playground-worker \
  apt-get install -y nfs-common

docker exec nfs-failover-playground-worker2 \
  apt-get update

docker exec nfs-failover-playground-worker2 \
  apt-get install -y nfs-common

Verify the mount helper.

docker exec nfs-failover-playground-worker \
  which mount.nfs

docker exec nfs-failover-playground-worker2 \
  which mount.nfs

Expected path:

/usr/sbin/mount.nfs

Deploy the Namespace

kubectl apply -f manifests/namespace.yaml

Deploy the NFS Server

Deploy the OpenEBS NFS server on the dedicated storage worker.

The previous manifests/nfs/rbac.yaml manifest is no longer used. The OpenEBS container only runs the NFS server; dynamic provisioning is handled separately by the official NFS CSI driver.

kubectl apply -f manifests/nfs/server.yaml

Wait for the server.

kubectl rollout status \
  deployment/nfs-server \
  -n nfs-failover-playground \
  --timeout=180s

Verify its placement.

kubectl get pods \
  -n nfs-failover-playground \
  -l app=nfs-server \
  -o wide

The NFS pod must run on:

nfs-failover-playground-worker3

The Deployment is named nfs-server. Commands referring to deployment/nfs-provisioner belong to the previous architecture and must not be used.


Install the Official Kubernetes NFS CSI Driver

Install the official driver, pinned to version v4.13.2.

curl -skSL \
  https://raw.githubusercontent.com/kubernetes-csi/csi-driver-nfs/v4.13.2/deploy/install-driver.sh \
  | bash -s v4.13.2 --

Wait for its controller and node components.

kubectl rollout status \
  deployment/csi-nfs-controller \
  -n kube-system \
  --timeout=180s

kubectl rollout status \
  daemonset/csi-nfs-node \
  -n kube-system \
  --timeout=180s

Verify the registered CSI driver.

kubectl get csidriver nfs.csi.k8s.io

Create the NFS StorageClass

kubectl apply -f manifests/nfs/storage-class.yaml

Verify it.

kubectl get storageclass nfs-rwx

🕒 Failover Scenario

T0 — Create the Persistent Counter

Create the ReadWriteMany PVC.

kubectl apply -f manifests/workload/pvc.yaml

Wait until it is bound.

kubectl wait \
  --for=jsonpath='{.status.phase}'=Bound \
  pvc/timer-data \
  -n nfs-failover-playground \
  --timeout=180s

Deploy the timer.

kubectl apply -f manifests/workload/timer.yaml

Wait until the Deployment is available.

kubectl rollout status \
  deployment/timer \
  -n nfs-failover-playground \
  --timeout=180s

Verify the initial placement.

kubectl get pods \
  -n nfs-failover-playground \
  -l app=timer \
  -o wide

The scheduler should prefer:

nfs-failover-playground-worker

Follow the logs.

kubectl logs \
  -n nfs-failover-playground \
  deployment/timer \
  -f

Example output:

2026-07-26T10:00:00Z node=nfs-failover-playground-worker pod=timer-... counter=1
2026-07-26T10:00:10Z node=nfs-failover-playground-worker pod=timer-... counter=2
2026-07-26T10:00:20Z node=nfs-failover-playground-worker pod=timer-... counter=3

Press Ctrl+C after several increments.

Read the counter directly from the PVC.

kubectl exec \
  -n nfs-failover-playground \
  deployment/timer \
  -- cat /data/counter

T1 — Stop the First Application Worker

Stop the kind container hosting the timer.

docker stop nfs-failover-playground-worker

Watch the node state.

kubectl get nodes -w

In another terminal, watch the pods.

kubectl get pods \
  -n nfs-failover-playground \
  -o wide \
  -w

The stopped node first becomes NotReady. The timer pod is then evicted and recreated.

The workload has short not-ready and unreachable tolerations for a faster laboratory demonstration. The full transition can still take several tens of seconds because Kubernetes must detect the failed node.


T2 — Observe the Application Failover

Verify the new placement.

kubectl get pods \
  -n nfs-failover-playground \
  -l app=timer \
  -o wide

The new timer pod must run on:

nfs-failover-playground-worker2

Follow the new pod logs.

kubectl logs \
  -n nfs-failover-playground \
  deployment/timer \
  -f

Expected behavior:

Before failure on worker1: counter=41
Before failure on worker1: counter=42
Before failure on worker1: counter=43

After failover on worker2: counter=44
After failover on worker2: counter=45
After failover on worker2: counter=46

The pod name and node name change, but the counter does not reset.


T3 — Verify the Persisted History

Display the persisted counter.

kubectl exec \
  -n nfs-failover-playground \
  deployment/timer \
  -- cat /data/counter

Display the latest persisted events.

kubectl exec \
  -n nfs-failover-playground \
  deployment/timer \
  -- tail -n 20 /data/events.log

The log must contain entries written by pods running on both application workers.


T4 — Restart the Failed Worker

docker start nfs-failover-playground-worker

Wait for it to become ready.

kubectl get nodes -w

The existing timer remains on worker2. Kubernetes does not automatically move a healthy running pod back to its preferred node.

To demonstrate preferred placement again, delete the current timer pod after worker1 is ready:

kubectl delete pod \
  -n nfs-failover-playground \
  -l app=timer

The replacement should return to the preferred worker while continuing from the same counter value.


Split-Brain Window During Restart

A short overlap between the old and replacement timer pods is possible during node recovery, eviction or manual pod deletion. For example, the persisted history may temporarily contain interleaved entries from worker1 and worker2:

node=worker2 pod=timer-old counter=18
node=worker1 pod=timer-new counter=19
node=worker2 pod=timer-old counter=20
node=worker1 pod=timer-new counter=21

This is a transient split-brain window. The Deployment uses replicas: 1 and the Recreate strategy, but these settings express the desired state; they do not provide distributed fencing. The NFS volume uses ReadWriteMany, so both pods are technically allowed to mount and write to it while they overlap.

The timer performs a read, increment and write sequence without a distributed lock. Sequential counters during an overlap are therefore not guaranteed: two pods could read the same value and one increment could be duplicated or lost.

This behavior is acceptable for this failover playground because its purpose is to demonstrate shared NFS persistence, not strict single-writer semantics. A production workload requiring a single writer should use a distributed lease or another fencing mechanism. Changing the PVC to ReadWriteOnce or using a StatefulSet alone would not provide that guarantee.


✅ Result

The final state demonstrates application failover with persistent shared storage:

Timer pod on worker1
counter=41
counter=42
counter=43

worker1 stopped

Timer pod recreated on worker2
counter=44
counter=45
counter=46

Kubernetes maintains the Deployment's desired replica count. NFS maintains the application state independently of the timer pod.


💡 Why Does the Counter Survive?

The counter is not stored in the container filesystem.

Timer container
      │
      ▼
timer-data PVC
      │
      ▼
nfs-rwx StorageClass
      │
      ▼
Kubernetes NFS CSI driver
      │
      ▼
OpenEBS NFS server on worker3
      │
      ▼
/var/local/nfs-failover

When worker1 is stopped:

  • the original timer pod disappears;
  • Kubernetes creates a new timer pod on worker2;
  • the new pod mounts the same PVC;
  • the counter is read from NFS;
  • the next increment continues from the persisted value.

The NFS worker is deliberately kept online during this scenario.


⚠️ Storage Limitation

This playground demonstrates application failover, not highly available storage.

The third worker is a single NFS server:

worker1 failure → application can fail over
worker2 failure → application can fail over
worker3 failure → NFS storage becomes unavailable

The NFS data is stored in a hostPath on the third kind node. It survives NFS pod recreation and application failover, but deleting the kind cluster deletes the node and its data.

Production environments should use a highly available storage backend instead of this laboratory architecture.


🔍 Explore the Resources

kubectl get nodes -L app-role,storage-role

kubectl get storageclass

kubectl get pvc \
  -n nfs-failover-playground

kubectl get pv

kubectl get pods \
  -n nfs-failover-playground \
  -o wide

kubectl describe deployment timer \
  -n nfs-failover-playground

kubectl describe pvc timer-data \
  -n nfs-failover-playground

kubectl logs \
  -n nfs-failover-playground \
  deployment/nfs-server

Inspect the NFS data directly inside the storage node.

docker exec nfs-failover-playground-worker3 \
  find /var/local/nfs-failover \
  -maxdepth 3 \
  -type f \
  -print

🛠️ Troubleshooting

PVC Remains Pending

Verify the NFS server and CSI driver.

kubectl get pods \
  -n nfs-failover-playground \
  -l app=nfs-server

kubectl logs \
  -n nfs-failover-playground \
  deployment/nfs-server

kubectl get pods \
  -n kube-system \
  -l app.kubernetes.io/name=csi-driver-nfs

kubectl get csidriver nfs.csi.k8s.io

If the server logs report that a FILEPERMISSIONS_* variable is unbound, reapply manifests/nfs/server.yaml. The manifest configures the export as root:root with mode 0777, allowing the CSI driver to create a separate subdirectory for each PVC.

Verify that the StorageClass uses nfs.csi.k8s.io.

kubectl get storageclass nfs-rwx -o yaml

Timer Pod Reports an NFS Mount Error

Verify that the NFS client is installed on both application workers.

docker exec nfs-failover-playground-worker \
  which mount.nfs

docker exec nfs-failover-playground-worker2 \
  which mount.nfs

Inspect pod events.

kubectl describe pod \
  -n nfs-failover-playground \
  -l app=timer

Timer Does Not Move Immediately

Kubernetes must first detect that the node is unreachable. The timer uses 15-second NoExecute tolerations, but node detection adds additional time.

Watch both resources:

kubectl get nodes -w
kubectl get pods \
  -n nfs-failover-playground \
  -o wide \
  -w

🧹 Cleanup

Delete the workload.

kubectl delete -f manifests/workload/timer.yaml
kubectl delete -f manifests/workload/pvc.yaml

Delete the NFS resources.

kubectl delete -f manifests/nfs/storage-class.yaml
kubectl delete -f manifests/nfs/server.yaml
curl -skSL \
  https://raw.githubusercontent.com/kubernetes-csi/csi-driver-nfs/v4.13.2/deploy/uninstall-driver.sh \
  | bash -s v4.13.2 --

Delete the namespace.

kubectl delete -f manifests/namespace.yaml

Delete the kind cluster.

kind delete cluster \
  --name nfs-failover-playground

📚 What You Will Learn

After completing this playground you will understand:

  • How a Deployment replaces a pod after a node failure
  • How node affinity restricts and prefers workload placement
  • How Kubernetes detects unreachable nodes
  • How pod tolerations influence eviction timing
  • How a ReadWriteMany PVC provides node-independent application state
  • How NFS can expose shared storage to multiple Kubernetes workers
  • Why a recreated pod can resume from persisted data
  • Why application failover does not automatically imply storage high availability
  • Why a healthy pod is not moved back to its preferred node automatically

📖 References


🏛 About OpenMind Systems Lab

OpenMind Systems Lab is an independent French non-profit association dedicated to research, experimental development and technical benchmarking in Cloud Native technologies.

Our mission is to produce practical, reproducible and educational Open Source Proofs of Concept covering Kubernetes, Platform Engineering, Distributed Messaging, Infrastructure Security and Artificial Intelligence.

GitHub Organization:

https://github.com/openmind-systems-lab


Made with ❤️ by OpenMind Systems Lab

About

A hands-on Proof of Concept demonstrating Kubernetes application failover with persistent state stored on an NFS-backed `ReadWriteMany` volume.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors