How we patch a production GPU cluster with an AI agent
We patch our production AI cluster one node at a time. No node is touched until the previous one has passed a fixed set of checks. The cluster is a Kubernetes control plane in front of a fleet of multi-GPU NVIDIA servers, which serve large open-weight models and notebook workspaces. The operator that runs the round is an AI agent, working from intent documents in Git.
A kernel fix only takes effect when the node reboots, so every round is a sequence of planned reboots. This post covers how we run that OS and kernel round, why each step is there, and where the agent and the people fit. Kubernetes version upgrades are a separate change and we've left them out.
Intent documents first, then the agent
Every component of the platform has its own folder in a Git repository. Each folder holds an intent document covering why the component exists, the properties it must keep, what is out of scope, and the open questions. Next to it are a runbook, a changelog and a verify.sh that exits non-zero when something is wrong. For the OS layer, the intent says kernel upgrades are always deliberate. They go one node at a time, each node's checks must be green before the next, and a fleet-wide kernel-consistency check closes the round.
We use Anthropic's Claude Code as the operator. It reads the intent and the lessons from our test environment, then inspects the live nodes. From that it writes the plan: which packages move, what runs where, and what each drain will evict. Commands and versions come from the hosts as they are today, and a plan is re-checked against the live hosts before it runs.
The agent also works under standing rules for production. It defaults to caution, snapshots before any significant change, confirms with a person before anything destructive or hard to reverse, and treats a task as unfinished until the docs are updated. People review the plan and own the decisions that carry risk: accepting residual risk, taking hypervisor snapshots, and merging changes. The agent runs each window and its checks. It then writes what it found back into the repo before the next window starts.
Where the live system contradicts the intent, the agent flags it against the document instead of working around it. Our intent assumes Canonical Livepatch covers kernel CVEs between rounds. The agent found it reports nothing to apply on our kernel line, and that is now an open question against the intent. Every script still runs without the agent.
Order is a risk decision
- Control-plane nodes first. They carry the least risk and prove the drain procedure before anything expensive is involved.
- The node holding the Postgres primary and the API virtual IP goes last within the control plane. That way the primary moves once, onto a node that is already patched.
- The least critical GPU server goes next, as a canary. It proves the new kernel boots on bare metal and NVIDIA's driver builds against it. If it fails, the rest of the fleet stays put.
- The remaining GPU servers follow, one at a time.
We never run two windows at once. CloudNativePG switches the database over when the primary's node is cordoned, so nobody promotes by hand.

Preflights before anything is installed
Most of the risk sits in what apt decides to do, so we find out before the drain.
- Confirm the holds on containerd, kubelet, kubeadm, kubectl and cri-tools. Kubernetes must not move as a side effect of OS patching.
- Run
apt-get -d full-upgradewhile workloads still run, which shortens the outage. - Run
apt-get -s full-upgrade. Any removal stops the round. - Snapshot what an OS round can break: etcd,
/etc/kubernetesand host configuration.
The drain happens before the install, because needrestart can restart services mid-upgrade and we would rather that happen on an empty node. The old kernel stays installed as the rollback. On the GPU servers, unattended upgrades can't touch the kernel or NVIDIA packages. The driver must match whatever kernel boots.
A trial boot for bare metal
A kernel that fails to boot on a bare-metal GPU server is a bigger problem than on a VM. So we boot the new kernel once, and the old kernel stays the default:
# Old kernel stays the GRUB default; the new one is queued for the next boot only
echo "GRUB_DEFAULT=\"$OLD\"" | sudo tee /etc/default/grub.d/99-trial-boot.cfg
sudo update-grub
sudo grub-reboot "$NEW" && sudo systemctl rebootWith kernel.panic=10 set, a panic once sysctls have loaded sends the node back to the known-good kernel. So does any unplanned power cycle. Once the checks are green, the drop-in goes and the new kernel becomes the default. $OLD and $NEW must be GRUB entry IDs, not menu titles. For a kernel under Advanced options, that means the full path, such as gnulinux-advanced-<uuid>>gnulinux-<version>-advanced-<uuid>. The $menuentry_id_option lines in /boot/grub/grub.cfg list them.
A trial boot doesn't help with a hang or a very early panic. That still needs out-of-band console access.

What "done" means for each node
A node is finished when its checks pass, not when it answers SSH.
- Node verify: the kernel, network bonds and RDMA rails, and sysctls, with a separate check that identity lookups work.
- The gate before the next cordon: etcd has quorum and a leader, every Postgres instance is ready, the admission webhook is at full replicas, and the stack health check shows no new failures.
- GPU acceptance: the driver is loaded, every GPU is visible, Fabric Manager is up, and a real completion comes back through the model.
Uncordon before the GPU tests, or their pods can't schedule. If every GPU is allocated, a smoke test fails on capacity, which is not a regression.
A reboot is also when some components re-read their configuration for the first time in weeks. The Kubernetes API server fetches its OIDC issuer only at start, so a DNS change can break logins long after it was made. The cluster verify now checks the OIDC authenticator for that reason.
When is patching finished?
Old kernels stay for a 48-hour soak before apt autoremove, and a fleet check confirms every node runs the same kernel. NVIDIA's driver container ships its own userland, which we can't patch without rebuilding the vendor's image. We record a risk acceptance and review it at every driver change.
NCSC NZ's patching minimum standard for government agencies asks for critical patches within two days on external-facing systems and two weeks on internal ones. A reboot round across a GPU fleet only fits those windows if the runbook exists and has been run. A monthly round won't meet a two-day deadline by itself. It keeps the runbook practised, so an urgent round between windows follows the same steps. With the intent documented and an agent to do the work, we're proposing a monthly window that follows this runbook every time.
Next, we want the agent to trigger the trial-boot fallback on purpose on the canary, before a bad kernel does it for us. If you're planning something similar, message me and I'll walk you through the setup.