KAI Resource Management accepts GPU quotas that do not add up
I sent KAI Resource Management v0.18.2 a series of GPU quota changes by server dry-run, each one wrong on purpose. Projects were guaranteed more cards than their department holds. A guarantee sat above its own limit. A department was promised more cards than the node pool has. The API server accepted every change that was wrong only in its numbers.
KAI Resource Management (KRM) is the layer that adds node pools, departments and projects on top of the KAI Scheduler. A department holds a guarantee on a node pool. Its projects, each tied to a Kubernetes namespace, divide that guarantee between them. KRM turns those objects into the scheduler's queues. When the numbers do not add up, the change still syncs cleanly, and some team's guarantee stops being one without anything saying so.
So I wrote krm-quota-planner, an open-source tool that runs on your laptop. It draws the split, checks the arithmetic the cluster does not, and hands the change back for you to apply. This post covers what it checks and how it writes the change. It is not a guide to KRM itself.
What the cluster accepts
The dry-runs showed where KRM draws the line. It refused a queue on a node pool that does not exist, two queues on one pool inside the same object, and a second pool selecting the same node label.
Everything involving a number went through:
- Projects whose guarantees add up to more than their department's.
- A guarantee (
deserved) above the same queue'slimit, so the queue can never reach what it is promised. - Departments guaranteed more cards than the pool has.
- Negative values other than
-1, which means unlimited. - Two objects using the same queue name. Queue names are shared by the whole cluster, so the two would be one queue.
- Two projects claiming one namespace.
KAI's own Queue webhook has a parent and child check, but it can only warn, and it said nothing in any of these cases.
None of this produces a failure you can alert on. The scheduler carries on with numbers that no longer mean what the tenancy file says. The planner treats every item in that list as an error, and an error blocks the export.
Seeing the split before changing it
The planner is a local Node server with a page in the browser. With no flags it reads the node pools, departments and projects through your current kubectl context. Each pool gets a bar showing how its cards are divided between departments, and each department opens to its projects. Sliders change the guarantee, the limit and the over-quota weight.

The check messages name the queue and the size of the gap. The repo's tests use a small sample tenancy with one department of four cards. Give one of its projects an extra card without taking it from another, and the planner says:
platform-rtx: the projects are guaranteed 5 cards in total, but the department only has 4 cards — 1 card of those guarantees cannot be metWhen the cluster is reachable, the planner also compares the plan with what is running. A queue using more than its planned limit gets a warning rather than an error. So do non-preemptible workloads holding more than the planned guarantee. Nothing is evicted, so the new split does not take effect until those workloads stop. I would rather find that out before promising the cards to another team.
What does a safe quota change look like?
The tool never applies anything. From a cluster it prints kubectl patch commands for you to review and run. Moving one card from a project called workspace to one called benchmark, in the same sample tenancy, produces two commands. This is the first:
kubectl patch projects.kai.resources workspace --type=json -p '[{"op":"test","path":"/spec/queues/0/name","value":"workspace-rtx"},{"op":"test","path":"/spec/queues/0/resources/gpu/deserved","value":3},{"op":"replace","path":"/spec/queues/0/resources/gpu/deserved","value":2}]'Every replace is preceded by a test of the value the planner read. If someone changed that guarantee in the meantime, the patch fails instead of overwriting their change.
The commands are also ordered. The project giving up a card is patched before the project gaining one, and a department grows before its projects do. The commands are separate requests, so the order is what keeps each one acceptable at the moment it runs.
The catch is GitOps. If Argo CD or Flux owns the objects, a patch on the live cluster is put back at the next sync. The planner looks for the Argo CD and Flux tracking annotations and labels on what it reads, and warns you. In that case the change belongs in git.
One changed line per value
In git mode the planner reads the tenancy YAML from a ref with git show, not from your working tree. Tenancy files tend to be formatted by hand, with numbers in aligned columns and the reasoning in comments. Parsing the file and writing it back out would rewrite most of it.
The planner does not re-serialise. Each change replaces the bytes of one value and adjusts the padding after it, so the columns and comments stay where they were. The same one-card move, as a diff with the surrounding lines trimmed:
- gpu: {deserved: 3, limit: 3, overQuotaWeight: 1}
+ gpu: {deserved: 2, limit: 3, overQuotaWeight: 1}
...
- gpu: {deserved: 1, limit: 4, overQuotaWeight: 1}
+ gpu: {deserved: 2, limit: 4, overQuotaWeight: 1}The result is one commit on a new local branch, written with git plumbing, so your checkout and index are untouched. Nothing is pushed.
A quota often lives in more than one file. A profile can name the others:
- Variables that a verify script compares with the live queues. These are rewritten in the same commit.
- ResourceQuotas that mirror a queue's limit in whole cards, updated if you tick the box.
- Files you update by hand, such as a changelog or a runbook. The plan lists them as reminders.
What it leaves out
Version 0.1.0 changes GPU guarantees, limits, over-quota weights and queue priority. It adds node pools, departments, projects and queues. It does not delete or rename objects, move a project between departments, or edit CPU and memory quotas in the page. It does not create namespaces or label nodes for a new pool either, though it tells you which are still missing.
It is also not a gate. A planner helps the person making the change, and does nothing about a change somebody makes without it. If your tenancy lives in git, the same rules should run as a check on pull requests that touch it. The tool does not do that yet.
To try it you need Node 20 or later, plus kubectl access to nodepools, departments and projects in kai.resources:
git clone https://github.com/aleibovici/krm-quota-planner
cd krm-quota-planner
npm install
npm start # http://localhost:4780Without a cluster, npm run demo:generate followed by npm run demo starts the 160-GPU demo. The server binds to 127.0.0.1, and against the cluster it only runs kubectl get and an optional server-side dry-run.
The source is on GitHub under the MIT licence. Every rule was measured against KRM v0.18.2 and nothing else. If your version of KRM behaves differently from what I describe here, open an issue and tell me what it did.