Skip to main content

MachinePool and MachineClaim

A MachinePool keeps spec.replicas Machines booted and unclaimed. A MachineClaim takes one of them over. The bind is a label change on an already-Running Machine, so it completes in one kairon-controller reconcile tick instead of a cold boot. The pool then boots a replacement.

Use it for agent sandboxes, CI runners, or anything that needs a VM now and can tolerate it having been booted a few minutes earlier.

Example​

kubectl apply -f examples/machinepool.yaml # pool "agents" + claim "job-42"
kaironctl get machinepools
kaironctl claim agents --label team=ml --ttl 1h
kaironctl get machineclaims

kaironctl claim creates the claim, waits for it to bind (--wait, default 60s) and prints the Machine name and bind time:

machineclaim/agents-claim-3fa91c bound to machine/agents-8c21d0e4 in 412ms

How it works​

Pool members carry these labels:

LabelValue
kairon.zyvor.dev/machinepoolpool name
kairon.zyvor.dev/machinepool-template-hashhash of spec.template
kairon.zyvor.dev/pool-statewarm or claimed
kairon.zyvor.dev/machineclaimclaim name, once claimed

Each tick, kairon-controller:

  1. Binds pending claims, oldest first, to the oldest Running warm member of their pool. The label patch carries the Machine's resourceVersion, so a member changed since the listing is skipped (409) instead of being bound twice. spec.labels from the claim are added to the Machine.
  2. Brings each pool to spec.replicas warm members. Claimed members don't count, so a claim triggers a replacement in the same tick. Warm members built from an older template are deleted and recreated at once (nobody is using them). Surplus members are trimmed, booting ones first.

A claim with no Running warm member stays Pending with a message like no Running warm machine in pool "agents" (2 warming) and binds as soon as one is ready.

Claim lifecycle​

FieldEffect
spec.poolNamePool in the claim's namespace. Immutable.
spec.labelsAdded to the Machine at bind.
spec.reclaimPolicyDelete (default): deleting the claim deletes the Machine. Retain: the pool and claim labels are removed and the Machine is kept as a plain Machine.
spec.ttlSecondsDeletes the claim this long after it binds, which releases the Machine per reclaimPolicy.
spec.egressEgress allowlist for the claimed Machine (see below).
status.phasePending, Bound, or Lost (the Machine was deleted out from under the claim).
status.bindMillisClaim creation to bind, in milliseconds.
status.egressPolicyName of the MachineNetworkPolicy enforcing spec.egress.

Per-claim egress​

spec.egress confines the claimed Machine to the listed destinations for as long as the claim is bound. Everything else is dropped at the VM's edge.

spec:
poolName: agents
egress:
allowFqdns: [pypi.org, files.pythonhosted.org]
allowPorts: ["443"]
kaironctl claim agents --allow-fqdn pypi.org --allow-fqdn files.pythonhosted.org --allow-port 443

At bind, kairon-controller creates a default-deny MachineNetworkPolicy named claim-egress-<claim>, targeting the Machine by name, with the fields below. Editing spec.egress updates the policy and removing it deletes the policy. Releasing the claim deletes the policy before the Machine is deleted or handed back.

FieldMeaning
allowFqdnsHostnames the guest may resolve and reach.
allowSNITLS server names allowed, exact or *.suffix.
allowCidrsDestination CIDRs allowed.
allowPortsDestination ports or ranges (443, 8000-8100).
allowDNSDNS query names allowed; empty falls back to allowFqdns.
allowIcmpAllow ICMP.

Deleting a pool deletes its warm members only. Claimed Machines belong to their claims and are left running.

Scaling​

MachinePool and MachineSet both have the scale subresource, so these work:

kubectl scale machinepool agents --replicas 10
kubectl scale machineset web --replicas 5
kaironctl scale machinepool agents --replicas 10

A HorizontalPodAutoscaler (or KEDA) can target either kind through the same subresource.

AI agents​

kaironctl mcp serve --allow-write exposes claim_machine, release_claim and delete_machine; list_machine_pools is read-only. See hermes-mcp.md.

Limits​

  • Readiness is status.phase == Running, not a guest-agent heartbeat.
  • Members boot through the normal Machine path (scheduler, quota, webhook), so a pool counts against MachineQuota for every warm member.
  • FluxVM's node-local warm pools (/v1/pools/{name}/claim) are a separate, lower-level mechanism; a MachinePool works with any backend.