What ships today
Full feature inventory. The root README is the short product landing; guides linked below remain authoritative for how-to detail.
Machine lifecycle & placement
- Machine CRD — CPU, memory, image, network, power state, volumes, DRA device claims
MachineSet— a Deployment/ReplicaSet-shaped fleet of identical Machines from one template, withRollingUpdate/Recreaterollout — guideMachineInstanceType— a reusable named CPU/memory shape (spec.instanceTypeName), resolved intospec.resourcesonce instead of inlining it every time — guide- NUMA topology, CPU set, hugepages (
spec.resources.numaNode/.cpuSet/.hugepages, qemu-only) — direct passthroughs to FluxVM's own existing support, refused with a clear error on a non-qemu backend rather than silently ignored — guide - Real CPU pinning (
spec.resources.cpuPinning, qemu-only) — exclusive host-core allocation, a real scheduler capacity model against operator-assertedkairon.zyvor.dev/pinnable-cpusnode labels — atomically patched alongside node assignment, applied via the same cgroupcpuset.cpuswritespec.resources.limitsalready uses — guide - Windows guests — legacy-BIOS Windows Server/10 with a pre-built, virtio-driver + cloudbase-init image works today (
spec.cloudInitunchanged);spec.security.secureBoot/.tpm(Windows 11) now pass through to FluxVM's own real OVMF Secure Boot (qemu-only) and emulated TPM 2.0 (qemu or cloud-hypervisor) — Kairon enforces the same backend restriction FluxVM's own scheduler does, refusing an unsupported combination with a clear error rather than relaying a bare FluxVM 400. A real Secure Boot chain still needs the FluxVM node itself configured with an OVMF vars template (an operator responsibility, not something Kairon manages) — guide MigrationPolicy— scopes migration bandwidth/concurrency to Machines matching a selector, independent of the globalmigration.maxConcurrentPerNode/maxConcurrentClustercaps — guide- PVC-backed boot disk —
spec.volumes[0]resolves through a BoundPersistentVolumeClaimto a real host directory (hostPath/localPersistentVolumes directly, or a network-block volume via Kairon's own first-cut CSI driver) instead of a hand-placed image file — guide · CSI guide - Multi-volume (first cut) —
spec.volumes[1+]attach as FluxVM virtiofs shared folders (QEMU); boot remainsvolumes[0]/spec.image - CSI node CHAP (
node.csi.chap.enabled, opt-in) — resolve a PV'snodeStageSecretReffor iSCSI login, Secrets allowlisted to the release namespace only — CSI guide - Third-party CSI client (
node.thirdPartyCSIDrivers, first cut) —kairon-nodeacts as its own generic CSI client against an operator-allowlisted third-party driver's socket (validated conceptually against Ceph-CSI/RBD), includingattachRequireddrivers viaVolumeAttachmentand node-stage/publish secrets from one allowlisted namespace — guide - Atlas-provisioned volumes (
atlas.enabled, first cut) —spec.volumes[].atlasasks the Atlas storage control plane for a disk by size and policy; the controller creates it, waits on its job, writes the PVC intoclaimName, schedules only once Ready, and deletes it after node runtime cleanup.atlas.mode: rbdboots the Atlas image in place through FluxVMceph-rbd-in-place(node pool allowlist, image name bound to the Machine) — guide - Image import (
spec.image.source, opt-in vianode.imageCacheDir) —kairon-nodedownloads a remote qcow2/raw image URL into its own digest-keyed cache instead of requiring a hand-placed file, shared across every Machine naming the same digest — guide - Image catalog (
spec.image.catalogName, admin-only for mutations, API-only) — reference a named, checksummed (optionally signed) FluxVM image-catalog alias instead of a raw path; node-scoped register/rename/clone/export/read-only API — guide - Egress check (admin-only, API-only) — ask a node whether a sandbox's outbound request to a host would be allowed under its static egress allowlist, with any matched credential-vault secret reported only as a boolean — guide
- Warm pools (admin-only, API-only) — create a FluxVM pool of pre-booted,
PausedVMs for instant claiming instead of a cold create; node-scoped create/list/get/delete/claim API — guide - Runtime diagnostics (freeze/thaw admin-only, rest any-operator, API-only) — a node's real FluxVM capability manifest, a Machine's cgroup-derived PSI pressure stats and effective host CPU set, and a kernel-level cgroup freeze/thaw distinct from
spec.powerState: Paused— guide - Network observability (any-operator) — a Machine's fully-resolved effective network policy, real drop-reason/flow/stats data straight from FluxVM's eBPF dataplane —
kaironctl network flows|drop-reasons|stats|effective, dashboard Network panel, and the uiapi relays — guide ("Troubleshooting" section) - eBPF dataplane (recommended) —
spec.network.dataplaneMode: ebpf(production profile / docs default recommendation); Kairon declares, FluxVM owns TC/eBPF — reference - VM edge —
spec.network.antiSpoof/learnIP/qosplus a policy'sallowSNI/allowDNS, enforced by FluxVM's TC program (dataplane schema 12); policy allow lists also turn the edge on for a plain tap Machine; netns Machines get a generated MAC; attributed drops viakaironctl network dropsandkairon_net_drops_total; packet capture with pcap download (kaironctl network capture --output); conntrack moved on live migration — reference - AI agents over MCP —
kaironctl mcp serve: Model Context Protocol tools for Hermes Agent and other MCP clients (Machines, policies, VM-edge network data; power, snapshot and capture only with--allow-write) — guide - Placement — least-loaded scheduling across Ready, capable-labeled nodes, deterministic tie-break, required (hard)
spec.placement.affinity/antiAffinity, weighted soft scoring (preferredAffinity/preferredAntiAffinity,topologySpreadConstraints), a best-effort DRA topology-awareness hint, andtopologySpreadConstraints.maxSkewhard enforcement viawhenUnsatisfiable: DoNotSchedule— guide - Scheduling priority (
spec.priority, first cut) — when a reconcile tick can't fit every still-pending Machine (not enough eligible nodes, or aMachineQuotaat its cap), higher-priorityMachines are attempted first; purely an admission-order signal for that tick's pending Machines, never preemption of an already-scheduled one however low its own priority —kaironctl create/create machineset --priority N,kaironctl edit machine NAME --priority N— guide - DRA → VFIO — an allocated
ResourceClaim's PCI BDF against the node'svfio_devicesadministrator allowlist, fail-closed (an empty allowlist denies all passthrough); device-type-agnostic, so this also covers SR-IOV NIC passthrough once a VF is bound tovfio-pcion the host — the same mechanism GPU passthrough already uses, not a Multus integration (Machines have no Pod for Multus to attach to) — guide - CPU/memory hotplug — grow a running Machine's
spec.resourcesvia FluxVM's real QMPdevice_add/object-add, no reboot — guide - Pause and resume (
spec.powerState: Paused) — suspend a running Machine's guest CPUs via FluxVM's real QMPstop/cont(backend-agnostic), RAM and device state fully resident, Kairon'svirtctl pause/unpauseequivalent —kaironctl pause/resume, or the dashboard — guide - Halt (
spec.powerState: Halted) — power a Machine's VMM process off while FluxVM keeps its own VM record/disk intact, freeing node capacity likeStoppedbut resuming via FluxVM's own kept config instead of a full Kairon-side recreate —kaironctl halt, or the dashboard — guide - Live host resource limits (
spec.resources.limits) — a real, kernel-enforced cgroup v2 cap (CPU quota %, memory ceiling, I/O weight, PID count) on a running Machine's own VMM process, freely raisable/lowerable at any time, backend-agnostic — guide - Live resource usage (
status.resourceUsage) — real, cgroup-derived CPU/memory/disk-I/O usage refreshed every reconcile tick, shown inline in the dashboard next to each Machine — guide kaironctl top machines/top nodes— akubectl top-style live usage view over that samestatus.resourceUsagedata, no new metrics pipeline:top machinesprints each Machine's own CPU%/memory/disk-I/O,top nodesrolls those same per-Machine samples up byspec.nodeName(model.AggregateUsageByNode) into a per-host hotspot viewkaironctl get machinesalone can't answer without manually grouping rows. The same rollup backs the dashboard's Nodes page (GET /api/v1/nodes/usage) below — guide- Guest agent (
spec.guestAgent) — realqemu-guest-agent-reportedstatus.guestIP, including foruser/SLIRP networking, which has no DHCP lease to parse at all — guide - Guest exec (
spec.guestAgent.enabled, admin-only) — run a one-shot command inside the guest viakairon-ui's dashboard, no SSH/console needed — realqemu-guest-agentguest-exec, synchronous, exit code + stdout/stderr — guide - Text console (
spec.guestAgent.console) — an interactive shell in-browser (xterm.js), Kairon'svirtctl consoleequivalent, working on every backend — requires FluxVM's own proprietaryfluxvm-guest-agentbaked into the guest image — guide - Guest file access (
spec.guestAgent.console, admin-only) — read or write a file inside aRunningguest at runtime, no SSH/shared volume needed — samefluxvm-guest-agentchannel as the text console — guide - Backend-agnostic guest exec (
spec.guestAgent.console, admin-only, API-only) — run a command over FluxVM's own vsock agent instead ofqemu-guest-agent— the only guest-exec path that works on Cloud Hypervisor/Firecracker/FluxVm-backend sandboxes, not just QEMU — guide - Sandboxes (
spec.sandbox, admin-only, API-only) — run a Machine on FluxVM's own lightweight, fast-boot in-tree hypervisor track for short-lived ephemeral workloads, optionally from a pre-built template; includes an HTTP proxy relay straight into the guest's own web server — guide - Machine logs —
kubectl logs/-fequivalent: the VM's real captured serial console output, with a live-tailing Follow toggle, same any-operator authorization as the VNC console — guide - VM-state snapshot/restore (admin-only, API-only) — a full hypervisor-level checkpoint of a running Machine's RAM/CPU/device state via FluxVM's real
savevm/Cloud Hypervisor snapshot, restored back into the same Machine in place; distinct fromMachineSnapshot's disk-content-only CSI snapshot — guide
Migration
- Cold migrate & evacuate — stop → reassign → restart, restart-safe in the API
- Secure live handshake — TLS 1.3 mTLS
prepare → transfer → commit, never a user-suppliedtcp:URI - Split-brain guards — rollback on transfer failure,
NeedsRecoveryon ambiguous commit, adopt-only cutover - A real migration adapter ships —
cmd/kairon-migration-adapter-fluxvmagainst real FluxVM endpoints, not just a test double - Node fencing is detection + operator-attested action, never a guess —
kairon-controllerflags a Machine's nodeNodeUnreachableevery reconcile tick (Node-Ready, and whennode.livenessLease.enabledis on, Ready + stale agent Lease →AgentLivenessStale) but never reschedules it itself (that risks running the same VM twice if the node isn't actually dead);kaironctl fence --reason ...is the explicit, attested step that clears it for rescheduling — same shape asNeedsRecovery. The same opt-in Lease also lets fence refuse when kairon-node still looks alive despiteNodeUnreachable=True. Opt-inkairon.zyvor.dev/storage-domain/network-domainnode labels also let migration preflight block a target when it's confirmed to not share storage/network with the source — guide
Guarding the fleet
MachineDisruptionBudget—kaironctl evacuatethrottles itself againstminAvailable/maxUnavailableinstead of taking a whole node's Machines at once;kairon-controllernow reconciles realstatusevery tick too (kaironctl get budgets/kubectl get mdb), purely observational.kaironctl create budget NAME --selector k=v (--min-available X | --max-unavailable X)/edit budget NAME [--selector k=v] [--min-available X] [--max-unavailable X]create and mutate one directly — no more dropping tokubectl apply/YAML just to stand one up — guide- Cordon-triggered automatic evacuation (
controller.cordonEvacuation.enabled, opt-in) —kairon-controlleritself migrates every Machine off a Node the moment it's cordoned, budget-throttled exactly likekaironctl evacuate, mirroring KubeVirt'sLiveMigrateIfPossible— guide MachineQuota— a namespace-scopedmaxMachines/maxTotalCpu/maxTotalMemorycap.kaironctl create quota NAME [--max-machines N] [--max-total-cpu N] [--max-total-memory SIZE]/edit quota NAME [flags]create and mutate one directly — guide- Admission webhook (
webhook.enabled, opt-in) — both of the above can now be enforced at admission, not just in the reconcile loop orkaironctl: aMachinecreate that would blow a quota, or aMachineMigrationcreate that would violate a budget, gets rejected outright instead of just parkedPendingor silently allowed. See Guarding the fleet below. - Network Fabric —
spec.network,MachineNetworkPolicy,NetworkSecurityGroup, Service Fabric VIP membership → FluxVM eBPF edge — reference - Cilium CNI (opt-in, three paths) —
spec.network.dataplaneMode: ciliumrequests FluxVM's cilium dataplane;spec.network.ciliumAttach(Helmnetwork.ciliumAttach.enabled) reconciles a cluster-scopedCiliumExternalWorkloadand writesstatus.network.ciliumpluspodUID;MachineNetworkPolicy.spec.cilium.sync(Helmnetwork.ciliumPolicySync.enabled) upserts a namespacedCiliumNetworkPolicyso Hubble sees the same intent. Controller stays free of the Cilium Go SDK (raw REST).kaironctl network status MACHINEprints emoji lines for dataplane, attach, and CNP sync — reference · CLI
Snapshots
- CSI VolumeSnapshot —
MachineSnapshotorchestrates a standard snapshot object;MachineSnapshotRestorerestores one into a newPersistentVolumeClaimvia the standard CSIdataSourceflow (deliberately doesn't also create a Machine — see the guide for why) — guide - Guest quiesce for snapshots (
spec.guestAgent.enabled) — a realguest-fsfreeze/-thawaroundMachineSnapshot's VolumeSnapshot creation for an application-consistent snapshot, coordinated betweenkairon-controllerandkairon-nodesince only the node has a network path to the Machine's FluxVM instance — guide MachineSnapshotSchedule(first cut) — periodically creates aMachineSnapshotfor every Machine matching a label selector on a plainspec.intervalSeconds, Kairon's own backup-automation primitive — deliberately not real cron syntax.spec.keepLastoptionally retains only the N most recent ready-to-use snapshots per Machine, pruning older ones the schedule itself created (never a manually-created one).status.nextRunTimeprojects when a schedule's next round is expected (kaironctl get snapshotschedules' NEXTRUN column,kubectl get machinesnapshotschedules' own NextRun printer column, and the dashboard's Next run column all show it) — a simple as-of-last-fire projection, not a live countdown;kaironctl/the dashboard both check the livespec.suspendflag before trusting it, so a schedule suspended after its last run never shows a stale already-passed timestamp. The dashboard's own Snapshot schedules page lists every schedule and can suspend/resume one with a click — the one write action exposed there so far;spec.selector/.intervalSeconds/.keepLaststill needkaironctl/kubectlto edit.kaironctl describe snapshotschedule NAMEpreviews exactly which Machines its selector currently matches and whether the next reconcile tick would actually fire a new round of snapshots for them, evaluated with the sameSpec.Duecheck the controller itself uses — a real "what would this do right now" question worth answering before loosening/tightening a selector or interval, without waiting for the next tick to find out.spec.startingDeadlineSeconds(opt-in, unset by default) is Kairon's analog of KubernetesCronJob's field of the same name: a run discovered more than that many seconds late — e.g. afterkairon-controllerwas down, or this CRD's status was reset by a reinstall — is skipped rather than fired as a stale catch-up, though the schedule's own clock still advances so it doesn't re-detect the same missed window forever;status.lastRunErrorrecords the skip, surfaced the same way any other run error already is (kaironctl get/describe, the dashboard's Last error column).kaironctl trigger snapshotschedule NAMErequests an immediate, out-of-band run — "back up these Machines right now" without waiting forintervalSecondsto elapse, or temporarily touchingsuspend/intervalSecondsto force one — implemented as a durable request annotation (kairon.zyvor.dev/trigger-now) the controller notices and acts on its very next reconcile tick, bypassing bothspec.suspendandspec.startingDeadlineSeconds(a paused schedule can still be asked for one snapshot without permanently unpausing it, and "right now" is never "too late");status.lastHandledTriggerTimerecords which request was already handled so the same one is never re-fired — guide
Operate it
kaironctl— Cobra CLI with colored tables and emoji install progress;install/upgrade/uninstalluse the embedded Helm chart (no separatehelmbinary).kubectl kaironis the same tree (Krew manifest indeploy/krew/).network statuscovers Cilium attach/CNP sync. Full reference: CLIkairon-ui— an optional web dashboard, real per-operator login with in-place password change/reset (not "regenerate a hash and redeploy"), optional OIDC/SSO, aNeedsRecoveryrecovery workflow, an optional VNC console, and more than one replica once you need it — see The dashboardkairon-controllerHA —controller.replicaCountcan go above 1: acoordination.k8s.io/v1Lease (internal/leaderelection, on by default) coordinates replicas so only the elected leader reconciles, while the admission webhook and health/metrics endpoints keep serving from every replica regardless — see guide- Prometheus metrics + alert rules, a per-node/cluster migration concurrency quota,
status.dataPlaneEncryptedvisibility into whether a live transfer is actually encrypted - Opt-in OTel reconcile spans (
otel.enabled/KAIRON_OTEL_ENDPOINT) — tiny OTLP/HTTP JSON exporter on controller and node item loops, not the full OpenTelemetry SDK — guide PodDisruptionBudgeton by default, opt-inNetworkPolicy, digest-pinned + Trivy-scanned + cosign-signed + SBOM'd container images, Helm + raw manifests + CI
Live memory transfer needs a node-local migration adapter deployed and configured on each node — without one, a live request blocks before the source is touched. Cold migration needs nothing extra and works today.
Snapshots & devices (examples)
spec:
volumes:
- {name: data, claimName: database-data}
kaironctl snapshot database --name before-upgrade --class csi-snapclass
spec:
deviceClaims:
- name: gpu-claim # administrator allowlist on the node -- empty allowlist denies all
More worked examples: examples/.