Status
Release status and production gaps formerly maintained in the root README.
v0.6.0 is tagged and open source (production foundations: status skip-patch, node-scoped watches, capacity scheduling, production Helm profile, release artifacts, expanded CI). Prior v0.5.0 absorbed, absorbing everything the sections above describe: admission webhooks, dashboard password management, console TLS automation, Helm/CI hardening, MachineSet/instance types/Windows/NUMA parity, pause/resume/halt, VM-state snapshot/restore, a full second wave of FluxVM route wrapping (both guest-exec channels, guest file access, sandboxes/templates/warm pools/image catalog, runtime/network diagnostics), the MachineSnapshotSchedule CRD, kaironctl top, and a hardening pass aimed at untrusted multi-tenant traffic (opt-in namespace-scoped kairon-ui authorization, opt-in network default-deny) -- see RELEASE_NOTES.md for the full per-release changelog. Cold relocation, snapshots, DRA bridging, the secure live control plane, and a real FluxVM migration adapter are all real and tested — real two-host live migration has not yet been exercised against real hardware in this repository's own CI (see Operability).
Toward v0.7 (on main, not tagged yet)
Already on main (see ROADMAP.md): eBPF operator surface (kaironctl network flows|drop-reasons|stats|effective, production dataplaneMode: ebpf), dashboard Network panel, multi-volume Machines (volumes[1+] → virtiofs), lease-aware NodeUnreachable (AgentLivenessStale), opt-in CSI node CHAP (node.csi.chap.enabled), opt-in OTel reconcile spans (otel.enabled). The root README is the product landing for this track.
Still required before cutting v0.7.0:
- Green Zyvor lab matrix — cold + live + live-eBPF (and the failure/NeedsRecovery/failover cases) recorded in
COMPATIBILITY.md. Workflow:.github/workflows/hardware-migration.yml. Blockers today: self-hostedkairon-labrunner must be online, and repo secrets/vars (KAIRON_HW_LAB,KAIRON_KUBE_*, optionalKAIRON_HW_EBPF_MACHINE) must be set. Single-host systemd deploy smoke on80.79.5.173is already green (create/stop/start, UI,network-effective) — see that file's "Single-host deploy smoke" table; it does not unlock the multi-host live claim. - Cut v0.7.0 — bump
VERSION/ ChartappVersion,RELEASE_NOTES.md, README maturity / comparison Status once live rows are green.
Out of band (not a code gate): CII Best Practices badge still needs human OAuth (Scorecard / issue tracker).
Production gaps
Closed in the v0.6 foundations track (agent status path): kairon-node no longer patches Machine status every reconcile tick. Conditions use shared helpers that preserve LastTransitionTime unless Status/Reason/Message actually change; observedGeneration is set on meaningful writes; volatile ResourceUsage alone does not force an etcd write (live samples go to kairon_machine_resource_usage Prometheus gauges). kaironctl top / UI node-usage therefore reflect the last meaningful status write.
Closed in the v0.6 foundations track (node-scoped lists): each kairon-node lists Machines with kairon.zyvor.dev/assigned-node=<thisNode> and migrations with kairon.zyvor.dev/migration-source-node=<thisNode>, driven by watches with a periodic safety resync (default 5m), instead of listing every Machine/migration in the cluster every few seconds. The controller stamps and repairs those labels whenever spec.nodeName / status.sourceNode change.
Closed in the v0.6 foundations track (capacity-aware scheduling): the scheduler hard-filters on node status.allocatable CPU/memory (and hugepages when requested), scores by remaining capacity percentage instead of equal VM counts, and the controller reserves capacity within a reconcile tick so simultaneous placements cannot overcommit one node. Unschedulable outcomes set a Scheduled=False Condition with a structured reason.
Closed in the v0.6 foundations track (packaging): Helm image tags default to .Chart.AppVersion (not latest); charts/kairon/values-production.yaml enables webhook, namespace isolation, network default-deny, and migration dataplane TLS. Release workflow publishes CLI binaries, checksums, OCI Helm chart, and a GitHub Release. Hardware migration evidence is tracked in docs/COMPATIBILITY.md and driven by .github/workflows/hardware-migration.yml.
Still genuinely open, and why:
- Cilium cluster-network attach is opt-in and not a Pod CNI.
network.ciliumAttach/network.ciliumPolicySyncdefault to off so bare-metal clusters without Cilium are unchanged. ExternalWorkload identity and IPv4 stay empty until Cilium's own agent registers the workload — Kairon projects status, it does not program BPF. Multus NAD is still not a primary attach path. There is no embedded Hubble UI;kaironctl network status --flowsis a pass-through to the existing uiapi whenKAIRON_UI_URLis set. Seedocs/network-fabric.md. - Fencing detection is dual-signal when opted in, still not exhaustive. With
node.livenessLease.enabled,kairon-controllerfeeds each node's owncoordination.k8s.io/v1Lease intoNodeUnreachabledetection (Ready + stale agent lease →AgentLivenessStale) in addition to Node-Ready, andkaironctl fence --liveness-lease-namespace ...still cross-checks before clearing runtime state. Lease lookup errors fail open (do not inventAgentLivenessStale). Migration preflight can only catch a confirmed storage/network mismatch when both nodes are labeled withkairon.zyvor.dev/storage-domain/network-domain— it can't prove compatibility when the labels are unset. Real multi-host fencing/preflight behavior hasn't been exercised against real hardware in this repo's own CI. Seedocs/guides/machine-fencing.md. topologySpreadConstraints.maxSkewhard enforcement is opt-in per constraint (whenUnsatisfiable: DoNotSchedule) — omitted orScheduleAnyway(the default) stays scoring-only, as before. ADoNotScheduledomain is only counted among currently-eligible candidate nodes, not every node cluster-wide, and — like affinity/anti-affinity — is evaluated against the reconcile-time Machine list, not a live watch, so a same-tick scheduling race can transiently see stale domain counts (resolved on the next tick). Seedocs/guides/machine-placement.md.- DRA topology-awareness is a best-effort scoring hint, not an allocation decision —
kairon-controllerhas no role in DRA device allocation itself; seedocs/guides/machine-placement.md. - Confidential-compute enforcement (SEV-SNP/TDX), large-scale hardware qualification — hardware-dependent, not exercisable in CI.
- PVC-backed boot disks are a first cut: one boot volume per Machine,
Filesystem-modePersistentVolumes only.hostPath/localsources resolve directly; a network-block volume now works too, through Kairon's own first-cut CSI driver (csiNode.enabled, iSCSI only) — dynamic provisioning, CHAP (controller path always; node-side consumption via opt-innode.csi.chap.enabledallowlisted to the release namespace), volume expansion, and volume snapshots are all supported (csiController.enabled/.snapshotter.enabled). No other backends (Ceph/EBS/etc.) yet — seedocs/guides/machine-storage-csi.md. A PV naming any other CSI driver is still refused unless listed innode.thirdPartyCSIDrivers. MachineSnapshotRestorerestores into a new PVC only, deliberately not also a Machine (see its guide for why), and needs a real CSI snapshotter behind your StorageClass — Rancher'slocal-path-provisioner, a common default, doesn't have one. Creating the restore before itsMachineSnapshotfinishes is not an error:status.phaseparks atPendingand retries automatically once the snapshot (and its underlyingVolumeSnapshot) becomes ready — this used to instead land in a permanent, non-retryingFailedrequiring a delete-and-recreate, an inconsistency with how every other "waiting on an external condition" status in this same feature (e.g. a PVC stillPendingunderWaitForFirstConsumer) was already handled correctly.- CPU/memory hotplug is grow-only (FluxVM has no CPU/DIMM unplug), bounded by headroom reserved at creation (
spec.resources.maxCpu/.maxMemory, itself fixed once set — not something a running Machine can grow), and lost across any stop/start — expected QEMU behavior, not a bug. kairon-uimulti-replica state propagation is eventually-consistent, not instant:ui.replicaCount > 1is now supported (session revocation, login lockout, console tickets, and password changes propagate via a sharedConfigMap, deliberately not Redis), but cross-replica visibility lands within ~15s, login-lockout's failure count is per-replica not cluster-wide-atomic, and concurrent password changes to two different accounts on two different replicas can still race. Seedocs/guides/kairon-ui-ha.md.kairon-controllerleader-election takeover is coarse, not instant:controller.replicaCount > 1is now safe (acoordination.k8s.io/v1Lease, on by default, ensures only one replica ever reconciles), but a crashed or partitioned leader costs up to ~15s before another replica takes over and reconciliation resumes. Seedocs/guides/kairon-controller-ha.md.- Distributed tracing is opt-in and narrow. Every component exposes reconcile-loop and apiserver-call health metrics.
kairon_reconcile_item_errors_total{kind}counts per-item failures. Opt-inotel.enabled(Helm) /KAIRON_OTEL_ENDPOINTposts OTLP/HTTP JSON reconcile spans fromkairon-controllerandkairon-nodevia a tiny exporter (internal/oteltrace), not the full OpenTelemetry Go SDK — enough to see which Machine/migration timed out without a third large stdlib exception. Seedocs/guides/observability.md. - Rate limiting covers
kairon-uionly.kairon-uinow rate-limits every route per client address (-rate-limit-rps/-rate-limit-burst, on by default), on top of the existing per-username login lockout — keyed by the direct TCP peer by default, butui.rateLimit.trustedProxyHeader/.trustedProxyCidrs(opt-in, both empty by default) let it key by a forwarded-for header instead, once trusted, so aService/Ingress/load balancer no longer collapses every real client into one bucket. The header is only ever trusted from a direct peer matchingtrustedProxyCidrs— set to the load balancer's own address range — otherwise any client could spoof it to evade throttling or collide with someone else's bucket. The admission webhook and inter-component RPCs (migration control-plane mTLS, the console relay) deliberately have no rate limiting added here: the webhook backstops real Kubernetes writes underfailurePolicy: Fail, so throttling it risks rejecting a legitimate bulkkubectl apply; the RPCs are already restricted to mTLS/bearer-token-authenticated peers, not open to arbitrary traffic. See SECURITY.md. - Every
kairon-uiJSON request body is now size-capped. Nothing inkairon-ui's HTTP stack ever bounded request body size before this — a caller, even a legitimately authenticated one, could send an arbitrarily large body and havejson.Decoderbuffer all of it into memory.decodeJSON(internal/uiapi/server.go) now wraps every request body inhttp.MaxBytesReaderat a 1MiB default, comfortably above every ordinary request this API accepts; guest file writes (POST .../agent-file/put) get an explicit 6MiB override, since a real file's content travels base64-encoded (~4/3 expansion) and the read side of the same feature was already capped around 3MB of real content. Rejected with a plain 400, not a broken connection — seedocs/guides/machine-guest-agent-files.md. - Live-migration data-plane identity is opt-in, and still two places to configure at once.
migration.dataplaneTlsSecretName(opt-in) gives every node a real, distinguishable cert for the QEMU RAM/state stream instead of the shared control-plane cert every node could otherwise present there —scripts/gen-migration-mtls-certs.shgenerates both the certs and a ready-to-apply Secret. Left unset (the default), the data plane still falls back to the shared cert, same as before this existed. Either way, actually turning data-plane encryption on also requires each node'skairon-migration-adapter-fluxvmsystemd unit to pass-migration-data-tls=trueindependently — outside Helm's control, easy to half-upgrade a fleet. The control-plane cert (migration.tlsSecretName) stays deliberately shared regardless — it proves cluster membership, not host identity, by design. Seedocs/guides/machine-migration-tls.md. - Backup/restore covers Kairon's Kubernetes-level state only (
scripts/backup-crds.sh/restore-crds.sh, seedocs/runbook-backup-restore.md) — the eightkairon.zyvor.devCRDs and, opt-in, the chart-managed Secrets. It does not back up VM disk content (your CSI driver's/storage backend's own job) or FluxVM's own per-host runtime state; whether a restore gets you working VMs back, not just Kubernetes objects, depends on whether those survived independently. Not yet drilled against a real full cluster-loss scenario in this repo's own CI. - OIDC/SSO group-to-admin mapping is opt-in and login-time-fresh, not IdP-live:
ui.oidc.adminGroups(unset by default, matching prior behavior exactly) grants admin capability to an OIDC session whose ID token carries one of the configured groups, re-checked against current Kairon config on every request — but the group membership itself only reflects the IdP's state as of the caller's last login, not a live check, so removing someone from an IdP group doesn't revoke their admin session early; it takes effect at their next sign-in. Seedocs/guides/kairon-ui-oidc.md. - No CRD actually runs more than
v1alpha1yet, but the conversion webhook scaffold now exists:kairon-controller's webhook server exposes a testedPOST /convert/machinequotas(internal/conversion), proven end to end against a real workedMachineQuotafield-rename example — on the same TLS listener/certificate/Serviceas the existing admission webhook, no new trust boundary. Cutting a realv1beta1still needs a CRD manifest change and a real converter for whatever's actually changing, plus solving one still-open gap:charts/kairon/crds/*.yamllives in Helm's special, never-templatedcrds/directory, so wiring a livecaBundleinto a CRD'sspec.conversionneeds a deliberate choice (move that CRD intotemplates/, or a separate patch step) not yet made. Seedocs/guides/crd-versioning.md. - The admission webhook's
MachineDisruptionBudgetand Machine-CREATE-MachineQuotachecks only ever evaluateCREATE, matching the reconcile-loop/kaironctlchecks they backstop exactly.MachineQuotaalso evaluatesUPDATE(a hotplug resize past quota) at the webhook, but only whenwebhook.enabledis true. Independent of that:kairon-nodenow also self-checks a hotplug resize againstMachineQuotaitself before actually growing a running Machine's live vCPUs/memory — reusing the exact same admission logic as the webhook (internal/controller.QuotaTrackersForNamespace/AdmitQuotaResize) — so a resize past quota is caught even on a deployment that has never turned the webhook on. This self-check is effectively on by default (no new Helm flag) but fails open: any lookup error, most commonly akairon-nodeServiceAccount that hasn't yet been granted the new read-onlymachinequotasRBAC (a binary upgraded ahead of its chart), causes it to log a warning and proceed with the resize exactly as before this existed, rather than newly blocking on account of infrastructure the check itself depends on. - The VNC console inherits FluxVM's own unauthenticated VNC socket as-is — Kairon can't add auth/encryption FluxVM itself doesn't have. Its real security rests on kairon-ui's operator auth, a single-use per-session ticket now bound to the one Machine it was issued for, an opt-in per-Machine allowlist (
kairon.zyvor.dev/console-allowed-users, unset means unchanged all-operator access), and the shared cluster-wide token to kairon-node (optionally now over TLS). See SECURITY.md. - Every feature added since VM-state snapshot/restore (Halted, both guest-exec channels, sandboxes/templates/HTTP-proxy/warm pools, the image catalog, egress check, runtime and network diagnostics) is API-only — no dashboard button yet, matching how earlier features like
qga/fsfreeze-status/firewall shipped kubectl/API-only first before getting a UI.kaironctlalso has no dedicated verbs for most of these. kairon-uinow has a dashboard page forMachineQuota/MachineDisruptionBudget/MachineSet/MachineInstanceType/MigrationPolicy.GET /api/v1/quotas/disruption-budgets/machinesets/instancetypes/migration-policieswrap the same read-onlyList*Namespacecallskaironctl getalready uses — any-authenticated-operator, matching every other read-only cross-checkable resource — and each now has a matching list page (web/src/pages/{Quotas,DisruptionBudgets,MachineSets,InstanceTypes,MigrationPolicies}.tsx), closing what was previously a real visibility gap: an operator using only the dashboard had no way to see quota usage or disruption-budget state at all, both first-class admission-time concerns (see Guarding the fleet). Four of the five stay list-only — no create/update/delete route or dashboard form;kaironctl/kubectlremain how they get mutated.MachineSetis the one exception:DELETE /api/v1/machinesets/{namespace}/{name},PATCH /api/v1/machinesets/{namespace}/{name}/scale, and matching "Delete"/"Scale" row actions now let an operator remove or resize one straight from the dashboard, the same any-authenticated-operator gatehandleDeleteMachinealready has (no separate admin check). Picked over the other four because it's the one an operator manages as a day-to-day fleet-sizing operation (kaironctl create/scale/edit machinesetall shipped this same session) rather than a GitOps-managed policy object or a mostly-static reference value. Deleting a MachineSet from the dashboard also cascades to delete every Machine it created — see Deleting a MachineSet deletes its replicas too — so the dashboard's confirm prompt says so explicitly.kairon-uinow has a dashboard page for the real KubernetesNode, not just a kairon.zyvor.dev CRD. Node taints started actually gating scheduling, thenkaironctl get/describe nodegave an operator a way to see them from the CLI (see Guarding the fleet) -- butGET /api/v1/nodes(added earlier for the Overview tile's plain node count) had no dashboard page rendering the list itself until now. The new Nodes page shows the same columnskaironctl get nodesprints -- name,Readycondition,spec.unschedulable, akey[=value]:Effecttaint summary, and addresses -- read-only likekaironctl's own support, since Node's create/delete lifecycle stayskubectl's job either way. No new REST route, no new RBAC: the route and the ClusterRole grant it needs both already existed for the Overview tile. Seedocs/guides/machine-placement.md's "Taints and tolerations" section.- Fixed a real gap:
kairon-ui's ownClusterRolenever actually grantedget/list/watchonMachineQuota/MachineDisruptionBudget/MachineInstanceType/MigrationPolicy, or the extra verbsMachineSet's delete andMachineSnapshotSchedule's suspend/resume route need (delete,patch) — every one of those dashboard/REST routes above has 403'd against a real Kubernetes apiserver since the day it shipped. Nothing caught it because everyinternal/uiapitest runs its handlers against a fakehttptest.Serverdouble, which never enforces RBAC at all — a real cluster was the only way this would surface.charts/kairon/templates/all.yaml'skairon-uiClusterRolenow grants exactly the verbs each route actually needs, no more. kairon-uiandkaironctlnow also coverMachineNetworkPolicy/NetworkSecurityGroup— the two CRDs that drive FluxVM's real eBPF/TC network enforcement (see the network-policy guide) had zero human-facing tooling beyond rawkubectl, unlike every other kind in this project: nokaironctl get/describe/delete, and no dashboard/REST route at all (the existing network-observability endpoints only ever show a Machine's effective, already-merged policy, not the policy objects that produced it).kaironctl get networkpolicies/securitygroups,describe, anddeletenow work like every other kind;GET /api/v1/network-policies/security-groupsback two new list-only dashboard pages (Network policies, Security groups). The dashboard stays list-only;kaironctl create networkpolicy/create securitygroupandkaironctl edit networkpolicy/edit securitygroup(see CLI) have since closed the other half of this gap —kaironctl/kubectl(not the dashboard) are how these get created or edited.- A warm-pool claim only becomes a Kairon-managed
Machinewhen the caller opts in. By default, claiming hands back FluxVM's own VM record directly with no automaticMachineobject — a claimed VM sits outside admission-time guards (MachineQuota, the webhook) until one exists.POST .../claim'screateMachine: true(plus a requiredmachineName) closes that gap for a given claim: it forces the FluxVM-side VM name to matchMachine.RuntimeName()and creates the Machine object from the pool's own template, so kairon-node's existing adoption path picks up the exact runtime on its next reconcile tick. Covers image/CPU/memory/backend only — VFIO/NUMA/cloud-init and other template fields aren't carried onto the new Machine, a first cut. Seedocs/guides/machine-sandboxes.md's "Warm pools" section. - Real CPU pinning (
spec.resources.cpuPinning) depends entirely on an operator keeping thekairon.zyvor.dev/pinnable-cpusnode label accurate. Kairon deliberately doesn't read kubelet's own internal CPU Manager state file to auto-discover which cores are safe (a known, fragile, version-dependent, unsupported community pattern) — if the label includes a core kubelet's static CPU Manager later exclusively grants to a real Guaranteed-QoS Pod, nothing here detects or prevents that collision. Seedocs/guides/machine-cpu-pinning.md. - The third-party CSI client (
node.thirdPartyCSIDrivers) is a first cut, not a general "any CSI driver" integration.attachRequireddrivers work through aVolumeAttachment(the driver's external-attacher must be running) and node secrets resolve from one allowlisted namespace, but no live third-party driver has been run end to end yet; behavior is unit-tested against fakes. Seedocs/guides/machine-storage-thirdparty-csi.md. kairon-uinow has opt-in per-namespace authorization (ui.auth.namespaceScoping.enabled, defaultfalse) — closes a real gap:kairon-ui's ownClusterRoleis cluster-scoped and its HTTP layer previously had no namespace authorization axis at all, onlyisAdminIdentity's admin-vs-not split, so every authenticated operator (password, legacy shared token, or OIDC) could see and mutate every namespace's Machines/quotas/policies regardless of intent. Once enabled, a non-admin session-token operator is restricted to the namespaces in their ownui.auth.users[].namespacesor reachable via a matchingui.oidc.namespaceGroupsentry — everything else 403s. Off by default: flipping it on for an existing deployment with nonamespaces/namespaceGroupsconfigured moves every non-admin operator from "sees every namespace" to "sees none," a deliberate fail-closed default, not a silent no-op. Honest first-cut limits: the dashboard's namespace inputs stay free-text (no namespace-picker UI), a scoped-out operator gets the existing generic 403 rather than a constrained dropdown;GET /api/v1/overviewstill aggregates cluster-wide regardless;POST /api/v1/migrations/evacuate(a node-shaped bulk action spanning whatever namespaces that node's Machines happen to live in) isn't scoped by it either; and Node-scoped routes aren't namespace-scoped at all, since Node isn't a namespaced kairon object. The legacy shared token has no per-caller identity to scope by, so it stays fully unrestricted regardless of this setting, by design.Machine.spec.tenantis not read or enforced anywhere. It exists in the CRD as free-form metadata only — no admission check, no controller logic, no kaironctl flag, no dashboard display reads it. Kubernetes Namespace is kairon's one real multi-tenancy boundary (MachineQuota,MachineDisruptionBudget, and the namespace-scoped authorization above are all keyed on Namespace, never on this field).kubectl explain machine.spec.tenantandkaironctl describe machinenow both say so explicitly rather than leaving an operator to assume the field does something. Seedocs/guides/machine-quotas.md's "Real limits today" section.
Report vulnerabilities privately to security@zyvor.dev — see SECURITY.md.