Skip to main content

Status

Release status and production gaps formerly maintained in the root README.

v0.6.0 is tagged and open source (production foundations: status skip-patch, node-scoped watches, capacity scheduling, production Helm profile, release artifacts, expanded CI). Prior v0.5.0 absorbed, absorbing everything the sections above describe: admission webhooks, dashboard password management, console TLS automation, Helm/CI hardening, MachineSet/instance types/Windows/NUMA parity, pause/resume/halt, VM-state snapshot/restore, a full second wave of FluxVM route wrapping (both guest-exec channels, guest file access, sandboxes/templates/warm pools/image catalog, runtime/network diagnostics), the MachineSnapshotSchedule CRD, kaironctl top, and a hardening pass aimed at untrusted multi-tenant traffic (opt-in namespace-scoped kairon-ui authorization, opt-in network default-deny) -- see RELEASE_NOTES.md for the full per-release changelog. Cold relocation, snapshots, DRA bridging, the secure live control plane, and a real FluxVM migration adapter are all real and tested — real two-host live migration has not yet been exercised against real hardware in this repository's own CI (see Operability).

Toward v0.7 (on main, not tagged yet)​

Already on main (see ROADMAP.md): eBPF operator surface (kaironctl network flows|drop-reasons|stats|effective, production dataplaneMode: ebpf), dashboard Network panel, multi-volume Machines (volumes[1+] → virtiofs), lease-aware NodeUnreachable (AgentLivenessStale), opt-in CSI node CHAP (node.csi.chap.enabled), opt-in OTel reconcile spans (otel.enabled). The root README is the product landing for this track.

Still required before cutting v0.7.0:

  1. Green Zyvor lab matrix — cold + live + live-eBPF (and the failure/NeedsRecovery/failover cases) recorded in COMPATIBILITY.md. Workflow: .github/workflows/hardware-migration.yml. Blockers today: self-hosted kairon-lab runner must be online, and repo secrets/vars (KAIRON_HW_LAB, KAIRON_KUBE_*, optional KAIRON_HW_EBPF_MACHINE) must be set. Single-host systemd deploy smoke on 80.79.5.173 is already green (create/stop/start, UI, network-effective) — see that file's "Single-host deploy smoke" table; it does not unlock the multi-host live claim.
  2. Cut v0.7.0 — bump VERSION / Chart appVersion, RELEASE_NOTES.md, README maturity / comparison Status once live rows are green.

Out of band (not a code gate): CII Best Practices badge still needs human OAuth (Scorecard / issue tracker).

Production gaps​

Closed in the v0.6 foundations track (agent status path): kairon-node no longer patches Machine status every reconcile tick. Conditions use shared helpers that preserve LastTransitionTime unless Status/Reason/Message actually change; observedGeneration is set on meaningful writes; volatile ResourceUsage alone does not force an etcd write (live samples go to kairon_machine_resource_usage Prometheus gauges). kaironctl top / UI node-usage therefore reflect the last meaningful status write.

Closed in the v0.6 foundations track (node-scoped lists): each kairon-node lists Machines with kairon.zyvor.dev/assigned-node=<thisNode> and migrations with kairon.zyvor.dev/migration-source-node=<thisNode>, driven by watches with a periodic safety resync (default 5m), instead of listing every Machine/migration in the cluster every few seconds. The controller stamps and repairs those labels whenever spec.nodeName / status.sourceNode change.

Closed in the v0.6 foundations track (capacity-aware scheduling): the scheduler hard-filters on node status.allocatable CPU/memory (and hugepages when requested), scores by remaining capacity percentage instead of equal VM counts, and the controller reserves capacity within a reconcile tick so simultaneous placements cannot overcommit one node. Unschedulable outcomes set a Scheduled=False Condition with a structured reason.

Closed in the v0.6 foundations track (packaging): Helm image tags default to .Chart.AppVersion (not latest); charts/kairon/values-production.yaml enables webhook, namespace isolation, network default-deny, and migration dataplane TLS. Release workflow publishes CLI binaries, checksums, OCI Helm chart, and a GitHub Release. Hardware migration evidence is tracked in docs/COMPATIBILITY.md and driven by .github/workflows/hardware-migration.yml.

Still genuinely open, and why:

  • Cilium cluster-network attach is opt-in and not a Pod CNI. network.ciliumAttach / network.ciliumPolicySync default to off so bare-metal clusters without Cilium are unchanged. ExternalWorkload identity and IPv4 stay empty until Cilium's own agent registers the workload — Kairon projects status, it does not program BPF. Multus NAD is still not a primary attach path. There is no embedded Hubble UI; kaironctl network status --flows is a pass-through to the existing uiapi when KAIRON_UI_URL is set. See docs/network-fabric.md.
  • Fencing detection is dual-signal when opted in, still not exhaustive. With node.livenessLease.enabled, kairon-controller feeds each node's own coordination.k8s.io/v1 Lease into NodeUnreachable detection (Ready + stale agent lease → AgentLivenessStale) in addition to Node-Ready, and kaironctl fence --liveness-lease-namespace ... still cross-checks before clearing runtime state. Lease lookup errors fail open (do not invent AgentLivenessStale). Migration preflight can only catch a confirmed storage/network mismatch when both nodes are labeled with kairon.zyvor.dev/storage-domain/network-domain — it can't prove compatibility when the labels are unset. Real multi-host fencing/preflight behavior hasn't been exercised against real hardware in this repo's own CI. See docs/guides/machine-fencing.md.
  • topologySpreadConstraints.maxSkew hard enforcement is opt-in per constraint (whenUnsatisfiable: DoNotSchedule) — omitted or ScheduleAnyway (the default) stays scoring-only, as before. A DoNotSchedule domain is only counted among currently-eligible candidate nodes, not every node cluster-wide, and — like affinity/anti-affinity — is evaluated against the reconcile-time Machine list, not a live watch, so a same-tick scheduling race can transiently see stale domain counts (resolved on the next tick). See docs/guides/machine-placement.md.
  • DRA topology-awareness is a best-effort scoring hint, not an allocation decision — kairon-controller has no role in DRA device allocation itself; see docs/guides/machine-placement.md.
  • Confidential-compute enforcement (SEV-SNP/TDX), large-scale hardware qualification — hardware-dependent, not exercisable in CI.
  • PVC-backed boot disks are a first cut: one boot volume per Machine, Filesystem-mode PersistentVolumes only. hostPath/local sources resolve directly; a network-block volume now works too, through Kairon's own first-cut CSI driver (csiNode.enabled, iSCSI only) — dynamic provisioning, CHAP (controller path always; node-side consumption via opt-in node.csi.chap.enabled allowlisted to the release namespace), volume expansion, and volume snapshots are all supported (csiController.enabled/.snapshotter.enabled). No other backends (Ceph/EBS/etc.) yet — see docs/guides/machine-storage-csi.md. A PV naming any other CSI driver is still refused unless listed in node.thirdPartyCSIDrivers.
  • MachineSnapshotRestore restores into a new PVC only, deliberately not also a Machine (see its guide for why), and needs a real CSI snapshotter behind your StorageClass — Rancher's local-path-provisioner, a common default, doesn't have one. Creating the restore before its MachineSnapshot finishes is not an error: status.phase parks at Pending and retries automatically once the snapshot (and its underlying VolumeSnapshot) becomes ready — this used to instead land in a permanent, non-retrying Failed requiring a delete-and-recreate, an inconsistency with how every other "waiting on an external condition" status in this same feature (e.g. a PVC still Pending under WaitForFirstConsumer) was already handled correctly.
  • CPU/memory hotplug is grow-only (FluxVM has no CPU/DIMM unplug), bounded by headroom reserved at creation (spec.resources.maxCpu/.maxMemory, itself fixed once set — not something a running Machine can grow), and lost across any stop/start — expected QEMU behavior, not a bug.
  • kairon-ui multi-replica state propagation is eventually-consistent, not instant: ui.replicaCount > 1 is now supported (session revocation, login lockout, console tickets, and password changes propagate via a shared ConfigMap, deliberately not Redis), but cross-replica visibility lands within ~15s, login-lockout's failure count is per-replica not cluster-wide-atomic, and concurrent password changes to two different accounts on two different replicas can still race. See docs/guides/kairon-ui-ha.md.
  • kairon-controller leader-election takeover is coarse, not instant: controller.replicaCount > 1 is now safe (a coordination.k8s.io/v1 Lease, on by default, ensures only one replica ever reconciles), but a crashed or partitioned leader costs up to ~15s before another replica takes over and reconciliation resumes. See docs/guides/kairon-controller-ha.md.
  • Distributed tracing is opt-in and narrow. Every component exposes reconcile-loop and apiserver-call health metrics. kairon_reconcile_item_errors_total{kind} counts per-item failures. Opt-in otel.enabled (Helm) / KAIRON_OTEL_ENDPOINT posts OTLP/HTTP JSON reconcile spans from kairon-controller and kairon-node via a tiny exporter (internal/oteltrace), not the full OpenTelemetry Go SDK — enough to see which Machine/migration timed out without a third large stdlib exception. See docs/guides/observability.md.
  • Rate limiting covers kairon-ui only. kairon-ui now rate-limits every route per client address (-rate-limit-rps/-rate-limit-burst, on by default), on top of the existing per-username login lockout — keyed by the direct TCP peer by default, but ui.rateLimit.trustedProxyHeader/.trustedProxyCidrs (opt-in, both empty by default) let it key by a forwarded-for header instead, once trusted, so a Service/Ingress/load balancer no longer collapses every real client into one bucket. The header is only ever trusted from a direct peer matching trustedProxyCidrs — set to the load balancer's own address range — otherwise any client could spoof it to evade throttling or collide with someone else's bucket. The admission webhook and inter-component RPCs (migration control-plane mTLS, the console relay) deliberately have no rate limiting added here: the webhook backstops real Kubernetes writes under failurePolicy: Fail, so throttling it risks rejecting a legitimate bulk kubectl apply; the RPCs are already restricted to mTLS/bearer-token-authenticated peers, not open to arbitrary traffic. See SECURITY.md.
  • Every kairon-ui JSON request body is now size-capped. Nothing in kairon-ui's HTTP stack ever bounded request body size before this — a caller, even a legitimately authenticated one, could send an arbitrarily large body and have json.Decoder buffer all of it into memory. decodeJSON (internal/uiapi/server.go) now wraps every request body in http.MaxBytesReader at a 1MiB default, comfortably above every ordinary request this API accepts; guest file writes (POST .../agent-file/put) get an explicit 6MiB override, since a real file's content travels base64-encoded (~4/3 expansion) and the read side of the same feature was already capped around 3MB of real content. Rejected with a plain 400, not a broken connection — see docs/guides/machine-guest-agent-files.md.
  • Live-migration data-plane identity is opt-in, and still two places to configure at once. migration.dataplaneTlsSecretName (opt-in) gives every node a real, distinguishable cert for the QEMU RAM/state stream instead of the shared control-plane cert every node could otherwise present there — scripts/gen-migration-mtls-certs.sh generates both the certs and a ready-to-apply Secret. Left unset (the default), the data plane still falls back to the shared cert, same as before this existed. Either way, actually turning data-plane encryption on also requires each node's kairon-migration-adapter-fluxvm systemd unit to pass -migration-data-tls=true independently — outside Helm's control, easy to half-upgrade a fleet. The control-plane cert (migration.tlsSecretName) stays deliberately shared regardless — it proves cluster membership, not host identity, by design. See docs/guides/machine-migration-tls.md.
  • Backup/restore covers Kairon's Kubernetes-level state only (scripts/backup-crds.sh/restore-crds.sh, see docs/runbook-backup-restore.md) — the eight kairon.zyvor.dev CRDs and, opt-in, the chart-managed Secrets. It does not back up VM disk content (your CSI driver's/storage backend's own job) or FluxVM's own per-host runtime state; whether a restore gets you working VMs back, not just Kubernetes objects, depends on whether those survived independently. Not yet drilled against a real full cluster-loss scenario in this repo's own CI.
  • OIDC/SSO group-to-admin mapping is opt-in and login-time-fresh, not IdP-live: ui.oidc.adminGroups (unset by default, matching prior behavior exactly) grants admin capability to an OIDC session whose ID token carries one of the configured groups, re-checked against current Kairon config on every request — but the group membership itself only reflects the IdP's state as of the caller's last login, not a live check, so removing someone from an IdP group doesn't revoke their admin session early; it takes effect at their next sign-in. See docs/guides/kairon-ui-oidc.md.
  • No CRD actually runs more than v1alpha1 yet, but the conversion webhook scaffold now exists: kairon-controller's webhook server exposes a tested POST /convert/machinequotas (internal/conversion), proven end to end against a real worked MachineQuota field-rename example — on the same TLS listener/certificate/Service as the existing admission webhook, no new trust boundary. Cutting a real v1beta1 still needs a CRD manifest change and a real converter for whatever's actually changing, plus solving one still-open gap: charts/kairon/crds/*.yaml lives in Helm's special, never-templated crds/ directory, so wiring a live caBundle into a CRD's spec.conversion needs a deliberate choice (move that CRD into templates/, or a separate patch step) not yet made. See docs/guides/crd-versioning.md.
  • The admission webhook's MachineDisruptionBudget and Machine-CREATE-MachineQuota checks only ever evaluate CREATE, matching the reconcile-loop/kaironctl checks they backstop exactly. MachineQuota also evaluates UPDATE (a hotplug resize past quota) at the webhook, but only when webhook.enabled is true. Independent of that: kairon-node now also self-checks a hotplug resize against MachineQuota itself before actually growing a running Machine's live vCPUs/memory — reusing the exact same admission logic as the webhook (internal/controller.QuotaTrackersForNamespace/AdmitQuotaResize) — so a resize past quota is caught even on a deployment that has never turned the webhook on. This self-check is effectively on by default (no new Helm flag) but fails open: any lookup error, most commonly a kairon-node ServiceAccount that hasn't yet been granted the new read-only machinequotas RBAC (a binary upgraded ahead of its chart), causes it to log a warning and proceed with the resize exactly as before this existed, rather than newly blocking on account of infrastructure the check itself depends on.
  • The VNC console inherits FluxVM's own unauthenticated VNC socket as-is — Kairon can't add auth/encryption FluxVM itself doesn't have. Its real security rests on kairon-ui's operator auth, a single-use per-session ticket now bound to the one Machine it was issued for, an opt-in per-Machine allowlist (kairon.zyvor.dev/console-allowed-users, unset means unchanged all-operator access), and the shared cluster-wide token to kairon-node (optionally now over TLS). See SECURITY.md.
  • Every feature added since VM-state snapshot/restore (Halted, both guest-exec channels, sandboxes/templates/HTTP-proxy/warm pools, the image catalog, egress check, runtime and network diagnostics) is API-only — no dashboard button yet, matching how earlier features like qga/fsfreeze-status/firewall shipped kubectl/API-only first before getting a UI. kaironctl also has no dedicated verbs for most of these.
  • kairon-ui now has a dashboard page for MachineQuota/MachineDisruptionBudget/MachineSet/MachineInstanceType/MigrationPolicy. GET /api/v1/quotas/disruption-budgets/machinesets/instancetypes/migration-policies wrap the same read-only List*Namespace calls kaironctl get already uses — any-authenticated-operator, matching every other read-only cross-checkable resource — and each now has a matching list page (web/src/pages/{Quotas,DisruptionBudgets,MachineSets,InstanceTypes,MigrationPolicies}.tsx), closing what was previously a real visibility gap: an operator using only the dashboard had no way to see quota usage or disruption-budget state at all, both first-class admission-time concerns (see Guarding the fleet). Four of the five stay list-only — no create/update/delete route or dashboard form; kaironctl/kubectl remain how they get mutated. MachineSet is the one exception: DELETE /api/v1/machinesets/{namespace}/{name}, PATCH /api/v1/machinesets/{namespace}/{name}/scale, and matching "Delete"/"Scale" row actions now let an operator remove or resize one straight from the dashboard, the same any-authenticated-operator gate handleDeleteMachine already has (no separate admin check). Picked over the other four because it's the one an operator manages as a day-to-day fleet-sizing operation (kaironctl create/scale/edit machineset all shipped this same session) rather than a GitOps-managed policy object or a mostly-static reference value. Deleting a MachineSet from the dashboard also cascades to delete every Machine it created — see Deleting a MachineSet deletes its replicas too — so the dashboard's confirm prompt says so explicitly.
  • kairon-ui now has a dashboard page for the real Kubernetes Node, not just a kairon.zyvor.dev CRD. Node taints started actually gating scheduling, then kaironctl get/describe node gave an operator a way to see them from the CLI (see Guarding the fleet) -- but GET /api/v1/nodes (added earlier for the Overview tile's plain node count) had no dashboard page rendering the list itself until now. The new Nodes page shows the same columns kaironctl get nodes prints -- name, Ready condition, spec.unschedulable, a key[=value]:Effect taint summary, and addresses -- read-only like kaironctl's own support, since Node's create/delete lifecycle stays kubectl's job either way. No new REST route, no new RBAC: the route and the ClusterRole grant it needs both already existed for the Overview tile. See docs/guides/machine-placement.md's "Taints and tolerations" section.
  • Fixed a real gap: kairon-ui's own ClusterRole never actually granted get/list/watch on MachineQuota/MachineDisruptionBudget/MachineInstanceType/MigrationPolicy, or the extra verbs MachineSet's delete and MachineSnapshotSchedule's suspend/resume route need (delete, patch) — every one of those dashboard/REST routes above has 403'd against a real Kubernetes apiserver since the day it shipped. Nothing caught it because every internal/uiapi test runs its handlers against a fake httptest.Server double, which never enforces RBAC at all — a real cluster was the only way this would surface. charts/kairon/templates/all.yaml's kairon-ui ClusterRole now grants exactly the verbs each route actually needs, no more.
  • kairon-ui and kaironctl now also cover MachineNetworkPolicy/NetworkSecurityGroup — the two CRDs that drive FluxVM's real eBPF/TC network enforcement (see the network-policy guide) had zero human-facing tooling beyond raw kubectl, unlike every other kind in this project: no kaironctl get/describe/delete, and no dashboard/REST route at all (the existing network-observability endpoints only ever show a Machine's effective, already-merged policy, not the policy objects that produced it). kaironctl get networkpolicies/securitygroups, describe, and delete now work like every other kind; GET /api/v1/network-policies/security-groups back two new list-only dashboard pages (Network policies, Security groups). The dashboard stays list-only; kaironctl create networkpolicy/create securitygroup and kaironctl edit networkpolicy/edit securitygroup (see CLI) have since closed the other half of this gap — kaironctl/kubectl (not the dashboard) are how these get created or edited.
  • A warm-pool claim only becomes a Kairon-managed Machine when the caller opts in. By default, claiming hands back FluxVM's own VM record directly with no automatic Machine object — a claimed VM sits outside admission-time guards (MachineQuota, the webhook) until one exists. POST .../claim's createMachine: true (plus a required machineName) closes that gap for a given claim: it forces the FluxVM-side VM name to match Machine.RuntimeName() and creates the Machine object from the pool's own template, so kairon-node's existing adoption path picks up the exact runtime on its next reconcile tick. Covers image/CPU/memory/backend only — VFIO/NUMA/cloud-init and other template fields aren't carried onto the new Machine, a first cut. See docs/guides/machine-sandboxes.md's "Warm pools" section.
  • Real CPU pinning (spec.resources.cpuPinning) depends entirely on an operator keeping the kairon.zyvor.dev/pinnable-cpus node label accurate. Kairon deliberately doesn't read kubelet's own internal CPU Manager state file to auto-discover which cores are safe (a known, fragile, version-dependent, unsupported community pattern) — if the label includes a core kubelet's static CPU Manager later exclusively grants to a real Guaranteed-QoS Pod, nothing here detects or prevents that collision. See docs/guides/machine-cpu-pinning.md.
  • The third-party CSI client (node.thirdPartyCSIDrivers) is a first cut, not a general "any CSI driver" integration. attachRequired drivers work through a VolumeAttachment (the driver's external-attacher must be running) and node secrets resolve from one allowlisted namespace, but no live third-party driver has been run end to end yet; behavior is unit-tested against fakes. See docs/guides/machine-storage-thirdparty-csi.md.
  • kairon-ui now has opt-in per-namespace authorization (ui.auth.namespaceScoping.enabled, default false) — closes a real gap: kairon-ui's own ClusterRole is cluster-scoped and its HTTP layer previously had no namespace authorization axis at all, only isAdminIdentity's admin-vs-not split, so every authenticated operator (password, legacy shared token, or OIDC) could see and mutate every namespace's Machines/quotas/policies regardless of intent. Once enabled, a non-admin session-token operator is restricted to the namespaces in their own ui.auth.users[].namespaces or reachable via a matching ui.oidc.namespaceGroups entry — everything else 403s. Off by default: flipping it on for an existing deployment with no namespaces/namespaceGroups configured moves every non-admin operator from "sees every namespace" to "sees none," a deliberate fail-closed default, not a silent no-op. Honest first-cut limits: the dashboard's namespace inputs stay free-text (no namespace-picker UI), a scoped-out operator gets the existing generic 403 rather than a constrained dropdown; GET /api/v1/overview still aggregates cluster-wide regardless; POST /api/v1/migrations/evacuate (a node-shaped bulk action spanning whatever namespaces that node's Machines happen to live in) isn't scoped by it either; and Node-scoped routes aren't namespace-scoped at all, since Node isn't a namespaced kairon object. The legacy shared token has no per-caller identity to scope by, so it stays fully unrestricted regardless of this setting, by design.
  • Machine.spec.tenant is not read or enforced anywhere. It exists in the CRD as free-form metadata only — no admission check, no controller logic, no kaironctl flag, no dashboard display reads it. Kubernetes Namespace is kairon's one real multi-tenancy boundary (MachineQuota, MachineDisruptionBudget, and the namespace-scoped authorization above are all keyed on Namespace, never on this field). kubectl explain machine.spec.tenant and kaironctl describe machine now both say so explicitly rather than leaving an operator to assume the field does something. See docs/guides/machine-quotas.md's "Real limits today" section.

Report vulnerabilities privately to security@zyvor.dev — see SECURITY.md.