Compatibility matrix
Short, dated evidence for Kairon’s differentiating features on real hardware. Historical narrative belongs in runbooks; this file is the pass/fail surface operators and releases should cite.
Do not claim production live migration in README until the live row below is green from an automated nightly run.
Lab profile (Zyvor)
| Component | Expectation |
|---|---|
| Nodes | ≥3 KVM hosts with FluxVM |
| Storage | Shared filesystem or Ceph RBD at identical paths |
| Network | L2/L3 reachability for guest + migration data plane |
| CNI | Cilium (when network-policy / identity tests run) |
| Runner | Self-hosted GitHub Actions runner with lab access |
Automation entry points:
- Runbook:
docs/runbook-multi-host-migration-test.md - Matrix driver:
scripts/hardware-migration-matrix.sh - Workflow:
.github/workflows/hardware-migration.yml(workflow_dispatch+ nightly) - Failure drills the matrix runs (each also works standalone):
scripts/recovery-drill.sh(firewalls the destination's control port mid-migration, then runskaironctl recover),scripts/lab-inject-source-failure.sh(stopskairon-nodeon the source during transfer), andscripts/lab-inject-controller-failover.sh(restarts the controller mid-migration)
Lab enablement (current blocker for the multi-host matrix): the
matrix job runs only on runs-on: [self-hosted, kairon-lab]. Bring that
runner online and set repository secrets/vars KAIRON_HW_LAB=1,
KAIRON_KUBE_URL, KAIRON_KUBE_TOKEN, KAIRON_KUBE_CA (PEM content;
the workflow writes it to a file), KAIRON_HW_MACHINE,
KAIRON_HW_TARGET_NODE, plus optional KAIRON_HW_EBPF_MACHINE (live-eBPF
case), KAIRON_HW_IMAGE (real create/stop/start/delete smoke),
KAIRON_HW_SOURCE_SSH / KAIRON_HW_TARGET_SSH (failure drills; the runner
needs passwordless SSH and sudo for iptables / systemctl) and
KAIRON_HW_CONTROLLER_SSH (systemd controller instead of Pods). Until then,
push-triggered runs succeed on the validate job only and leave every
migration row below as not run.
Single-host deploy smoke (manual)
Bare-metal systemd path (scripts/deploy-remote.sh), not a substitute for
the multi-host matrix above.
| Case | Last result | Date | Notes |
|---|---|---|---|
Deploy node+controller+ui (v0.6.0-24-g0020758) | pass | 2026-09-22 | Host 80.79.5.173 / nldw4-4-16-36 (k3s). Health: node :8081, controller :8083 (host :8080 occupied by krytond), ui :8082 |
| CRD server-side apply | pass | 2026-09-22 | deploy/crd.yaml |
| Machine create / Running | pass | 2026-09-22 | v07-smoke then deleted |
| Machine stop → start | pass | 2026-09-22 | deploy-smoke |
| UI overview + machines API | pass | 2026-09-22 | http://80.79.5.173:8082 |
Network effective (console relay) | pass | 2026-09-22 | Needs matching KAIRON_NODE_CONSOLE_TOKEN in both env files; redeploy leaves existing env untouched |
| Network flows/stats/drops | n/a | 2026-09-22 | deploy-smoke uses network.mode: user — FluxVM eBPF flow/drop/stats need tap+dataplaneMode: ebpf |
| VM edge, Kairon-scheduled Machine | pass | 2026-10-04 | Host 175.110.122.71 (k3s, FluxVM cilium mode, schema 12). Machine edge-demo (tap+netns, dataplaneMode: ebpf, antiSpoof, learnIP, qos) Running; FluxVM loaded flags, egress 2000 pps, 3 allow-list names and a 100 Mbit tbf from Kairon's edge post; status.network.edge projected |
| VM edge enforcement (FluxVM VMs on the same host) | pass | 2026-10-04 | DNS and SNI allow/deny, spoof_ip and spoof_mac (bridged tap), egress pps rate_limit, ingress police drops, learned IP via ARP, live conntrack export (14 entries), and edge + pending conntrack reapplied after FluxVM restart and VM stop/start |
| VM edge packet capture via kairon-ui | pass | 2026-10-04 | kaironctl network capture edge-demo --filter icmp -o icmp.pcap through kairon-ui, the node relay (KAIRON_NODE_CONSOLE_ADDR=:8092) and FluxVM: pcap with ICMP echo request and reply; captures lists done with packet counts; unknown token 404 |
| VM edge metrics | pass | 2026-10-04 | kairon-node /metrics exports kairon_net_conntrack_restored_total and kairon_net_migration_blackhole_ms; kairon_net_drops_total appears on the first drop (unit-tested) |
| Policy-only edge, netns MAC default | unit tests | 2026-10-04 | Not yet run against a live Machine |
Migration matrix
| Case | Last result | Date | Notes |
|---|---|---|---|
| Create / restart / delete smoke (≤100 VMs scaled down in CI) | not run | — | Multi-host matrix; single-host create/stop/start smoke is green above |
| Cold migration | not run | — | |
| Live pre-copy under load | not run | — | |
| Guest agent exec after migration | not run | — | echo through kairon-ui → kairon-node → qemu-guest-agent in the migrated Machine; needs spec.guestAgent.enabled and KAIRON_UI_URL/KAIRON_UI_TOKEN |
| Live + eBPF dataplane | not run | — | Machine with spec.network.dataplaneMode=ebpf; set KAIRON_HW_EBPF_MACHINE. Needs two hosts; also exercises the VM-edge conntrack export/restore |
| Source failure during transfer | not run | — | |
Ambiguous commit → NeedsRecovery | not run | — | |
| Controller failover mid-migration | not run | — |
hardware-migration-matrix.sh --write-compat rewrites these rows from a
run (the workflow uploads the result as the compatibility-md artifact);
without the flag it prints markdown rows for copy/paste. Every migration
case checks that the Machine ends up Running and Ready on the expected node,
and the script exits non-zero if any case fails. For the eBPF live case, the lab
Machine must set dataplaneMode: ebpf and dataplaneRequired: true so
migration network quiesce/export/restore exercises FluxVM's TC/eBPF path
(see network-fabric.md).
Component versions (fill per run)
| Piece | Version |
|---|---|
| Kairon | |
| Kubernetes | |
| FluxVM | |
| Cilium | |
| CSI / storage |