Skip to main content

Next features (ranked)

This ranked set is complete on the code + gate path as of 2026-09-26.

It covers Secure Containers Set 19 and the P0 multi-VMM GA close (Cloud Hypervisor virtiofs, Multus hybrid/direct, Calico/Flannel, hostPath broker, Firecracker ext4 block shares β€” schema-v10, fluxvm_pridx, IPv6 extension walk, guest policy mirror, Observer ops), Sets 16–18, Sentinel operational tooling (Set 12E GA certification, Set 13E–17E migrate/upgrade/fleet/drift/ admission), in-tree KVM FC/CH parity (P0–P2 / H1–H5), and the post-GA closes below. Any future ranked set should still prefer proven live gates over inventing new wire formats β€” open a new section only when product asks.

Sentinel / Secure Containers (highest leverage)​

PriorityFeatureWhySource
S1Live stateful conntrack-bypass proof under a real TCP handshakeDone (depth) β€” mid-flow kill + live SYN anti-replay under load: scripts/e2e-networkpolicy-s1-depth.sh (plus set16/set17 starters)set14, set16, PRODUCTION
S2Multi-node NetworkPolicy conformanceDone β€” REQUIRE_MULTI_NODE=1 Set 17 + scripts/evidence-networkpolicy-second-cni.sh (second kubeconfig or nftables policy-engine stand-in)set14 #4, set13
S3True per-direction counters + rule-hit identityDone (live) β€” Set 17 fluxvm_prhit + observer scrape; evidence sc-live-s3-prhit-20260926.txt (directional_counters=1, ingress+egress hooks attached) via scripts/evidence-s3-prhit-scrape.shset14 #5, set15, set17
S4EndpointSlice-aware Service VIP policyDone (live) β€” Set 18 EndpointSlice routing; evidence sc-live-s4-endpointslice-20260926.txt (Set 18 live EndpointSlice Service-VIP gate: PASS) via scripts/test-networkpolicy-endpointslice-set18.shset14 #3, set18
S5Indexed LPM/L4 (or verifier-budget raise)Implemented by Set 19 via fluxvm_pridx candidate bitmap over the existing 64 fluxvm_prules slots (fluxvm_prules remains authoritative)set14 #1, set19
S6IPv6 extension-header walkingImplemented by Set 19 (schema-v10 bounded walk in both TC directions; fragments never conntrack-learned)set14 #2, set16, set19
S7Wire Set 8S guest cgroup policy from Pod Set 6S/14Implemented by Set 19 (create-time + live CIDR/direction mirror; host TC remains protocol/port SoT)PRODUCTION, set19
S8Policy Observer prod sizing + Prometheus scrapeDone (live) β€” ServiceMonitor + sizing + scrape path; schema recognition tracks dataplane (CurrentSchema=11); live sc-live-observer-schema11-20260926.txt (schema_compatible=1) and earlier RSS policy-observer-scrape-20260918T035554Z.txtset15, set19
S9Kata P0/P1 gates (OCI fixtures, Multus, hostPath broker, NOTIF_ADDFD, TTY churn, warm-pool claim, FC block shares)Done (matrix) β€” scripts/evidence-kata-p0p1-matrix.sh + tests/oci-fixtures/; hostPath allowlist broker; SECCOMP_IOCTL_NOTIF_ADDFD; Firecracker ext4 share packing (FLUXVM_CONTAINER_BACKEND=firecracker); RuntimeClass is GA vs full Kata product claimsecure-containers.md
S10Real multi-node fleet rollout runDone (live) β€” scripts/evidence-fleet-multihost.sh canaryβ†’approveβ†’resume on β‰₯2 SSH hosts (StrictHostKeyChecking=yes); evidence sc-live-s10-fleet-20260926.txt (S10 REAL MULTI-HOST FLEET: PASS)set15e
S11Migration orchestrator against a real attached VMDone (live) — scripts/evidence-migration-attached-vm.sh quiesce→export→restore→resume on schema-11 attached dataplane; evidence sc-live-s11-migration-20260926.txt (S11 ATTACHED MIGRATION: PASS)set13e
CNICilium-like CNI for Secure ContainersDone (live under load) β€” providers + Multus hybrid/direct; evidence sc-live-cni-under-load-20260926.txt (CNI UNDER LOAD: PASS: calico/flannel churnΓ—80 + Cilium/fluxvm ReadyΓ—5 + second-cluster runcΓ—5) via scripts/evidence-cni-under-load.shsecure-containers P0
DPDirect (bridge-less) datapath: Cilium-style redirect between the Pod veth / host uplink and the guest tapDone, default auto for Secure Containers β€” shim FLUXVM_CONTAINER_CNI_DATAPATH=auto|direct|bridge (default auto), daemon loader in the Pod netns (TCX + legacy tc), l2-uplink standalone mode (shared per-uplink maps, ARP steering, MicroVM CRD networkMode: direct), warm-pool hotplug via QMP fd-passing, Linux 7.x verifier fix, hop view, capability probe, bench + evidence scripts. Live Cilium + KVM evidence passed (FLUXVM_DIRECT_LIVE=1, direct-datapath-live-20260919T202359Z.txt); real-guest run (direct-datapath-live-realguest-20260920.txt): default/auto fallback/l2-uplink/warm-pool hotplug verified, bridge-vs-direct shows no detectable guest-visible difference; Multus secondaries support hybrid/direct opt-indirect-datapath.md

In-tree hypervisor / agent sandboxes​

PriorityFeatureWhySource
H1Virtio device live-state in FLUXKVM1Done (v3) β€” pack/restore queue rings + status/features; backends still re-attach paths from boot config; Firecracker remains production snap for FC guestshypervisor README, agent-sandbox-gaps
H2Full virtio-pci BAR / Windows pathDone (tables) β€” ECAM virtio-net @ 00:01.0, BAR0, MSI-X table at BAR0+0x800, DSDT PCI0, firmware entry at 0x01000000, ACPI reset I/O 0xCF9. virtio-win boot is unproven; Cloud Hypervisor stays the production Windows VMMDESIGN / README
H3vhost-net multi-queue bind + virtio-net guest RXDone β€” mem-table + VRING NUM/ADDR/BASE/KICK/CALL + TAP backend; kernel datapath kicks when rings programmed; a failed bind or late ring-program falls back to the userspace pump with one log line instead of a half-configured device. The userspace pump now also delivers tap frames into guest RX queue 0 (bounds-checked chain walk, used-ring publish, interrupt suppression); verified live with DHCP/ping/TCP transfer (scripts/test-kvm-net.sh, 5 runs x 1/2 vCPUs) and through the egress proxy (scripts/test-egress-guest.sh stage 2, see http-acl.md)hypervisor boot smoke / DESIGN
H4P3 on demand: virtio-fs, live migration, hotplugDone β€” FLUXKVM1 migrate export/import bundles; real guest-visible CPU hotplug (#128): every vCPU up to max_cpus is pre-created parked at boot with final-topology CPUID, maxcpus=<cpus> caps auto-onlining, guest brings the rest online itself via sysfs (echo 1 > .../cpuN/online), verified live 2β†’4 vCPUs 5/5 (scripts/test-kvm-hotplug.sh); disk hotplug validation; virtio-fs attach (virtiofsd spawn + MMIO device/tag). Control API: MigrateExport/Import, HotplugCpu/Disk. Production shared-FS still prefers CH when vhost-user queue bind depth mattersDESIGN
H5Published kvm/FC density figuresDone (archive) β€” lab run on 80.79.5.173 archived in benchmarks/evidence/density-20260918-80.79.5.173.txt; cold flux-vm create is not a Track B density claimROADMAP-DENSITY, benchmarks

Firecracker adoption (general CP β€” Track A follow-ups)​

FluxVM is a general VM control plane; FC figures are isolation layers + optional density. Track A isolation, virtio rate limiters, and FC1–FC3 are shipped:

PriorityFeatureStatus
FC1Oversubscription policy knobsDone β€” default_cpu_quota_percent / memory_max_equals_guest + oversubscription.md
FC2Static CPU templatesDone β€” cpu_template on Firecracker + FluxVm(firecracker); rejected on CH/QEMU/kvm
FC3/v1/vms/{id}/snapshot multi-backendDone β€” QEMU/CH (existing) + plain Firecracker + FluxVm control path

Custom CPUID templates and Track B density publication remain out of scope here.

Fabric / density (adjacent)​

PriorityFeatureStatus
F1Hubble SID attribution for VM trafficDone β€” from_flow_record_attributed overlays CEP / Cilium-agent SID on hubble observe flows (src + peer VM dst); never writes Cilium private maps
F2Scrape MICROVM_METRICS_ADDR (:9108) from PrometheusDone β€” node-agent annotations + port, deploy/k8s/microvm/servicemonitor.yaml, deploy/prometheus/fluxvm-scrape.yaml

Post-GA closes (2026-09-26)​

Work closed after the numbered S/H/FC/F rows, still part of this ranked set:

CloseStatus
Concurrent NBD device allocDone β€” guestkit flock serialize (#104); lab evidence nbd-concurrent-alloc-20260926.txt
Keep create concurrencyDone β€” ZYVOR_AGENT_SANDBOX_CREATE_CONCURRENCY default 4 + GitHub concurrency CI
Remote seccomp NOTIFY RPCDone — io.zyvor.seccomp.notify.mode=remote guest→host AF_VSOCK (#106); fail closed; optional FLUXVM_SECCOMP_POLICY_RPC_URL
AppArmor fluxctl build-imageDone β€” enforce-mode profile covers guestkit path (#108); smoke scripts/test-apparmor-build-image.sh
H4Done β€” see H table above (#111); not duplicated here

This ranked set β€” closed​

S1–S11 / CNI / DP / F1–F2 / H1–H5 / FC1–FC3 plus the post-GA closes above are closed on the code + gate path. Secure Containers P0/P1 multi-VMM (incl. Firecracker block shares) is shipped; RuntimeClass is GA. The next ranked set is the cross-area backlog below.

Cleared live evidence (archive)​

  1. S2 second-cluster kubeconfig path live on 2026-09-26 β€” benchmarks/evidence/sc-live-s2-real-kubeconfig-20260926.txt (S2 SECOND CNI (kubeconfig): PASS against 80.79.5.173, RuntimeClass runc). Portable nftables stand-in remains β€” sc-live-s2-second-cni-20260926.txt. Lab drop-in /etc/fluxvm-second-cni.env auto-sourced by the evidence script.
  2. Live SC matrix + day-0 wow: sc-live-matrix-20260926.txt (pass=10 skip=0 fail=0, includes s2 + wow); earlier day-0 wow sc-live-wow-demo-20260926.txt; post-#93 soft finish sc-live-phases-post93-summary-20260926.txt (pass=10 skip=1, s2 soft-skip later cleared). Adoption packaging: secure-containers-supported-profile.md, secure-containers-15min-lab.md, sentinel-wedge.md.
  3. S10 live multi-host fleet cleared 2026-09-26 β€” sc-live-s10-fleet-20260926.txt (lab 175.110.122.71 ↔ 80.79.5.173, canary approve/resume, both nodes healthy). S11 live attached migration cleared same day β€” sc-live-s11-migration-20260926.txt (schema-11 cilium-mode fabric, quiesce/export/restore/resume). S3 live prhit directional scrape cleared same evening β€” sc-live-s3-prhit-20260926.txt (directional_counters=1). Multi-CNI under load cleared same evening β€” sc-live-cni-under-load-20260926.txt. S4 EndpointSlice Service-VIP live gate cleared same evening β€” sc-live-s4-endpointslice-20260926.txt. Observer CurrentSchema bumped to 11 with live sc-live-observer-schema11-20260926.txt (schema_compatible=1).

Still open honesty​

  1. H2 virtio-win guest boot is unproven; Cloud Hypervisor stays the production Windows VMM.
  2. mode=remote seccomp live smoke (code Done in #106): the fail-closed deny path passes on the lab (sc-live-seccomp-remote-20260927.txt, chmod: /tmp: Operation not permitted), but a deny can't be told apart from the guest broker's own fail-closed fallback. The run that would prove the RPC path (EXPECT=continue with containerd FLUXVM_SECCOMP_POLICY_RPC=continue) hangs: ctr times out with no guest output. Still open. Suspects: stale containerd-shim-fluxvm-v2 processes that ignore SIGTERM and may still hold vsock port 17780, or a stall in the broker's continue path.
  3. In-tree virtio-fs depth under load still prefers Cloud Hypervisor (H4 honesty: attach/API shipped; full vhost-user queue bind under load is CH SoT).

Next set β€” cross-area backlog (opened 2026-09-27)​

Ranked by operator value vs effort. Tier 1 is being built on feat/fluxctl-cli-parity; Tiers 2–3 are the queue after it.

Tier 1 β€” build now​

#ItemSurface
N1Restart as one opscheduler restart, POST /v1/vms/{id}/restart, fluxctl restart
N2VM metadata PATCH (rename, labels)PATCH /v1/vms/{id}, fluxctl label, fluxctl rename-vm
N3Snapshot list / deleteGET/DELETE /v1/vms/{id}/snapshots[/{tag}], fluxctl snapshot-list, snapshot-delete
N4Events API + watchGET /v1/events, GET /v1/events/stream (SSE), fluxctl events --follow
N5Tenant quota usageGET /v1/quotas/me, fluxctl quota
N6VM name / UUID-prefix resolutionevery fluxctl VM verb
N7Output formatsglobal -o json|table|wide
N8Shell completionsfluxctl completions bash|zsh|fish
N9Wait for statefluxctl wait <vm> --for running|stopped|agent
N10Console resizeTIOCGWINSZ + SIGWINCH β†’ PtyFrame::Resize
N11Doc honesty fixessecure-containers.md remote RPC, PRODUCTION.md CH Windows / 13E
N12Seccomp mode=remote live smokeMODE=remote scripts/e2e-secure-containers-seccomp-notify.sh + SC live matrix (deny passes; continue hangs, see honesty below)

Tier 2 β€” shipped (see operations.md β†’ Day-2 VM operations)​

  • fluxctl --server / FLUXVM_URL remote mode for the core VM verbs.

  • Disks: /v1/vms/{id}/disks list/attach (QMP blockdev-add + scsi-hd hotplug)/resize (block_resize or qemu-img resize)/detach.

  • VMβ†’VM clone (POST /v1/vms/{id}/clone, stopped source, data disks + labels copied).

  • Serial console: QEMU chardev socket teeing into console.log, websocket /v1/vms/{id}/serial, fluxctl serial.

  • Root-disk backup to standalone qcow2 (live via temporary snapshot); label-driven scheduled snapshots with retention.

  • Label selectors (GET /v1/vms?label=) and bulk start/stop/restart/delete -l.

  • GET /v1/openapi.json (hand-maintained OpenAPI 3.1, VM surface).

  • Follow-ups shipped: fluxctl context (named remotes), VM templates (/v1/vm-templates, fluxctl vm-template), backup --all-disks, and serial / events -f / create over --server.

Still open from this tier: backup to S3, console (agent PTY) over --server, VM lifecycle pages in /console, a generated (utoipa) spec covering every route.

Sandlock gap follow-ups​

  • procbox sandboxes: closed by per-sandbox uid, private mount/pid/ipc/uts/net namespaces and seccomp socket-arg filters (see procbox-backend.md). Still open: a live run of a root daemon with the uid pool through the HTTP API, and disk quotas for commands.
  • Dry-run for VM sandboxes: done for native flux-vm, QEMU and Firecracker (snapshot/restore, POST /v1/vms/{id}/restore, live-verified 2026-09-29, see sandbox-changes.md). Cloud Hypervisor sandboxes are unreachable at all (create_sandbox never selects that backend), so there's nothing to extend there. Each dry-run still costs a full-disk copy per snapshot and restore on filesystems without reflink for flux-vm, or a full VMM relaunch for QEMU/Firecracker.
  • HTTPS interception: HTTP/2, a body cap and a transparent (redirect) mode are implemented and verified with real curl in a throwaway namespace (scripts/test-egress-guest.sh, see http-acl.md), including inside the golden KVM guest now that virtio-net has a tap-to-guest receive path (devices/virtio_net.rs::rx_deliver, queue worker RX polling in queue_service.rs); vhost-net's own bind still returns EFAULT before the guest programs its vrings, so this runs over the userspace pump, with a clean fallback log line instead of a half-configured device. The tap-address byte swap (192.168.100.1 becoming 1.100.168.192) is also fixed, with a unit test.
  • procbox learn: policy inference beyond one observed run (merge multiple runs); seccomp-notify variant for non-ptrace environments.
  • Live-verified 2026-09-28: change-set on a real VM guest, both SDKs against real daemons, procbox as an unprivileged daemon, and the egress ACL over HTTP/HTTPS. HTTP/2 and transparent interception were then verified with real curl (see above).

Tier 3 β€” evidence / hardening (lab hardware or long runs)​

  • H2 gate scripts/test-kvm-windows-boot.sh (Tiny11 on fluxvm_engine=kvm + OVMF).
  • In-tree virtio-fs under load bench vs CH (scripts/bench-kvm-virtiofs.sh).
  • Multi-container Pod Sentinel proof; aya ELF parse flake; full-fidelity KVM snapshots (done: FLUXKVM1 v5 captures XSAVE/MSR/LAPIC/TSC/clock/irqchip/PIT; cross-CPU-model portability remains); Firecracker warm-pool claim; OIDC/mTLS; GPU-aware placement.

Portable CI maps each code-side use case to a test target in secure-containers-use-case-matrix.md (enforced by scripts/check-use-case-matrix.sh and .github/workflows/secure-containers-coverage.yml). Live rows stay opt-in via FLUXVM_SECURE_CONTAINERS_LIVE_CI=1.

Hypervisor work should stay Firecracker/CH-matched (no novel device models).