FluxVM Secure Containers — containerd runtime v2
A Pod. Its own guest kernel. (GA)
FluxVM Secure Containers is the first container-runtime layer on top of the existing FluxVM VM lifecycle, VSOCK guest agent and QEMU virtiofs support. The goal is the same security shape users expect from Kata Containers: a Pod or container group gets its own guest kernel instead of sharing the node's host kernel — plus, starting with Set 6, FluxVM Sentinel: an eBPF-based policy substrate spanning the host VMM edge and (in later Sets) the guest kernel with one shared schema and identity model, which Kata's namespace + seccomp + static-policy-file model does not attempt.
Architecture
Kubernetes / ctr
|
containerd CRI
|
containerd-shim-fluxvm-v2
|
+---- (optional) CNI L2 bridge prep on Pod netns
+---- FluxVM REST API ----> QEMU/KVM microVM
| |
| +-- virtiofs Pod share (+ Pod-UID volumes)
| +-- TAP on host CNI bridge (when netns present)
| +-- fluxvm-guest-agent :17777
| +-- fluxvm-container-agent :17778 (lifecycle)
| +-- stdio streams :17779 (TTY/pipes)
| |
+---------------- VSOCK ------------------+
|
OCI processes + guest cgroup v2
The shim groups tasks using the same annotations containerd uses for a
Kubernetes Pod (io.containerd.runc.v2.group first, then
io.kubernetes.cri.sandbox-id). The first task creates one QEMU FluxVM with a
Pod-scoped virtiofs directory. Later tasks in the group reuse the same VM.
Why a separate container agent?
FluxVM's existing fluxvm-guest-agent is a stable VM-management interface for
ping, command execution, file transfer, console and shutdown. Secure
Containers does not expand that stable protocol. Instead the shim uploads a
small fluxvm-container-agent into the guest through the existing authenticated
agent, launches it on VSOCK port 17778, and uses a dedicated lifecycle protocol.
That makes the feature removable and independently versionable while reusing the VM's existing per-instance authentication token.
Implemented through Set 19
Everything through Set 13 below, plus:
- Set 14 — NetworkPolicy v2 directional CIDR+L4 rules (
ipBlock.except, SCTP, named/endPortports) and a separate stateful Pod-ingress BPF object — see docs/secure-containers-set14.md. - Set 15 — directional attachment health (
pod_ingress_*onNativeAttachmentStatus) and the read-only Sentinel Policy Observer — see docs/secure-containers-set15.md. - Set 16 — schema-v9 timestamped conntrack with idle expiry, synchronous
clear on policy update, SYN/INIT anti-replay, and VM-level
sctp/PORT— see docs/secure-containers-set16.md. - Set 17 — optional
fluxvm_prhitrule-attributed directional telemetry (no schema-v8 ABI bump; Observer keeps Set 15 shared-counter fallback) — see docs/secure-containers-set17.md. - Set 18 — EndpointSlice-aware opt-in Service VIP admission for egress peers (controller-only; no dataplane ABI change) — see docs/secure-containers-set18.md.
- Set 19 — GA completion candidate: schema-v10,
fluxvm_pridxrule index, IPv6 extension-header walk, Set 8S guest policy mirror, Observer ServiceMonitor/sizing, fail-hard GA runner — see docs/secure-containers-set19.md.
Implemented through Set 13
- containerd runtime-v2 binary:
containerd-shim-fluxvm-v2 - one Kubernetes Pod/containerd shim group -> one FluxVM QEMU VM
- CNI L2 Pod IP/MAC handoff into the guest
- Pod/sandbox and per-container guest cgroup-v2 limits, stats, pids and updates
- containerd task events + asynchronous exit publication
- OCI create/start/state/wait/kill/pause/resume/delete and exec lifecycle
- uid/gid, supplementary groups, cwd, argv, PATH, environment, umask and rlimits
- Linux capability sets, noNewPrivileges, seccomp syscall + argument-comparator rules
- read-only rootfs, masked/read-only paths, device nodes and Pod-VM sysctls
- containerd snapshot rootfs staging into the VM-visible virtiofs share
- Pod-UID-scoped write-through Kubernetes
volumes/volume-subpaths - authenticated VSOCK lifecycle RPC on port 17778
- authenticated raw stdin/stdout/stderr streaming on port 17779
- real guest PTY for terminal init/exec processes
- containerd
ResizePtyandCloseIOforwarding - legacy regular-file virtiofs stdio fallback for non-TTY debugging
- RuntimeClass/containerd examples, source/unit gates and KVM smoke scripts
- per-container PID/mount/IPC/UTS namespace isolation (
pivot_root, notchroot); containers in a Pod no longer rely on the VM boundary alone - sandbox/pause-container convention: the CRI sandbox container's IPC/UTS
(and PID, when
shareProcessNamespace: true) are joined by sibling containers, matching runc/Kata's "join the pause container" model - opt-in
CLONE_NEWUSERper container (FLUXVM_CONTAINER_USERNS=1, off by default) for an additional capability-scoping boundary - Sentinel: identity-aware, Kubernetes-NetworkPolicy-shaped Pod network
policy on the VM edge (
bpf/fluxvm_pod_policy.bpf.h, additive on top of the existing CIDR/L4 dataplane), with a stable per-Pod identity and an independent policy-content API (/v1/vms/{id}/network/pod-policy) — see docs/secure-containers-set6s.md. - Sentinel Set 13: a standalone node-local
fluxvm-networkpolicy-controller(dependency-free Go, DaemonSet, read-only RBAC) that compiles selectedNetworkPolicyobjects into FluxVM Pod policy. Superseded for ingress/CIDR/endPort/named-port/SCTP by Set 14 — see docs/secure-containers-set13.md and docs/secure-containers-set14.md. - Sentinel Set 15: directional
fluxvm_pod_ingressattachment health onNativeAttachmentStatus, plus the read-only Policy Observer (tools/fluxvm-policy-observer) — see docs/secure-containers-set15.md. - Sentinel Set 16: schema-v9 revocation-safe conntrack + VM-level SCTP — see docs/secure-containers-set16.md.
- Sentinel Set 17: optional
fluxvm_prhitrule-hit / directional counters on the Policy Observer — see docs/secure-containers-set17.md. - Sentinel Set 18: EndpointSlice-proven Service VIP admission for
--include-service-clusterips— see docs/secure-containers-set18.md. - Sentinel Set 19: schema-v10 +
fluxvm_pridx+ IPv6 ext-hdr walk + guest Pod-policy mirror + Observer ops / GA runner — docs/secure-containers-set19.md. - restart-safe runtime ownership journal and IPv4/IPv6 dual-stack CNI replay on the primary interface — see docs/secure-containers-set6.md
- Sentinel: per-VM QEMU process hardening — a
BPF_CGROUP_DEVICEallowlist (only/dev/kvm//dev/vhost-vsock//dev/net/tun/configured VFIO devices) and acgroup_skbegress filter (loopback-only outbound IP) on the samefluxvm.slice/{id}.scopecgroup every VM already gets — see docs/secure-containers-set7s.md. Applies to every VM, not only Secure Containers Pods. - containerd
TaskOOMpublication from guest cgroup-v2memory.events, CPU throttling + detailed memory/swap/fault task metrics, OCI memory reservation/swap translation, allowlisted cgroup-v2unifiedupdates, and fail-closed raw block/special-device handling — see docs/secure-containers-set7.md - VM boot pipelined with the host-side rootfs mount+copy instead of strictly sequential — see docs/secure-containers-set7r.md.
- Sentinel: in-guest, per-container
cgroup_skbnetwork policy (fail-closed by default) so one container in a multi-container Pod can no longer reach another's local ports unchecked — the first Sentinel piece Kata has no equivalent of. Usesayaas a deliberate, narrowly-scoped exception to the project's "no aya/libbpf" rule. Not yet wired to Set 6S's Pod policy as an automatic per-container default — see docs/secure-containers-set8s.md. - Pod-scoped raw block QMP hotplug and allowlisted VFIO PCI passthrough,
reference-counted per container with QEMU
DEVICE_DELETED-confirmed hot-unplug — see docs/secure-containers-set9.md - fail-closed AppArmor/SELinux process labels and generated per-container
cgroup-v2
BPF_PROG_TYPE_CGROUP_DEVICEpolicy from OCIlinux.resources.devices— see docs/secure-containers-set10.md - a bounded in-guest seccomp user-notification broker for explicit
SCMP_ACT_NOTIFYrules (listener fd handed off viaSCM_RIGHTSto a supervisor thread outside the workload's own filter/chroot, fail-closed by default), OCIlinux.mountLabelSELinux mount labels, and aSecurityStatslifecycle counter snapshot — see docs/secure-containers-set11.md
Current limitations
Status: GA. Scope boundaries below — not a claim of full Kata Containers compatibility.
Supported production profile: see secure-containers-supported-profile.md (QEMU/CH primary, allowlisted hostPath, in-cluster NetworkPolicy → Sentinel, isolation without TEE). Flip runbook: secure-containers-flip-runtimeclass.md. Sentinel wedge: sentinel-wedge.md.
- QEMU, Cloud Hypervisor, and Firecracker are supported Secure Containers
VMMs. Set
FLUXVM_CONTAINER_BACKEND=qemu|cloud-hypervisor|firecracker. Firecracker has no virtio-fs upstream: Pod shares are packed to ext4 block images at launch (write-through hostPath/live sync is QEMU/CH only). Firecracker also needs a Linux kernel viaFLUXVM_CONTAINER_KERNELor the daemon'sfirecracker_kernelconfig. - CNI coverage is Cilium-, Calico-, and Flannel-aware on the primary
interface. The L2 handoff auto-detects the provider
(
FLUXVM_CONTAINER_CNI_PROVIDER=auto|cilium|calico|flannel|generic), preferseth0, and treats MultusnetNsecondaries as extra guest NICs. Multus secondaries ride a hybrid path with a direct primary (bridge by default; setFLUXVM_CONTAINER_CNI_MULTUS_DATAPATH=direct|autofor per-secondary direct when the iface is a veth). Unusual non-veth layouts still fall back to the bridge chain. Firecracker cannot attach a bridge-less direct tap by fd — use tap/bridge or QEMU/CH for direct. - Arbitrary hostPath passthrough is allowlisted. Set
FLUXVM_HOSTPATH_ALLOW(colon-separated roots) at sandbox create, or use the dynamic broker (FLUXVM_HOSTPATH_BROKER=1, default) to hotplug an annotation allowlist root into a running sandbox (QEMU/CH). Outside the allowlist, binds remain copied or fail closed when an allowlist is in force. - Device-cgroup enforcement landed in Set 10. PID/mount/IPC/UTS
namespace isolation landed in Set 6 (see
docs/secure-containers-set6r.md), and user
namespace isolation is available opt-in. Per-container device access
control is now a real
BPF_PROG_TYPE_CGROUP_DEVICEprogram compiled from OCIlinux.resources.devicesand attached to the per-container guest cgroup, not just the existing static device-node bind mounts. - Seccomp argument comparators and user notification are both
implemented. All OCI comparison operators (
SCMP_CMP_*) are enforced viaseccomp_rule_add_array.SCMP_ACT_NOTIFYis served by a bounded in-guest broker (Set 11);SCMP_ACT_NOTIFYasdefaultActionand NOTIFY onsendmsgare both rejected at OCI-parse time rather than risking a listener-bootstrap deadlock.SECCOMP_IOCTL_NOTIF_ADDFDis available viaio.zyvor.seccomp.notify.mode=addfd; remote policy RPC viaio.zyvor.seccomp.notify.mode=remote(guest→host AF_VSOCK; fail closed). - TTY streaming is new in Set 5 and requires real KVM/containerd churn and resize testing on self-hosted nodes before production claims.
CLONE_NEWUSERis opt-in. Live gates (scripts/e2e-secure-containers-userns-load.sh) prove the shim forwardsFLUXVM_CONTAINER_USERNS=1and the guest agent creates a user namespace. Cross-VM inode numbers often collide; setFLUXVM_USERNS_REQUIRE_DISTINCT=1for multi-container-in-one-VM layouts (FLUXVM_USERNS_MODE=k8s-multi). RuntimeClass load:FLUXVM_USERNS_MODE=k8s(lab evidence 2026-09-26).- Raw block/device-plugin passthrough is scoped, not general. Pod-scoped
raw block volumes are hotplugged over QMP from the owning Pod's
volumeDevicestree (or an explicit operator allowlist prefix); VFIO PCI character-device passthrough requires an exact BDF allowlist and a device already bound tovfio-pci. Anything outside those allowlists still fails closed instead of silentlymknod-ing an unrelated node.
Next gates to production Kata-style Kubernetes support
Code-side P0/P1 phases below are closed. Remaining work is live-node
evidence under FLUXVM_SECURE_CONTAINERS_E2E=1 / lab scripts — not new VMM
surface.
| Phase | Status | Notes |
|---|---|---|
| P0 OCI fixtures + parsers | Done | tests/oci-fixtures/ + agent parser tests |
| P0 CNI (Cilium/Calico/Flannel/Multus) | Done | providers + Multus hybrid/direct; live under load sc-live-cni-under-load-20260926.txt |
| P0 hostPath broker | Done | create-time allowlist + QEMU/CH hotplug broker |
| P0/P1 multi-VMM (QEMU/CH/FC) | Done | FC = ext4 block shares + FLUXVM_CONTAINER_KERNEL |
| P1 seccomp NOTIFY + ADDFD | Done | mode=remote guest→host AF_VSOCK RPC |
| P1 warm-pool claim | Done | QEMU/CH; FC cold-create only |
| Live e2e matrix (TTY, SELinux, CNI load, userns) | Lab archived | docs/benchmarks/evidence/sc-live-*.txt + sc-hotcake-bundle-20260926.txt: ctr/TTY/namespaces/CNI-churn/NOTIFY; SELinux mountLabel green; k8s userns Succeeded; k8s-multi distinct_user_ns=2; guestkit Fedora/btrfs token inject. Multi-CNI under load prefers FLUXVM_SECOND_CNI_KUBECONFIG (nftables stand-in portable default) |
| Remote policy RPC | Done | mode=remote; FLUXVM_SECCOMP_POLICY_RPC / _URL; fail closed |
| Phase orchestrator | Done | scripts/complete-secure-containers-phases.sh |
P0 — OCI conformance fixtures
PID/mount/IPC/UTS namespace isolation (opt-in CLONE_NEWUSER) and
per-container device-cgroup enforcement (BPF_PROG_TYPE_CGROUP_DEVICE) have
both landed, with checked-in fragments under tests/oci-fixtures/. Remaining:
live-node matrix coverage under FLUXVM_SECURE_CONTAINERS_E2E=1 (combined
controls on real KVM/containerd), not more fragment files.
P0 — CNI conformance
Cilium / Calico / Flannel providers auto-detect
(FLUXVM_CONTAINER_CNI_PROVIDER); Multus netN secondaries attach as extra
guest NICs (hybrid bridge default, optional per-secondary direct via
FLUXVM_CONTAINER_CNI_MULTUS_DATAPATH). Portable churn:
scripts/evidence-cni-churn.sh. Live under-load evidence:
scripts/evidence-cni-under-load.sh /
sc-live-cni-under-load-20260926.txt.
P0 — volume broker / explicit hostPath
Pod-scoped kubelet volume exports remain the safe default. Create-time
FLUXVM_HOSTPATH_ALLOW (colon-separated) plus optional dynamic broker
(FLUXVM_HOSTPATH_BROKER=1, QEMU/CH hotplug) are implemented. Firecracker packs
allowlisted roots at launch (no hotplug).
P1 — seccomp notification broker (landed, Set 11)
Argument comparators (Set 10), bounded fail-closed SCMP_ACT_NOTIFY (Set 11),
and SECCOMP_IOCTL_NOTIF_ADDFD (io.zyvor.seccomp.notify.mode=addfd) are
implemented — see docs/secure-containers-set11.md.
Remote policy RPC landed later as io.zyvor.seccomp.notify.mode=remote
(guest→host AF_VSOCK, port 17780, fail closed); live smoke via
MODE=remote scripts/e2e-secure-containers-seccomp-notify.sh (fail-closed deny
passes live; the EXPECT=continue proof is still open — see NEXT-FEATURES.md). Live-node gates (enforcing-SELinux
guest mount label, mid-flight kill through full create/delete) still need
containerd/Kubernetes Pod runs, not just the guest-agent binary in isolation.
P1 — streaming/TTY conformance
Run repeated init + exec terminal resize, stdin close, high-volume stdout/stderr
and abrupt-exit tests on self-hosted KVM/containerd nodes
(scripts/e2e-secure-containers-tty.sh).
P1 — performance
Set 7 (Runtime) pipelines VM boot with the host-side rootfs copy (no longer
strictly sequential) — see
docs/secure-containers-set7r.md. Warm-pool
claim is opt-in: set FLUXVM_CONTAINER_WARM_POOL to a pool whose template
uses network.mode=none. The shim claims a paused member, then
POST /v1/vms/{id}/hotplug/nic for the Pod CNI bridge (and each Multus
netN). Warm-pool claim requires QEMU or Cloud Hypervisor (share hotplug);
Firecracker uses cold create with ext4-packed shares. RuntimeClass is GA,
not a Kata-equivalent product.
P1 — additional VMMs
QEMU, Cloud Hypervisor, and Firecracker are all selectable via
FLUXVM_CONTAINER_BACKEND. Firecracker packs Pod shares to ext4 at launch
(no live virtio-fs) and needs FLUXVM_CONTAINER_KERNEL (or daemon
firecracker_kernel). See deploy/containerd/env.example. Remaining honesty
bounds are live e2e gates — not a fourth VMM.
Installation
sudo ./scripts/install-secure-containers.sh
Set the guest image explicitly when needed:
export FLUXVM_CONTAINER_GUEST_IMAGE=/var/lib/fluxvm/images/secure-container.qcow2
export FLUXVM_API_URL=http://127.0.0.1:7788
# export FLUXVM_API_TOKEN=... # if FluxVM API auth is enabled
# Firecracker: also set a Linux kernel (or configure firecracker_kernel on the daemon)
# export FLUXVM_CONTAINER_BACKEND=firecracker
# export FLUXVM_CONTAINER_KERNEL=/var/lib/fluxvm/images/vmlinux
Merge deploy/containerd/fluxvm-runtime.toml into containerd configuration,
restart containerd, and install deploy/containerd/runtimeclass.yaml. Prefer
secure-containers-flip-runtimeclass.md
or scripts/provision-secure-containers-lab.sh for Absolute BinaryName on k3s.
Test levels
Level 1 — host-independent
./scripts/test-secure-containers.sh
This builds both binaries and runs all new unit tests.
Level 2 — KVM/containerd host
Set FLUXVM_SECURE_CONTAINERS_E2E=1; the script validates KVM, containerd and
FluxVM API prerequisites and runs scripts/e2e-secure-containers-ctr.sh. Run
scripts/e2e-secure-containers-volume.sh for write-through PVC coverage,
scripts/e2e-secure-containers-tty.sh for VSOCK stdio + init/exec PTY/resize,
and scripts/e2e-secure-containers-namespaces.sh for Set 6 PID/mount/IPC/UTS
isolation (both need a Kubernetes cluster with the fluxvm RuntimeClass
applied, not just bare ctr). Set
FLUXVM_SECURE_CONTAINERS_SECURITY_E2E=1 to additionally run
scripts/e2e-secure-containers-security.sh (Set 10 seccomp/AppArmor/SELinux/
device-cgroup enforcement), and FLUXVM_SECURE_CONTAINERS_SET11_PREFLIGHT=1
to run scripts/preflight-secure-containers-set11.sh (guest kernel/
libseccomp/libselinux readiness for Set 11 NOTIFY/mountLabel).
The GitHub-hosted CI job intentionally does not claim KVM end-to-end coverage, because ordinary hosted runners do not provide the nested virtualization and node configuration needed for an honest containerd+FluxVM test.
Set 4 addendum — write-through Pod volumes + OCI security
Set 4 adds Pod-UID-scoped virtiofs exports for kubelet volumes/ and
volume-subpaths/, so matching Kubernetes volume bind mounts are write-through
instead of copied. Arbitrary host binds remain snapshot-based by default.
The guest agent also enforces read-only rootfs, masked/read-only paths, OCI
device nodes, Pod-VM sysctls and libseccomp syscall-name rules. Seccomp profiles
using unsupported argument comparators fail closed. See
docs/secure-containers-set4.md for the exact support boundary and test gates.
Set 5 addendum — VSOCK stdio + TTY/PTY
Set 5 moves normal runtime stdio off virtiofs polling files. Lifecycle RPC stays
on port 17778 while raw process streams attach on authenticated VSOCK port
17779. Non-TTY tasks use guest pipes; terminal tasks use a real guest PTY and
containerd ResizePty maps to TIOCSWINSZ. See
docs/secure-containers-set5.md for protocol, fallback and test details.
Runtime recovery + dual-stack CNI addendum
Persists an atomic versioned sandbox/task/exec ownership journal (including
ready=false provisional VM ownership), validates it against the live FluxVM
before reuse, and fails closed rather than risking a duplicate Pod VM when
recovery cannot be proven safe. Restores task/exec exit watchers and attempts
VSOCK stdio reattachment after a replacement shim launches. The primary CNI
interface now captures/replays IPv4 + IPv6 addresses and family-specific
routes; additional routable interfaces are rejected by default
(FLUXVM_CONTAINER_CNI_STRICT_MULTI_INTERFACE=0 to relax in controlled labs
only). See docs/secure-containers-set6.md.
OOM + cgroup metrics + cleanup-safety addendum
Publishes containerd TaskOOM from guest cgroup-v2 memory.events with a
journaled oom_kill cursor that survives replacement-shim recovery. StatsTask
gains CPU throttling and detailed memory/swap/fault/pids statistics. OCI
memory.reservation maps to memory.low, OCI memory+swap is converted to
cgroup-v2 swap-only memory.swap.max, and LinuxResources.unified updates are
restricted to a small explicit allowlist. Raw block and special host-device
binds are rejected instead of copied, and character devices are validated
against the guest's actual attached major/minor. Sandbox teardown retains
journal/CNI ownership state when FluxVM VM deletion fails, so cleanup can be
retried safely. See docs/secure-containers-set7.md.
Set 8 addendum — raw block + VFIO devices
Set 8 replaces the prior raw/special-device fail-closed placeholder with real
QEMU hotplug for Pod-scoped raw block volumes and explicitly allowlisted VFIO
PCI devices. Raw block sources are resolved only from the owning Pod's
volumeDevices tree (or an explicit operator prefix), attached as SCSI devices,
and rediscovered inside the guest by stable serial. VFIO character-device
passthrough requires an exact BDF allowlist and a device already bound to
vfio-pci; host drivers are never detached automatically. See
docs/secure-containers-set8.md.
Set 9 addendum — passthrough device lifecycle
Set 9 reference-counts raw-block/VFIO attachments by container, waits for QEMU DEVICE_DELETED before final backend cleanup, journals delayed unplug for retry, validates raw-block identity and complete VFIO IOMMU groups on recovery, and supports guest-driver GPU companion nodes without copying host major/minor numbers. See secure-containers-set9.md.
Set 10 addendum — guest security enforcement
Set 10 adds all OCI/libseccomp argument comparison operators, fail-closed
AppArmor and SELinux process labels, and generated per-container cgroup-v2
device eBPF (BPF_PROG_TYPE_CGROUP_DEVICE). Unsupported seccomp notify and
SELinux mount labels remain explicit fail-closed boundaries. See
secure-containers-set10.md.
Set 11 addendum — seccomp notification + SELinux mount labels
Set 11 adds a guest-agent seccomp user-notification supervisor for explicit
SCMP_ACT_NOTIFY syscall rules. Listener fds are transferred to the unfiltered
supervisor with SCM_RIGHTS; default behavior denies with EPERM, while the
vendor annotation io.zyvor.seccomp.notify.mode=continue explicitly opts into
kernel CONTINUE semantics. NOTIFY as the filter default and NOTIFY on sendmsg
remain rejected because they can deadlock listener bootstrap.
OCI linux.mountLabel is applied as context="<label>" to mounts created by the
guest OCI setup path after SELinux/libselinux preflight. It does not claim to
retroactively relabel the already-mounted virtiofs rootfs. The lifecycle
protocol also exposes VM-local cumulative security counters, which the shim logs
on init-task delete. See secure-containers-set11.md.
Set 14 addendum — NetworkPolicy v2
Set 14 replaces Set 13's exact-peer/exact-port model with directional CIDR+L4
tuple rules (ipBlock.except, SCTP, named/endPort ports) and a separate
fluxvm_pod_ingress.bpf.o sharing the main program's maps/conntrack. See
secure-containers-set14.md.
Set 15 addendum — policy observability
Set 15 makes attachment_status() fail unhealthy when Pod-ingress is pinned
but its TC hook is missing, and ships the read-only Policy Observer
(tools/fluxvm-policy-observer). See
secure-containers-set15.md.
Set 16 addendum — conntrack revocation safety
Set 16 moves shared conntrack to schema-v9 timestamped entries with
per-protocol idle timeout, clears the table on policy update, and forces
TCP SYN / SCTP INIT through current policy. VM-level allow_ports accepts
sctp/PORT. See secure-containers-set16.md.
Set 17 addendum — rule-attributed telemetry
Set 17 adds optional fluxvm_prhit so the Observer can export true
per-direction and per-rule packet counters without a schema-v8 ABI bump.
Live production scrape and multi-node / stateful CT proof remain open —
secure-containers-set17.md and
NEXT-FEATURES.md.
Set 18 addendum — EndpointSlice Service VIP policy
Set 18 makes opt-in --include-service-clusterips admit ClusterIPs only
when discovery.k8s.io/v1 EndpointSlice backends prove every routable
endpoint is an allowed non-terminal Pod. No dataplane schema change —
secure-containers-set18.md.
Set 19 addendum — GA completion candidate
Set 19 bumps the native dataplane to schema-v10 with fluxvm_pridx,
bounded IPv6 extension-header walking, Set 8S guest CIDR/direction mirror
(create + live), and Observer ServiceMonitor/sizing. External S1/S2/S9–S11
lab evidence remains mandatory —
secure-containers-set19.md,
secure-containers-use-case-matrix.md, and
NEXT-FEATURES.md.