FluxVM Secure Containers — Set 8 (Sentinel): in-guest per-container network policy
Set 8 (Sentinel track) is the first genuinely novel piece of the roadmap —
not something Kata attempts. Through Set 7, nothing stopped one container in
a multi-container Pod from reaching another container's local ports: the
only network enforcement was Set 6S's Pod-wide policy at the VM edge, which
by definition can't see traffic that never leaves the guest. Set 8S closes
that gap with per-container cgroup_skb programs enforced inside the
guest kernel.
The build/tooling exception: aya
The project-wide rule (crates/fluxvm-network/src/ebpf.rs: "bpftool+tc, not
aya/libbpf") is scoped to avoid linking libbpf into the shared host
dependency graph. fluxvm-container-agent is the opposite case: a single,
purpose-built binary the shim uploads fresh into the guest on every boot
(docs/secure-containers.md), not part of that shared graph — and the
minimal guest image cannot be relied on to have bpftool/iproute2/kernel
headers installed the way the host always does. This crate is a deliberate,
narrowly-scoped exception: it depends on aya (a pure-Rust eBPF loader) to
attach bpf/fluxvm_guest_cgroup.bpf.c — compiled with the exact same
clang/BTF-map conventions as every other bpf/*.bpf.c object in this repo,
not written in Rust against aya-ebpf.
That combination (C-compiled object, Rust-based loader) is not aya's
primary documented workflow, and an old GitHub issue suggested clang objects
might not load cleanly. Verified empirically rather than trusted from
that issue's age: a minimal clang-compiled, BTF-defined-map cgroup_skb
object was loaded, submitted to the real kernel verifier, attached to a
live cgroup, and round-tripped through a map — all successfully — using
aya 0.13.1 on a real Linux 6.8 host. The full Set 8S integration was then
built and tested the same way, not assumed to work by analogy.
What was built
bpf/fluxvm_guest_cgroup.bpf.c: two programs,cgroup_skb/egressandcgroup_skb/ingress, both keyed bybpf_skb_cgroup_id()(the inode number of the cgroup that owns the packet's socket).bpf_get_current_cgroup_id()was used originally, which is wrong for ingress: that hook runs in softirq context under whichever task was interrupted, so inbound packets were judged against an unrelated cgroup's (usually absent, i.e. deny-all) policy. It went unnoticed because container workloads were never actually inside their container cgroup (see the CHANGELOG "Fixed" entry) and so were never subject to this policy. Maps:fluxvm_cpol(per-container policy flags),fluxvm_cid4/fluxvm_cid6(peer allow/deny, mirroring Set 6S'sfluxvm_pid4/fluxvm_pid6shape),fluxvm_cdrops(per-container drop counters).- Fail-closed by default — the opposite of Set 6S's "unconfigured = allow"
default. Set 6S's default exists for backward compatibility with the
large existing population of VMs that never opt into Pod-scoped policy at
all. There's no equivalent concern here: a container's cgroup only ever
gets these programs attached because Set 8S is in use for it, so an
unconfigured
fluxvm_cpolentry denies non-loopback traffic rather than allowing it. - One shared loaded instance per Pod VM.
fluxvm-container-agentloadsfluxvm_guest_cgroup.bpf.o(compiled atcargo buildtime bycrates/fluxvm-container-agent/build.rsviascripts/build-ebpf-guest.shand embedded withinclude_bytes!, so the uploaded binary stays a single self-contained file) once, on the first container's creation, thenattach()s the same already-loaded programs again for every later container's own cgroup — safe because the maps are keyed by cgroup id. - Attach point:
create_container_cgroupinfluxvm-container-agent, right before the process is spawned, so policy is live from the container's first packet. Best-effort like the resource-limit cgroup itself: a guest kernel missingcgroup_skb/BTF support still runs the container, just without this layer (a platform-capability gap, not the fail-closed-policy-content design, which only governs what an attached program does when unconfigured). ContainerNetworkPolicy(fluxvm-container-protocol): a new, optional field onContainerRequest::Create—default_allow,audit_mode,allow_addresses,deny_addresses. Cleaned up on container delete (forget_container_policy) so the shared maps don't grow unboundedly across a long-lived Pod VM's container churn.
Deliberately out of scope for this Set (closed by Set 19)
At the time this Set shipped, the shim did not yet fetch a Pod's Set 6S
policy and forward it per container — every container got Set 8S's
fail-closed-by-default enforcement attached unconditionally, but with an
empty policy (network_policy: None from the shim), matching Set 6S's own
precedent of shipping the mechanism ahead of its primary population source.
This wiring has since landed, in Set 19 (fluxvm-containerd-shim's
guest_network_policy): the shim queries GET /v1/vms/{id}/network/ pod-policy, maps the response into a ContainerNetworkPolicy (inverting
default_deny into default_allow, de-duplicating rules, carrying
schema_version/ingress_isolated/egress_isolated), and passes it into
each container's own Create call — the call site's own comment states the
design plainly: "Set 19: fetch the current host Pod policy before Create.
Failure is fail-closed because None keeps Set 8S's deny-non-loopback
default." A second path, spawn_network_policy_watch, continuously polls
the same endpoint after creation and pushes a live ContainerRequest:: UpdateNetworkPolicy whenever the fetched policy changes, so a container's
enforcement doesn't go stale for the rest of its lifetime. See Set 19's own
notes for that work's full scope.
Validation performed
- Compatibility probe (see above): a standalone clang-compiled
cgroup_skbobject, loaded/attached/map-tested via a minimalayaprogram on a real Linux 6.8 host. - Kernel verifier: both real programs (
fluxvm_guest_egress,fluxvm_guest_ingress) passbpftool prog loadindependently of the aya path, confirming the C code itself is correct regardless of loader. - Full Rust integration, end-to-end, not mocked: a new
#[test] guest_cgroup_policy_attaches_and_enforcesinfluxvm-container-agent(root-only, self-skips otherwise) creates a real throwaway cgroup, calls the actualattach_guest_network_policy/configure_container_policy/forget_container_policyfunctions, moves the test process into the cgroup, and proves real enforcement: a non-loopback connect is blocked under a default-deny policy, a loopback connect (through the same process's own egress and ingress, both in the policed cgroup) succeeds, and cleanup leaves no stale map entries. Run withsudoon a real Linux host to exercise it. cargo build/cargo testand a full workspace build pass for every touched crate on a real host..github/workflows/secure-containers.ymlnow installsclang(previously not needed there) sincefluxvm-container-agent'sbuild.rscompiles a BPF object atcargo buildtime.- Not yet run: a live multi-container Pod e2e test proving container B
is denied from reaching container A's port by policy — the mechanism this
would exercise is the same one the root-only unit test above already
validates directly; the multi-container Pod scenario additionally needs a
live Kubernetes cluster with the
fluxvmRuntimeClass installed, which wasn't set up against the already-running FluxVM host used for other Set 6/7/8 validation.
Remaining gates
- The live multi-container Pod e2e test named in "Validation performed"
above (container B denied from reaching container A's port, proven
against a real Kubernetes cluster with the
fluxvmRuntimeClass installed) — the per-container policy-forwarding wiring itself is no longer an open gate (see "Deliberately out of scope" above), only its live multi-container-Pod proof is. CLONE_NEWUSER/full namespace isolation (Set 6R) and this Set's cgroup scoping are complementary, not sequenced — Set 6R's own doc already notes where a future per-container LSM identity model would hook in; this Set's cgroup-id-keyed scoping needs no changes when that lands.