FluxVM eBPF dataplane and Cilium coexistence
FluxVM ships a real TC/eBPF VM-edge dataplane (Network Fabric GA; dataplane schema v4) while keeping the existing nftables path as the backwards-compatible default.
Operator reference with full safety properties: network-fabric.md. Production runbook: production-dataplane.md. CNP / identities: network-policy.md. Diagrams: Packet-decision and control-plane diagrams.
Modesβ
| Mode | Behavior |
|---|---|
legacy (default) | Existing per-sandbox nftables SNAT + optional destination allowlist (CIDR + ports). Stats/flows API unavailable. |
ebpf | Load FluxVMβs TC classifier (bpf/fluxvm_tc.bpf.c), pin programs/maps under pin_root, attach to the host-visible VM interface. |
cilium | Same FluxVM VM-edge eBPF path, but only after verifying the Cilium agent socket and bpffs are visible. Never writes Cilium private BPF maps. |
For Secure Containers Pod CNI on Cilium (L2 handoff of eth0 into the
guest, Multus-safe filtering), see cilium-cni.md β that path
is separate from this VM-edge Fabric mode.
An opt-in bridge-less alternative to that bridge chain (a TC redirect between the Pod veth and the guest
tap, with Cilium's lxc* hooks untouched) is described in direct-datapath.md.
Default remains sandbox.dataplane.mode = "legacy". Existing configs that omit
[sandbox.dataplane] keep using nftables.
Dataplane attach / teardown / reconfigure / reconcile runs for all backends
(QEMU, Cloud Hypervisor, Firecracker, FluxVm) when a host-visible interface
exists. No host-visible iface (network.mode=none / user NAT) always
soft-skips, even when required = true. With an edge present,
required = true fail-closes on attach error (GA). IPv6 CIDRs and Mbps/PPS
limits are native-only and refuse silent nftables downgrade.
GA enable: sudo ./scripts/enable-network-fabric-ga.sh --restart or
configs/network-fabric-ga.toml. Production profile:
configs/network-fabric-prod.toml β production-dataplane.md.
How coexistence fitsβ
flowchart TB
Guest[FluxVm guest] --> Edge[Host VM iface]
Edge --> Fvm["FluxVM TC /sys/fs/bpf/fluxvm"]
Fvm --> Stack[Host routing]
Stack --> Cilium[Cilium / node CNI]
Cilium -.->|private maps untouched| Fvm
What the eBPF path providesβ
- Actual TC classifier source:
bpf/fluxvm_tc.bpf.c - Optional XDP guard:
bpf/fluxvm_xdp.bpf.c - Per-VM pins under
/sys/fs/bpf/fluxvm/vms/<uuid-simple>/(progs/,maps/) - Detach metadata under
/run/fluxvm/ebpf/vms/<uuid-simple>/(iface,prog_id,schema_version,policy_fingerprint) β bpffs cannot store regular files - XDP ownership markers under
/run/fluxvm/xdp/when XDP is enabled (disabled / refused inciliummode) - Host ifindex β stable FluxVM VM identity map
- IPv4 and IPv6 destination-CIDR (LPM) and TCP/UDP destination-port allowlists
- Optional Mbps/PPS fixed-window egress limits (
max_egress_mbps/max_egress_pps; native-only) - Per-CPU allow/drop counters, family-aware LRU flow table, drop/sampled-allow ring buffer
- Live fail-closed policy reconfigure (Running + Paused); schema/fingerprint heal on reconcile; orphan pin GC
GET /v1/vms/{id}/network/status- ARP/DHCP and IPv6 NDP/DHCPv6 always allowed so guests can bootstrap
- Attach without a known guest IP on direct TAP/macvtap (nftables still needs CIDR)
- Maps configured before TC attach (no initial allow window)
- Attach point:
- namespaced TAP β host-side veth
vh<short-id> - direct TAP/macvtap β that host-visible device
- namespaced TAP β host-side veth
- Teardown on VM network cleanup only when program ID still matches FluxVMβs
- Container image builds/installs both
.ofiles; DaemonSet mounts host bpffs and read-only/var/run/cilium; unit/container raise memlock (LimitMEMLOCK=infinity/SYS_RESOURCE)
Why Cilium coexistence (not private-map integration)β
FluxVM VM interfaces are not first-class Cilium endpoints today. Writing Ciliumβs internal maps would couple FluxVM to Cilium implementation details and release-specific map layouts.
Ownership boundary:
- Cilium owns Kubernetes/node CNI and its host datapath.
- FluxVM owns the VM TAP/veth edge (pins under
/sys/fs/bpf/fluxvmonly). - Both may use bpffs; pin namespaces stay separate.
- A later launcher-pod/CNI change can make each VM a native Cilium identity (Hubble-aware) without replacing the native FluxVM dataplane API.
Traffic pathβ
Namespaced VM:
VM -> TAP -> bridge -> namespace veth -> host veth [FluxVM TC/eBPF]
-> host routing -> Cilium/node dataplane
Non-namespaced TAP/macvtap: the classifier attaches directly to the host-visible device.
Build the BPF objectsβ
Debian/Ubuntu:
sudo apt-get install clang llvm libbpf-dev linux-tools-common \
"linux-tools-$(uname -r)" iproute2 nftables
./scripts/build-ebpf.sh
sudo install -D -m 0644 dist/bpf/fluxvm_tc.bpf.o \
/usr/lib/fluxvm/bpf/fluxvm_tc.bpf.o
sudo install -D -m 0644 dist/bpf/fluxvm_xdp.bpf.o \
/usr/lib/fluxvm/bpf/fluxvm_xdp.bpf.o
(bpftool often ships as linux-tools-* rather than a package named bpftool.)
The Dockerfile builder stage runs ./scripts/build-ebpf.sh and installs both
objects under /usr/lib/fluxvm/bpf/. Runtime image includes nftables and
bpftool. systemdβs ReadWritePaths includes /sys/fs/bpf and /run/fluxvm;
LimitMEMLOCK=infinity is required for map load under ProtectSystem=strict.
Configurationβ
# GA profile (see configs/network-fabric-ga.toml)
[sandbox.dataplane]
mode = "ebpf" # legacy | ebpf | cilium
bpf_object = "/usr/lib/fluxvm/bpf/fluxvm_tc.bpf.o"
pin_root = "/sys/fs/bpf/fluxvm"
required = true # fail-closed when a host VM edge exists
default_allow = false
allow_cidrs = ["10.0.0.0/8", "172.16.0.0/12", "192.168.0.0/16"]
allow_ports = ["tcp/443", "tcp/80", "udp/53"]
max_egress_mbps = 250 # native only
max_egress_pps = 100000
sample_rate = 100 # 0 = off; N β 1/N allowed-flow samples
# Optional node-ingress XDP blocklist (not with mode = "cilium")
# [sandbox.dataplane.xdp]
# enabled = true
# interface = "eno1"
# bpf_object = "/usr/lib/fluxvm/bpf/fluxvm_xdp.bpf.o"
# required = true
# block_cidrs = ["198.51.100.0/24", "2001:db8:bad::/48"]
Also merge allowlists from sandbox.egress_allow_domains (resolved to CIDRs) and
any CIDRs passed at apply time.
Kubernetes node with Cilium present:
[sandbox.dataplane]
mode = "cilium"
bpf_object = "/usr/lib/fluxvm/bpf/fluxvm_tc.bpf.o"
pin_root = "/sys/fs/bpf/fluxvm"
required = true
default_allow = true
The DaemonSet mounts host /sys/fs/bpf and read-only /var/run/cilium so FluxVM
can pin programs and see cilium.sock for coexistence checks. See
deploy/k8s/daemonset.yaml.
Policy semanticsβ
- IPv4/IPv6 destination-CIDR egress policy (LPM) and optional L4 destination ports.
- When both CIDR and port lists are non-empty, both must match.
- ARP, DHCP, NDP, and DHCPv6 always allowed for bootstrap.
- IPv6 extension headers fail closed under L4 policy in v3.
Per-VM maps (pinned under each VMβs maps/ directory):
| Map | Role |
|---|---|
fluxvm_id | ifindex β FluxVM identity + policy/rate flags |
fluxvm_v4 | LPM: identity + destination IPv4 prefix β allow |
fluxvm_v6 | LPM: identity + destination IPv6 prefix β allow |
fluxvm_l4 | identity + proto/port β allow |
fluxvm_rate | identity β fixed-window Mbps/PPS state |
fluxvm_stats | per-CPU allow/drop counters |
fluxvm_flows | LRU flow table (family 4/6) |
fluxvm_events | ring buffer for drop / sampled-allow events |
fluxvm_gid | ifindex β up to eight shared group identities |
fluxvm_ct | LRU established 5-tuple table |
fluxvm_deny4 / fluxvm_deny6 | LPM destination deny lists |
REST (see network-fabric.md, network-groups.md, network-policy.md):
GET /v1/vms/{id}/network/policy
POST /v1/vms/{id}/network/policy # admin role when auth is enabled
GET /v1/vms/{id}/network/status
GET /v1/vms/{id}/network/stats
GET /v1/vms/{id}/network/flows?limit=100
GET /v1/vms/{id}/network/effective
GET/POST /v1/network/groups
GET/DELETE /v1/network/groups/{name}
GET/POST /v1/network/cnp
GET/DELETE /v1/network/cnp/{name}
GET /v1/network/identities
GET /v1/network/observe
GET /v1/network/health
GET /v1/network/ipcache
POST /v1/network/refresh-dns
Validationβ
./scripts/validate-network-fabric.sh
FLUXVM_PRIVILEGED_SMOKE=1 ./scripts/validate-network-fabric.sh
sudo -E ./scripts/test-network-fabric.sh
sudo -E ./scripts/test-security-groups-e2e.sh # groups + deny/ICMP maps
python3 scripts/test-network-policy.py # CNP / identity / audit unit
python3 scripts/test-production-dataplane.py
sudo -E ./scripts/test-production-dataplane-e2e.sh
Security-group control plane: network-groups.md. Network policy (CNP / identities / audit): network-policy.md. Production dataplane: production-dataplane.md.
Privileged integration smoke (FluxVm + NetworkSpec::Tap { netns: true }):
- Set
[sandbox.dataplane] mode = "ebpf"(and install the.ofiles). - Create a FluxVm sandbox with netns networking.
- Confirm
tc filter show dev vh<short-id> ingressshowsfluxvm_egress. - Inspect pins under
/sys/fs/bpf/fluxvm/vms/<uuid-simple>/and meta under/run/fluxvm/ebpf/vms/<uuid-simple>/. - Exercise
GET β¦/network/status(schema_version=4,policy_synced). - Delete the VM; pins, meta, and the TC filter should be gone.
- With
required = falseand a missing.o, create should warn and fall back to nftables when fallback is safe.
Netns NAT tables (fluxvm_netns_*) remain independent of sandbox dataplane mode
and continue to use nftables helpers (apply_subnet_masquerade).