Skip to main content

⚡ Direct (bridge-less) datapath

Cilium delivers to a Pod without the host stack: bpf_host does bpf_redirect_peer(lxc*) and the packet lands in the Pod's eth0. A Secure Container's guest used to sit behind a bridge chain after that. The direct datapath removes the chain: a TC redirect moves frames straight between the outer device and the guest's tap, and Cilium keeps every hook it has on the host-side lxc* peer. FluxVM never writes Cilium maps (cilium-cni.md).

Two attach modes share one mechanism:

ModeOuter deviceWhereUse
peer-veththe Pod's veth eth0tap created inside the Pod netnsSecure Containers under Cilium/any veth CNI
l2-uplinka physical/bond NIC with no bridge mastertap in the host netnsstandalone VMs replacing vmbr0

✅ What works​

PieceBehavior
Shim knobFLUXVM_CONTAINER_CNI_DATAPATH=bridge|direct|auto (default auto)
autoDirect when every precondition holds, else the bridge chain with the reason logged; a create or warm-pool hotplug the daemon rejects is retried once on the bridge chain
directDirect or the Pod fails to start (no silent fallback)
Preconditionskernel ≥ 5.10 (bpf_redirect_peer), the primary interface is a veth, no Multus secondaries
Pod topologyeth0 keeps its MAC; its addresses/routes move to the guest (restored on teardown). No host bridge, no extra veth pair, no in-Pod bridge
Tap handoffCreated by the daemon inside the Pod netns and given to QEMU as an inherited fd (QEMU runs in the host netns and cannot open it by name). Non-persistent: it disappears with the VMM
Warm poolThe claimed VM (booted network.mode=none) gets its NIC over QMP: getfd (the tap as SCM_RIGHTS) → netdev_add fd=… → device_add, all on one QMP session because QEMU resolves a named fd through the monitor that received it. The dataplane is applied before QEMU sees the NIC
PolicyGuest→out runs fluxvm_egress (allow-list, L4, rate, conntrack, stats) before the redirect. Out→guest passes the tap's egress hook, so the Set 14 Pod-ingress program still runs
RecoveryThe wiring is recorded per VM (direct.json), so repair, reconfigure, status and remove re-enter the right netns and re-apply the redirect config
AttachTCX when available, else legacy clsact at a reserved pref with a per-VM handle. Both are tested
l2-uplinkSeveral VMs share one unbridged uplink: inbound frames are steered by destination MAC, ARP requests by their target IP (direct.guest_ips), guest↔guest is switched locally. The MAC/IP maps are shared per uplink and each VM adds/removes only its own entries
CRDMicroVM.spec.networkMode: direct with parent (the uplink), mac, guestIps
Hop viewhubble observe / packet-flow hops show the real path (no invented bridge)
Capability probetools/fluxvm-sentinel-certify.py reports redirect_peer / redirect_neigh from the kernel's own helper list
Linux 7.xThe VM-edge program loads on Linux 7.0 (see the verifier note)

⚠️ What does not work (honesty bounds)​

  • The shim default is auto, on the strength of two live runs. On a single-node k3s + Cilium v1.20.2 (veth mode) node, kernel 7.0.0-31, KVM, QEMU 10.2.1, two RuntimeClass fluxvm Pods in direct mode became Ready, reached each other, left no fvbh* bridge on the node, and a deny-all NetworkPolicy still blocked the client (evidence). A real QEMU also accepted both the launch-time tap descriptor and the warm-pool getfd/netdev_add fd= sequence. A second live run with real guests (evidence) exercised: the default with the env unset (8/8 concurrent Pods direct), the auto fallback (Pods came up on the bridge chain), standalone l2-uplink with two real guests, and warm-pool claim plus direct NIC hotplug. It also measured bridge vs direct with real guests: no detectable difference (p50 0.526 vs 0.462 ms and TCP 0.34 vs 0.37 Gbit/s medians, both inside the run-to-run spread on a shared node), so the datapath's value is fewer devices and less host state, not guest-visible speed. A third pass exercised Multus secondaries (Multus put in the CNI chain for the test: a Secure Container Pod got net1, auto fell back to the bridge chain with the reason logged, direct refused as documented) and a warm-pool Pod running containers (Ready in 7 s from a claim; the Pod's shares are hot-added over virtiofs, see POST /v1/vms/{id}/hotplug/share in api.md). A claimed member also carries the Pod's eBPF identity (pod_uid on the claim), and Pod-scoped policy was enforced against real peer traffic in both directions (evidence, section F). It also fixed /dev inside Secure Containers, which was empty. Deleting a Pod takes ~20 s in direct and bridge mode alike, warm-pool or cold: the grace period, then containerd's 5 s shim delete limit against the daemon's graceful guest power-off, then kubelet's retry. Pod teardown: deleting a Secure Container Pod originally hung (Terminating indefinitely) in both direct and bridge mode. That was three agent/shim bugs, not the datapath (see the CHANGELOG "Fixed" entry); with them fixed a direct-mode Pod is fully deleted in ~13 s and leaves no tap/veth links, netns aliases, QEMU or VM records or shim processes behind (a per-Pod shim leak, caused by the Shutdown handler being cancelled mid-teardown when containerd closed the connection, is fixed too). FLUXVM_CONTAINER_CNI_DATAPATH=bridge is the kill switch.
  • Multus secondary NICs use the bridge chain (auto falls back, direct fails).
  • Service Fabric is not applied to direct taps (its attaches take a bare interface name).
  • Cilium in netkit mode (or macvlan/ipvlan Pods) is not a veth → bridge chain.
  • Requires the daemon's sandbox.dataplane.mode = ebpf or cilium; legacy is rejected because the redirect is the only forwarding path. A failed dataplane attach is fatal for a direct VM.
  • l2-uplink bounds: IPv4 ARP only (no IPv6 NDP), no host↔guest traffic (same as macvtap; the host cannot reach a guest on its own uplink without an egress hook), DHCP broadcast replies and other broadcast/multicast do not reach guests, the uplink must not be a bridge/bond port, and at most 1024 MACs / IPs per uplink. Stale entries from a crashed daemon are overwritten on the next attach.
  • l2-uplink steering comes only from the operator-declared MAC and guest_ips, never learned from guest traffic, so a guest cannot poison the maps. Declaring the same MAC or IP for two VMs on one uplink is not detected: the last attach wins.
  • Live migration of a direct VM is not validated.
  • The shared per-uplink maps stay pinned under <pin_root>/uplinks/<nic>/ (a few KB) after the last VM leaves.
  • A failed create leaves a Failed VM record in the daemon (pre-existing behavior for any create failure).
  • Policy.allowed_network_modes sees a direct VM as tap.

🛠️ Operator setup​

# daemon: sandbox.dataplane.mode = "ebpf" (or "cilium"); /usr/lib/fluxvm/bpf must hold
# fluxvm_tc.bpf.o and fluxvm_direct.bpf.o (scripts/build-ebpf.sh, scripts/enable-network-fabric-ga.sh)
# TCX attach needs the helper: /usr/libexec/fluxvm/fluxvm-tcx (scripts/build-runtime-intelligence.sh);
# without it FLUXVM_TCX=auto quietly uses legacy tc
export FLUXVM_CONTAINER_CNI_DATAPATH=auto # or direct / bridge (shim, per node)

An old fluxvm_tc.bpf.o without the fluxvm_direct map cannot serve a direct VM: the attach fails, and auto falls back to the bridge chain.

Standalone uplink VM (POST /v1/vms; the uplink must be an unenslaved NIC):

{"network": {"mode": "tap", "mac": "02:00:00:00:0a:0a",
"direct": {"outer": "enp1s0", "mode": "l2-uplink", "guest_ips": ["192.168.1.50"]}}}

📡 Packet paths​

Pod (peer-veth)
in : NIC → bpf_host → redirect_peer → eth0 ingress [fluxvm_direct_in] → tap → guest
(tap egress hook: Pod-ingress policy)
out: guest → tap ingress [fluxvm_egress policy, then redirect_peer] → lxc* ingress → bpf_lxc → NIC

Standalone (l2-uplink)
in : LAN → NIC ingress [fluxvm_direct_in: ARP by target IP / MAC → tap] → tap → guest
out: guest → tap ingress [fluxvm_egress policy, then: local guest? → its tap : NIC] → LAN

bpf_redirect_peer only targets a veth/netkit peer, so the hop into the tap is a plain bpf_redirect; the guest→Pod hop uses redirect_peer, as Cilium does. "Not mine" is TC_ACT_UNSPEC, never TC_ACT_OK (which ends a TCX chain and would hide the frame from the next VM's copy).

📈 Measured forwarding cost​

Median of 6 interleaved rounds on Linux 7.0.0-31-generic x86_64 cpus=12 (load average at start 2.45 3.21 3.02), run 2026-09-19T19:18:54Z. Source: direct-datapath-interleaved-20260919T191854Z.txt / .json. A veth pair stands in for the tap in every row, so this is host forwarding cost only: no virtio, no QEMU. The + fluxvm_egress rows run the same real policy program the direct rows run, so the comparison is like for like.

TopologyDevicesPing p50 (ms)Ping p99 (ms)TCP (Gbit/s)TCP run spread64 B UDP (Mpps)
floor (one veth pair, no bridge)20.03650.06245.8±17%0.428
Pod, bridge chain (bare)80.05650.095536.6±18%0.227
Pod, bridge chain + fluxvm_egress80.0620.10737.1±18%0.222
Pod, direct40.04250.072543.4±14%0.355
Standalone, vmbr0 (bare)50.0510.086540.4±16%0.291
Standalone, vmbr0 + fluxvm_egress50.0570.09740.8±17%0.287
Standalone, direct (l2-uplink)40.0510.089543.5±18%0.35
  • Pod direct vs the bridge chain with the same policy program: -31% latency, +17% (within noise) TCP, +60% 64 B pps.
  • Standalone direct vs vmbr0 with the same policy program: -11% (within noise) latency, +6% (within noise) TCP, +22% 64 B pps.

How to read it: negative latency and positive throughput are improvements. The machine is a shared Kubernetes node: a delta is marked "within noise" when it is no larger than the half-range of the runs on either side, and those should not be read as improvements. Re-run BENCH_INTERLEAVE=1 ./scripts/bench-direct-datapath.sh on a quiet box before quoting an exact figure. These numbers say nothing about a real guest (BENCH_TARGET_IP measures one).

Takeaway. Removing the bridge chain clearly helps where the chain is long: on the Pod path (8 devices → 4) latency drops by about a third and small-packet throughput rises by about 60%, landing within roughly 15% of the no-bridge floor. Bulk TCP throughput is inside the noise on this shared machine, so no TCP claim is made. The standalone path already had a short chain (5 → 4 devices), and only its small-packet rate is clearly better. This is the cost of moving frames on the host, not of the guest, so a real VM will see a smaller fraction of it once virtio and the VMM are included.

🧯 Linux 7.x verifier note​

On 7.0.0-31 the VM-edge program was rejected (processed 1000001 insns, limit 1,000,000). The pod-policy rule scan inlined the whole rule match on each of 64 loop iterations in both address-family paths. The per-rule match is now a global BPF function, verified once (Linux ≥ 5.5), and the object uses about 17% of the limit. scripts/test-verifier-budget.sh fails above 50%.

🧪 Evidence​

./scripts/evidence-direct-datapath.sh # static + unit/integration suites
sudo FLUXVM_DIRECT_KERNEL=1 FLUXVM_BPF_DIR=dist/bpf ./scripts/evidence-direct-datapath.sh
FLUXVM_DIRECT_LIVE=1 FLUXVM_CONTAINER_CNI_DATAPATH=direct ./scripts/evidence-direct-datapath.sh # the gate for `auto`
sudo ./scripts/test-direct-datapath.sh # pod veth ⇄ tap ⇄ guest, policy, egress hook
sudo ./scripts/test-direct-uplink.sh # two VMs, one uplink (TCX and legacy tc)
sudo ./scripts/test-pod-policy-verdict.py # rule matching on the real object
sudo ./scripts/test-verifier-budget.sh # verifier complexity guard
sudo FLUXVM_BPF_DIR=dist/bpf ./scripts/bench-direct-datapath.sh