Direct datapath with REAL guests -- second live run (2026-09-20), same node as direct-datapath-live-20260919T202359Z.txt node: single-node k3s v1.35.5+k3s1, Cilium v1.20.2 (veth, tunnel), kernel 7.0.0-31, KVM, QEMU 10.2.1, 12 cores, shared with ~54 other pods build: origin/main at 130b7ac plus the fixes in this change (pod-id/ipcache race, VM-store update, failed-record cleanup, default devices, guest-NIC wait). Guests: Ubuntu 24.04, 2 vCPU, 1 GiB, virtio-net (no vhost). Every number below was typed from the run's own output; nothing is extrapolated. 1. Default (env UNSET) is auto -> direct - Two rounds of four Pods created at the same instant, FLUXVM_CONTAINER_CNI_DATAPATH unset: 8/8 took the direct datapath (VM records: 4 direct, 0 bridge, 0 failed per round), no fvbh* bridge, pod->pod ping OK. - The first attempt of this same test (two Pods at once) sent ONE of them down the bridge chain: the daemon rejected its create with a bare "No such file or directory". Cause: two creates raced on one fixed temp file (pod-ids.json.tmp). Fixed (lock + unique temp file; regression test failed 5/5 before, 15/15 pass after). - deny-all NetworkPolicy enforced (Cilium still sees the lxc* side). - Teardown: shims=0 qemu=0 VM records=0 after each round. 2. auto fallback loop (create rejected by the daemon -> retry once on the bridge chain) - Forced by removing /usr/lib/fluxvm/bpf/fluxvm_direct.bpf.o for the run: both Pods logged "direct datapath failed (... 400 ...); falling back to the bridge chain", came up on the bridge chain (2 fvbh bridges), pod->pod ping over it OK. - Each rejected attempt used to leave a `failed` VM record behind; the shim now removes it: the VM list held only the two running VMs, then nothing after teardown (shims=0 qemu=0 VM records=0). 3. Bridge vs direct, real guests, pod->pod on the same node (3 interleaved rounds, order alternating; one bridge pair and one direct pair booted together each round) ping, 200 echoes 20 ms apart, RTT ms round1 round2 round3 bridge p50 / p99 0.553/2.000 0.274/0.403 0.526/0.777 direct p50 / p99 0.462/0.947 0.477/2.390 0.432/0.918 TCP, single stream, 1 GiB through nc (Gbit/s) bridge 0.34 0.48 0.34 direct 0.37 0.33 0.41 Medians: p50 bridge 0.526 ms vs direct 0.462 ms; TCP 0.34 vs 0.37 Gbit/s. Both differences are inside the run-to-run spread (bridge p50 alone ranged 0.274-0.553), so this run does NOT show a measurable gain: with a real guest the virtio-net/QEMU path and the guest's own userspace dominate. The veth stand-in table in docs/direct-datapath.md is the datapath-only cost; the guest-visible effect here is "none detectable". Caveats: shared, loaded node; TCP is a guest-CPU-bound nc pipe, not iperf3 (an iperf3 attempt was invalid: see 7); 3 rounds; no UDP/pps figure. 4. Standalone l2-uplink with two real guests (private veth pair as the "uplink", no bridge master, node NIC untouched) - LAN peer (netns) -> guest 1 and -> guest 2 by ARP steered from the declared guest_ips: 3/3 each. - guest 1 -> LAN peer: 3/3. guest 1 <-> guest 2 (local switching on the uplink, no bridge): 3/3 both ways. - fvbh bridges: 0; uplink has no master. 5. Warm pool - API level (real QEMU, private netns): pool member booted and paused; claim took 24 ms and the VM was running; POST /v1/vms/{id}/hotplug/nic with a direct spec -> 204; guest saw the NIC; host <-> guest ping 3/3 both ways; no bridge. - From a Pod (FLUXVM_CONTAINER_WARM_POOL=): claim works and the direct NIC is hotplugged, but the Pod never became Ready. First blocker (fixed here): the shim ran the guest network script immediately, before the guest enumerated the hotplugged NIC, and failed with an empty "exit=1"; it now waits for the NIC. Second blocker (by design, documented in secure-containers-set7r.md): "mount: /run/fluxvm/pod: wrong fs type" -- a pool member keeps the virtiofs shares of its template, so a Pod's container rootfs cannot be delivered to it. So a warm-pool Pod cannot run containers today; the datapath half of it (claim + direct hotplug) does work. 6. Multus secondaries: NOT exercised. Multus is installed on the node (cluster-network-addons) but Cilium's cni-exclusive=true had renamed its config to 00-multus.conf.cilium_bak, so it is not in the CNI chain and a Pod annotated with a NetworkAttachmentDefinition got no net1 (and, correctly for that reason, went direct). Testing it needs the node's CNI chain reconfigured, which was not done on a shared node. 7. Other things found on the way (fixed unless noted) - Concurrent creates race (1). VM store update() resurrected deleted VMs (unit-tested; not reproduced live). - Failed VM records left by fallback (2). Guest NIC not yet enumerated after hotplug (5). - /dev inside Secure Containers is the image's empty directory: /dev/null, /dev/zero and /dev/urandom do not exist, so `iperf3` ("failed to open /dev/urandom") and `dd if=/dev/zero` fail. An A/B against the pre-cgroup-fix agent behaves the same, so it is pre-existing. NOT fixed. - The OCI device policy now being enforced (workloads really run in their cgroup) needs runc's default allowed devices; added (unit-tested; moot until the nodes exist). - Bridge-chain leftovers (an fvbh*/fvh* pair plus an fvcni-* netns alias) were seen after some bridge Pods in the first, interrupted benchmark attempts; not reproduced in the final three rounds. Not investigated further. ================================================================================ ADDENDUM (same day, later): items 5-7 above were re-run after fixes. This supersedes the "NOT exercised" / "NOT fixed" statements above. ================================================================================ A. /dev inside Secure Containers -- FIXED and verified live Cause: pivot_root_into bind-mounted the container rootfs onto itself non-recursively, which dropped the /dev tmpfs the OCI mounts had placed under it, so /dev was the image's empty directory. The self-bind is now MS_BIND|MS_REC, and the minimal device nodes are chmod'ed after mknod (umask). Live: /dev/null, /dev/zero and /dev/urandom present and usable (also by a non-root user); /dev, /dev/shm, /proc, /sys, a ConfigMap volume and resolv.conf survive the pivot; iperf3 loopback works. B. Multus secondaries -- EXERCISED (temporary, reverted) Multus was put into the CNI chain by renaming 00-multus.conf.cilium_bak back (Cilium's cni-exclusive=true renames it again within seconds, so it was set false for the test and restored afterwards, with a Cilium agent restart). Bridge/host-local plugins were copied into /opt/cni/bin for the NAD and removed afterwards. - ordinary runc Pods: primary networking works through Multus -> Cilium; a Pod annotated with the NAD got net1 (10.97.0.2); sec -> plain ping 0.14 ms avg. - Secure Container Pod with the NAD, default (auto): Ready; network-status shows eth0 (cilium) and net1 (10.97.0.3); shim logged "CNI datapath for eth0: bridge chain (auto: falling back to the bridge chain: 1 Multus secondary NIC(s) are not supported with the direct datapath yet)" and "Multus secondary net1 attached as guest NIC via bridge fvbh...". VM record: direct=0, extra NICs=1. - host -> guest net1: 3/3, 0.40 ms avg; guest -> host bridge 10.97.0.1: 3/3, 0.29 ms avg. - the same Pod with DATAPATH=direct: stays ContainerCreating, FailedCreatePodSandBox "FLUXVM_CONTAINER_CNI_ DATAPATH=direct but the direct datapath cannot be used: 1 Multus secondary NIC(s) are not supported" -- refuses instead of falling back, as documented. - after teardown: no shims, VMs, fvbh bridges or fvcni aliases; CNI dir back to 00-multus.conf.cilium_bak + 05-cilium.conflist, cni-exclusive=true. C. Warm-pool Pods running containers -- FIXED and verified live Root cause was a pool member keeping its template's virtiofs shares. Now the shim hot-adds the Pod's shares after the NIC (POST /v1/vms/{id}/hotplug/share; the pool template needs shared_memory=true). - API level (real QEMU): claim a paused member, hot-add two directories -> {"tag":"fs0"}, {"tag":"fs1"}; the guest mounted both (virtiofs) and read their files, a guest write to fs0 appeared on the host; two virtiofsd processes while running, zero after delete (and no QEMU left). Bug found by this test: the second share failed because the first port attempt (already taken by fs0) consumed virtiofsd's single vhost-user connection and the retry found no socket; virtiofsd is now respawned per attempt. - From a Pod (FLUXVM_CONTAINER_WARM_POOL=, pool of 2): Pod Ready after 7 s (a cold Pod takes ~140 s), `logs` printed "started", `exec` ran, the ConfigMap volume was readable (/cfg/greeting), /dev populated; the pool refilled to 2 members after the claim. - Namespace deletion took 54 s (a cold Pod's takes ~13 s): the extra time was not investigated; teardown completed with no shims left and no leaked VMs beyond the pool's own members. D. Lab node afterwards: RuntimeClass fluxvm, shim wrapper and pools removed; no fx* namespaces, QEMU, shims, VM records or fvbh/fvh/fvcb/fvn links; k3s and fluxvm active. The fixed agent and daemon stay installed. Unrelated and pre-existing: zynera/zynera-chat ImagePullBackOff and market-llama... Pending. E. Warm-pool Pods get their eBPF Pod identity (pod_uid on claim) -- verified live, direct and bridge Before: a claimed pool member had request.pod_uid = None, so /v1/vms/{id}/network/pod-policy failed with "VM has no associated Pod identity". Now POST /v1/pools/{name}/claim takes pod_uid (the shim sends the Pod UID). - direct Pod (Ready 12 s then 7 s on the rerun): VM record request.pod_uid == the Pod's UID; the eBPF meta pod_id (2757653943) equals the pod-ids.json entry; POST .../network/pod-policy -> HTTP 200. - bridge Pod (Ready 7 s / 6 s): same three checks pass (pod_id 3129108792), so the bridge-chain hotplug now attaches the dataplane with the Pod identity too. - a claim with pod_uid "../x" -> HTTP 400 and no pool member consumed. Not exercised: enforcement of a Pod-scoped policy against peer traffic; only that the identity exists and the policy maps accept a policy. Found on the way: the bridge hotplug never recorded the VM's primary tap, so delete leaked the hn0... tap (one stray link after the first bridge run). Fixed; after the rerun no fvbh/fvh/hn links remained. Lab node afterwards: clean (no fx* ns, QEMU, shims, VM records or links; CNI dir as before). F. Follow-ups checked after E (same day) - Pod-scoped policy ENFORCEMENT on a warm-pool Pod (direct datapath; secure Pod a, runc peers b and c on the same node, 2 pings each way): no policy .............................. a->b, a->c, b->a, c->a: all 0% loss default_deny + allow only c ............ a->b blocked, b->a 100% loss; a->c and c->a 0% loss policy cleared ......................... all four 0% loss again deny b only (no default_deny) .......... a->b blocked, b->a 100% loss; a->c and c->a 0% loss drop-reasons for the VM show pod-policy-deny for a->b (2 packets after the first policy, 4 after the second). The b->a drops do not appear there; the block is observed (0% -> 100% -> 0% with the policy), but I did not attribute it to a counter. In the a->b case ping exits at once with status 1 (no summary line) rather than timing out, so the guest reports an error for the blocked send; the host edge still counted the packets. - Pod delete timing, re-measured in ONE run with the same Pod spec (terminationGracePeriodSeconds 5): cold Pod 20.5 s, warm-pool Pod 20.3 s. The earlier "54 s warm vs 13 s cold" compared different things (a namespace delete against a Pod delete) and is retracted: warm-pool Pods are not slower to delete. Where the ~20 s goes (both): 5 s grace period; ~2 s container kill; then containerd runs `shim delete` for the dead sandbox shim, which waits for the daemon's VM delete (a graceful guest power-off, 7 s for an idle guest, ~11.5 s here) and is killed at containerd's 5 s limit ("failed to delete dead shim ... signal: killed"); kubelet's next StopPodSandbox retry follows ~6 s later. Not changed. - pod-ids.json: 110 stale entries had accumulated; fixed in this commit (see CHANGELOG). Existing stale entries are not pruned. Lab node afterwards: clean.