Skip to main content

AI-agent sandbox gaps

FluxVM is a multi-VMM control plane (QEMU, Cloud Hypervisor, Firecracker, and the in-tree FluxVM hypervisor) with CLI/REST, warm pools, and Kubernetes CRDs β€” a host-local libvirt/virsh-style lifecycle layer. Disposable patterns (TTL, CoW) are optional.

The FluxVM hypervisor track (backend: "flux-vm") is the AI-agent sandbox path.

Implemented on the FluxVm track​

FeatureStatus
BackendKind::FluxVm + fluxvm-hypervisor UDS control APIYes
Real guest boot via Firecracker engine under FluxVM controlYes
fluxvm_engine = "kvm" β€” pure in-tree KVM (no Firecracker child)Yes (opt-in)
Pause / resume / shutdown (proxied to guest engine)Yes
Memory+disk snapshot (Firecracker /snapshot/create + FICLONE)Yes
Fast restore via /snapshot/load (cold-boot fallback)Yes
/v1/sandboxes + fs/process APIsYes
Guest HTTP reverse proxy (/sandbox/{id}/…, AutoResume)Yes
Multi-port proxy defaults on sandbox create (http_proxy_port(s))Yes
AutoPause + activity tracking + wake-on-requestYes
Egress allowlist + credential vault + live L7 proxyYes
HTTP method/host/path ACL in the egress proxy (egress_http_rules, deny wins, normalized paths); HTTPS via opt-in TLS interception (egress_tls_intercept, HTTP/1.1, explicit-proxy clients) β€” http-acl.mdYes
File change-set (/v1/sandboxes/{id}/baseline + /changes: added/modified/deleted) β€” sandbox-changes.mdYes
Python and Go SDKs (python/, go/, stdlib only) for /v1/sandboxesYes
GPUs for sandboxes (gpus: N on POST /v1/sandboxes): FluxVM picks free VFIO-bound GPUs under a lock, NUMA-aware, 503 on a shortage; QEMU-backed templates only β€” sandbox-gpus.mdYes (picking logic tested; not run on real GPUs)
fluxvm-procbox rootless Landlock + seccomp process sandbox with TOML profiles and a ptrace learn mode β€” procbox.mdYes
procbox sandboxes via /v1/sandboxes (procbox create kind, opt-in [sandbox.procbox]): confined /process, workspace fs, baseline/changes, and a real-discard dry-run (POST /v1/sandboxes/{id}/dry-run) β€” procbox-backend.mdYes
VM sandbox dry-run and POST /v1/vms/{id}/restore: flux-vm, QEMU and Firecracker sandboxes are snapshotted, run, diffed and restored (memory and disk both revert, snapshot stays reusable) β€” sandbox-changes.mdYes (flux-vm, QEMU, Firecracker; Cloud Hypervisor sandboxes don't exist)
Sandbox dataplane β€” Network Fabric GA (schema v4): legacy nftables (default), ebpf TC IPv4/IPv6 L3+L4 + rate limits + groups/deny/CT, cilium coexistence, CNP/identities/observe, health/ipcache/refresh-dns, policy/status/stats/flows API, optional XDP, schema/fingerprint repairYes β€” network-fabric.md, network-groups.md, network-policy.md, production-dataplane.md
OCI β†’ template exportYes
Redis shared sandbox index (FLUXVM_SANDBOX_STATE_URL)Yes
/console ops UIYes
Benchmarks β€” scripts/bench-sandbox.sh, docs/benchmarks/README.mdYes

Dataplane (summary)​

  • Status: GA β€” enable with sudo ./scripts/enable-network-fabric-ga.sh --restart or merge configs/network-fabric-ga.toml. Production profile: configs/network-fabric-prod.toml β€” production-dataplane.md.
  • Default (upgrade-safe): sandbox.dataplane.mode = "legacy" (nftables).
  • ebpf: TC program from bpf/fluxvm_tc.bpf.c; pins under /sys/fs/bpf/fluxvm; iface/schema/fingerprint meta under /run/fluxvm/ebpf; IPv4/IPv6 L3+L4 allowlists (allow_cidrs, allow_ports); Mbps/PPS limits; stats/flows/events; schema v4 maps (fluxvm_gid, fluxvm_ct, fluxvm_deny4/6); ARP/DHCP/NDP bootstrap always allowed; fallback to nftables unless required = true and a host-visible edge exists (user NAT / mode=none soft-skip; IPv6/rate never silently downgrade). Optional node XDP (bpf/fluxvm_xdp.bpf.c, meta under /run/fluxvm/xdp/) β€” disabled by default and refused in cilium mode.
  • cilium: same FluxVM edge attach after verifying /var/run/cilium/cilium.sock + bpffs; does not write Cilium private maps (coexistence, not Cilium endpoint identity).
  • REST: per-VM GET/POST …/network/policy, …/status, …/stats, …/flows, …/effective; fabric-wide /v1/network/groups, /cnp, /identities, /observe, /health, /ipcache, POST /refresh-dns (native modes; writes need admin when auth is enabled).
  • v3 GA path / schema v4: dual-stack core ABI plus groups, CNP, deny CIDRs, conntrack, FQDN resolve-at-apply, ipcache, health; NDJSON flow exporter.

Applied on FluxVm create/start/restart on the host-visible interface (guest CIDR optional for native). See network-fabric.md, network-policy.md, production-dataplane.md, ebpf-cilium.md, and the packet-decision diagrams.

Remaining (optional hardening)​

  • Production-grade virtio device live-state inside in-tree KVM snapshots (FLUXKVM1 v3 packs queue rings; Firecracker remains production snap format for FC guests)
  • Hubble SID attribution for VM traffic beyond agent CEP enrichment Done (F1)
  • Optional: scrape MICROVM_METRICS_ADDR (default 127.0.0.1:9108) from Prometheus Done (F2)
  • Ranked backlog (Sentinel Set 16 candidates, vhost VRING GPA bind, etc.): NEXT-FEATURES.md
  • SMP: 2/4/8-vCPU guests boot to userspace on the in-tree KVM engine (fixed: guest CPUID topology, KVM_SET_IDENTITY_MAP_ADDR, ioeventfd-driven virtio queues; see crates/fluxvm-hypervisor/README.md). Pause and snapshot cover all vCPUs (APs join the pause barrier).

Resolved: in-tree KVM late-boot hang (init_zbud) β€” fixed by matching Firecracker/CH TSS, boot MSRs, FPU, LAPIC lint, and serial irqfd/THRE semantics. Linux guests now mount root=/dev/vda (auto virtio_mmio.device= cmdline) through /sbin/init.

Resolved: the real control-plane boot path (guest.rs, what fluxvm.service uses for fluxvm_engine = "kvm") had its vCPU execution silently freeze forever the instant the guest logged Run /sbin/init β€” a CLI demo/smoke-test convenience in run_until() was firing for production VMs too. Masked until now because the init_zbud hang meant no VM ever reached that point. Verified end-to-end: the real fluxvm-guest-agent (vsock ping/exec/shutdown) now actually starts inside the guest through the production boot path β€” [ OK ] Started Zyvor FluxVM in-guest agent. See crates/fluxvm-hypervisor/README.md.

Shipped: concurrent density (scripts/bench-density.sh), Cilium-agent CEP identity enrich (no private maps), in-tree KVM FLUXKVM1 v2 memory snapshots (all vCPUs), MicroVM schedule→Running histograms.

Resolved: the guest HTTP reverse proxy hung indefinitely for a tap+netns=true sandbox. guest_ip there is only routable from inside that VM's own network namespace β€” the daemon's default namespace has no path to it, so the proxy's plain reqwest::Client (which always connects from the calling thread's ambient namespace) silently hung until timeout. vsock-based guest-agent calls (exec, fs read/write) were unaffected, since vsock doesn't route through the guest's netns at all β€” a fully working guest still looked completely unreachable over this one path. Fix: a new connect_in_netns() setns()s a throwaway std::thread (never tokio::task::spawn_blocking, whose pool reuses threads and would leak the namespace change into unrelated work) before calling connect(), then hands the connected socket back to the async runtime β€” a namespace only governs which sockets a thread's syscalls create, not fds it already holds. Since reqwest has no way to accept a pre-connected socket, the proxy now speaks HTTP/1.1 directly over that stream via hyper instead.

Resolved: two more ways the same guest HTTP proxy could hang or return a broken response, found immediately after the netns fix above while auditing the same code path. First, every incoming header except Host was forwarded verbatim to the guest β€” but hyper's own body wrapper computes and sets its own Content-Length, so forwarding the original value duplicated the header; on any non-empty body, the resulting framing conflict left the guest's HTTP server waiting on body bytes that were never coming, hanging the request indefinitely instead of erroring (Transfer-Encoding has the same problem for a chunked-transfer client) β€” both are now stripped before forwarding, alongside Host. Second, the connect and request-send steps were already timeout-bounded, but the subsequent guest-response body read had no timeout at all β€” a guest that accepted the request but never finished (or never sent) its response body hung the whole proxy call forever; it now shares the same 30s tokio::time::timeout pattern already used for connect/send. The proxy also no longer forwards the guest response's Connection/keep-alive headers onto the caller: those describe the handler's own short-lived, one-shot connection to the guest (torn down right after this response regardless of what it says), not the caller's connection to fluxvm-api's own server β€” forwarding them could make the caller's connection pool try to reuse a connection axum's own server settings had already closed, surfacing as a bare "connection closed before message completed" with no indication why.

Resolved: a Firecracker guest kernel with no fast entropy source (no RDRAND passthrough, no other virtio-rng) boots with crng_init=0 β€” anything that calls getrandom() before enough environmental noise accumulates (in practice, the guest's own init/runtime startup) blocks forever, indistinguishable from a hung process from the outside: alive, no output, nothing listening. firecracker_config() now unconditionally requests Firecracker's built-in entropy device on every boot, regardless of what other boot options (tap, vsock) are set.

Host config​

# config.toml
fluxvm_engine = "firecracker" # default
# fluxvm_engine = "kvm" # no Firecracker child β€” in-tree KVM thread

[sandbox]
http_proxy_default_port = 8080

# [sandbox.dataplane]
# mode = "ebpf" # GA: enable-network-fabric-ga.sh
# bpf_object = "/usr/lib/fluxvm/bpf/fluxvm_tc.bpf.o"
# pin_root = "/sys/fs/bpf/fluxvm"
# required = true
# default_allow = false
# allow_cidrs = ["10.0.0.0/8", "172.16.0.0/12", "192.168.0.0/16"]
# allow_ports = ["tcp/443", "tcp/80", "udp/53"]
# max_egress_mbps = 250
# max_egress_pps = 100000
# sample_rate = 100 # 0 = off
# [sandbox.dataplane.xdp] # leave disabled with mode = "cilium"
# enabled = false
# block_cidrs = ["198.51.100.0/24", "2001:db8:bad::/48"]

Where FluxVM is ahead or different​

  • Three+ VMM backends and richer storage (LVM thin, NBD, Ceph RBD)
  • virtiofs, macvtap, per-VM netns, image catalog Ed25519 signing
  • Native TC/eBPF + safe Cilium coexistence without CNI lock-in
  • Suite fit: GuestKit β†’ FluxVM β†’ h2kvm / Ragnarok

CI​

  • agent-sandbox.yml (portable, PR + main): fmt for the touched crates, egress ACL + TLS interception, scheduler/API (change-set, procbox kind), procbox with an explicit Landlock check, the Python and Go SDKs, and both SDKs against a real fluxctl serve with procbox enabled.
  • native-kvm.yml: real guest boots on hosted runners (SMP, pause, snapshot; /dev/kvm required unless repo var FLUXVM_CI_KVM_OPTIONAL=1), the static musl guest agent, and, on the self-hosted lab runner (repo var FLUXVM_NATIVE_KVM_LIVE_CI=1, dispatch/weekly), the golden-template gate through fluxctl create.
  • all-features.yml runs the portable jobs on every push to main.