Linux VM with the in-tree KVM engine and no QEMU binaries
This profile boots a Linux guest with a direct kernel and a raw root disk.
It does not invoke qemu-system-x86_64, qemu-img, or qemu-nbd at create
time. The source image must already be raw and named .raw (or .ext4).
Use a kernel with the virtio-mmio and virtio-blk drivers and a root filesystem
matching the kernel command line (root=/dev/vda by default).
- Install the built
fluxctlandfluxvm-hypervisorbinaries and provide/dev/kvm. Setfluxvm_engine = "kvm"andfluxvm_kernel = "/var/lib/fluxvm/kernels/vmlinux"in/etc/fluxvm.toml. - Place a bootable raw image at
/var/lib/fluxvm/images/linux-rootfs.rawand the matching kernel at the path above. Keep the base image offline while cloning it. - Run
./scripts/preflight.sh --native-kvmand thensudo fluxctl --config /etc/fluxvm.toml create --spec examples/fluxvm-native-kvm.json. Checkfluxctl listand the VM's console log.
Create waits for the in-tree VMM to reach its running state and reports
initialization failures. This readiness check does not wait for the guest OS
to finish booting; inspect the console or guest agent for that.
On a Linux host with /dev/kvm, run the acceptance gate with
sudo KERNEL=/path/to/vmlinux ROOTFS=/path/to/root.raw ./scripts/test-native-kvm-no-qemu.sh.
Set FLUXVM_AGENT=1 when the image includes the enabled guest agent and
Cloud-init to check the vsock path and offline NoCloud injection too.
backend: "auto" selects the in-tree VMM only when fluxvm_engine = "kvm",
a direct kernel is available, and the image is named .raw or .ext4.
Use flux-vm explicitly for a predictable deployment.
The default fluxvm_engine is Firecracker, so set it to kvm explicitly.
The example disables the guest agent for a first boot. For agent control,
install and enable fluxvm-guest-agent in the guest image and set
"agent": {"enabled": true}. On a flat ext4 root disk, FluxVM injects its
per-VM token with debugfs instead of GuestKit's qemu-nbd. Cloud-init
requests are embedded as NoCloud files in the cloned root disk, without
cloud-localds; the guest must have Cloud-init installed. Install
e2fsprogs (debugfs) on the host for either feature. Injection supports
flat ext4 rootfs images only and fails for GPT-partitioned images. SELinux
enforcing guests need the token file labeled appropriately in the image.
qcow2 inputs require qemu-img conversion before launch.
Pause and memory snapshots work with any vCPU count: every vCPU (BSP and APs)
parks before a pause is reported complete, and the snapshot records each
vCPU's registers, MP state and (FLUXKVM1 v5) full-fidelity state: XSAVE/FPU,
XCRS, MSRs including the TSC, LAPIC, vCPU events, debug registers, TSC kHz,
kvmclock, PIC/IOAPIC and PIT. Older v1-v4 snapshots still load but restore
registers only, with a warning. Restore is intended for the same host CPU
class and the guest clock resumes from the pause instant -- restore calls
KVM_KVMCLOCK_CTRL on every vCPU (Firecracker does the same) so a
multi-vCPU guest's clocks stay consistent with each other after resume; skip
that call and the vCPUs' views of elapsed time can disagree, intermittently,
in a way a single-vCPU guest never shows (there is nothing to disagree
with). A native KVM restore now checks the snapshot before stopping a
running guest and returns a restore error rather than silently cold-booting
when the memory restore fails. Snapshot tags must be deleted before reuse.
Multi-vCPU guests see their own topology (one package, N cores, no SMT siblings) through CPUID leaves 1, 4, 0xB and 0x1F; scripts/test-kvm-topology.sh checks what the guest kernel derived. The topology is computed from the vCPUs present at boot: measuring with max_cpus headroom on 4.14 and 5.10 kernels showed no difference in the guest's view, and in-tree CPU hotplug is not yet wired end to end (no caller, no MPTABLE/MADT entries, no AP thread). See the hypervisor README for the evidence; the AMD leaves are unit-tested against the kernel source and the APM, and further validated by booting a real guest with a spoofed AuthenticAMD identity on this Intel lab host (scripts/test-kvm-topology-amd-spoof.sh) — real AMD-specific MSRs, cache/TLB leaves, errata and power management remain unverifiable without actual AMD hardware.
This is a Linux direct-kernel profile. Windows/UEFI, QEMU device parity,
cross-CPU-model snapshot portability, broad storage backends, and
a fully QEMU-free image build pipeline (partitioned cloud images still need
qemu-img once, to extract the root) are separate work.
Verified (2026-09-28) via a real fluxctl create, not just unit tests:
the auto-generated per-VM token showed up correctly at
/etc/fluxvm-guest-agent.token on the actual cloned instance disk
(debugfs -R 'cat ...' <instance>/root.raw), matching the token in the
create response — the offline debugfs injection path works end to end
through the real control plane.
Guest-agent binary compatibility gotcha, found the same way: a
fluxvm-guest-agent built on a modern host (glibc 2.39+) fails to exec in an
older guest (e.g. Ubuntu 18.04, glibc 2.27) with version 'GLIBC_2.28' not found. Build it statically instead:
rustup target add x86_64-unknown-linux-musl
scripts/build-guest-agent-static.sh # verifies the result is static
Golden template (recommended). Build the guest once and register it, so a VM is one command instead of a hand-built image and spec each time:
sudo scripts/setup-native-agent-template.sh # image + kernel + vm-template
fluxctl --config /etc/fluxvm.toml vm-template create native-agent my-vm
It installs /var/lib/fluxvm/images/linux-agent.raw (read-only base, cloned per
VM) and the 5.10 kernel, and saves examples/fluxvm-native-agent.json as the
native-agent template. The image bakes a systemd-networkd .network file
matching the virtio NIC (/etc/systemd/network/10-fluxvm.network, DHCP,
RequiredForOnline=no) and disables Cloud-init's own network rendering. Without
it, a VM with no network sat 120 s in systemd-networkd-wait-online before the
agent started (the acceptance gate's 90 s budget expired); with it the agent
answers about 20 s after create. FLUXVM_AGENT=1 scripts/test-native-kvm-no-qemu.sh
defaults to these paths.
Guest networking (verified 2026-09-29). The in-tree virtio-net device now
has a full tap<->guest datapath, not just transmit. The queue worker thread
(queue_service.rs) polls the tap fd alongside the ioeventfds; each frame is
copied into the next buffer chain the guest posted on RX queue 0
(devices/virtio_net.rs::rx_deliver), a bounds-checked descriptor walk with a
proper used-ring publish (release fence before the index bump) and interrupt
suppression when the guest sets VRING_AVAIL_F_NO_INTERRUPT. Polling
disarms itself once the guest has no free RX buffer, or while the VM is
paused, and re-arms on the next queue-0 kick or the 100 ms tick, so a slow
guest cannot spin the worker thread. vhost-net's VHOST_NET_SET_BACKEND
still runs before the guest programs its vrings and fails with EFAULT; that
failure (and a failed late program_vrings) now falls back to the userspace
pump with one log line instead of leaving the device half configured.
The host also used to address its own tap device wrong: Tap::open built the
sin_addr word with u32::from_be_bytes, which byte-swaps the octets on a
little-endian host, so 192.168.100.1/24 was actually applied as
1.100.168.192/8. It now uses native-endian conversion (the bytes are already
in network order in memory) via a small helper with a unit test
(tap::tests::ipv4_is_stored_in_network_byte_order).
scripts/test-kvm-net.sh (opt-in, root, lab; not in hosted CI — no golden
agent image on the runner) boots the golden agent guest on a tap in a
throwaway namespace, 5 runs at 1 vCPU and 5 at 2 vCPUs: DHCP lease, ping both
directions, and an 8 MiB HTTP upload/download each way with matching SHA-256.
All 10 runs passed with 0% packet loss and matching checksums; each run's
namespace, tap and processes are torn down and verified gone, and the host
nft ruleset and ip netns list are asserted unchanged at the end.
scripts/test-egress-guest.sh stage 2 (the golden guest through the egress
proxy) now runs its traffic cases instead of reporting BLOCKED; see
http-acl.md.
The rest of this section explains the pieces.
Cloud-init image and kernel (verified 2026-09-28). Firecracker's public
quickstart rootfs has no dpkg database and no Cloud-init, so it cannot verify
NoCloud. scripts/build-native-guest-image.sh builds a flat ext4 root with
Cloud-init and the static agent enabled from a Debian/Ubuntu base:
sudo scripts/build-native-guest-image.sh bionic-server-cloudimg-amd64.img \
/var/lib/fluxvm/images/linux-agent.raw
A partitioned cloud image (qcow2 or raw) has its root partition extracted into
a flat ext4 (a one-time step that uses qemu-img/losetup; the image and the
runtime profile need no QEMU). It also pins datasource_list: [NoCloud, None]
and drops /boot and /boot/efi fstab entries.
Use a kernel that seeds its entropy pool. The 4.14.174 quickstart kernel
never finished initialising it (random: fast init done only, no
crng init done), so Cloud-init's Python blocked in getrandom() and boot
stalled at "Starting Initial cloud-init job (pre-networking)"; the agent
never started. A 5.10 kernel with CONFIG_HW_RANDOM_VIRTIO=y,
CONFIG_RANDOM_TRUST_CPU=y, CONFIG_VIRTIO_MMIO=y and vsock (for example
Firecracker CI's vmlinux-5.10.225-no-acpi, under
s3.amazonaws.com/spec.ccfc.min/firecracker-ci/v1.11/x86_64/) boots it
cleanly. With that kernel and image,
sudo env KERNEL=/path/to/vmlinux-5.10.225-no-acpi ROOTFS=/path/to/linux-agent.raw \
FLUXVM_AGENT=1 ./scripts/test-native-kvm-no-qemu.sh
passed end to end through a real fluxctl create: native VMM running, the
static agent answering over vsock inside the glibc 2.27 guest, and the
Cloud-init hostname from the offline NoCloud injection applied.