Day-2 operations
Host setup, remote deploy, end-to-end verification, the Firecracker jailer, backend auto-selection, admission policy, pause/resume/exec, cgroup v2 resource control, warm VM pools, the image catalog, alternative storage backends, the distributed node-agent, and on-disk state layout. See also api.md for the REST API and PRODUCTION.md for the whole-project production checklist.
Host requirementsβ
Linux x86_64 with virtualization enabled and /dev/kvm available.
Typical packages/tools:
qemu-system-x86_64
qemu-img
cloud-localds
ip
cp
nft # netns NAT + legacy sandbox dataplane
# Optional native eBPF dataplane (sandbox.dataplane.mode = ebpf|cilium):
clang llvm libbpf-dev bpftool # build + load bpf/fluxvm_tc.bpf.c (+ optional fluxvm_xdp.bpf.c)
tc # iproute2 TC attach
Neither Cloud Hypervisor nor Firecracker is packaged by apt/dnf, so this repo ships installer
scripts that fetch the upstream release binary for your CPU architecture (x86_64 or aarch64) and
verify it against the SHA-256 digest GitHub records for that release asset before installing it.
For Firecracker, provide a compatible uncompressed guest kernel (vmlinux) and a Linux rootfs. For Cloud Hypervisor, use either direct kernel boot or firmware boot. The project's Rust Hypervisor Firmware (hypervisor-fw) is passed through the request's kernel field, matching the Cloud Hypervisor quick-start; firmware is reserved for firmware loaded through the VMM's --firmware option.
Installing (or updating) a single VMMβ
./scripts/install-cloud-hypervisor.sh # latest release, both cloud-hypervisor + hypervisor-fw
./scripts/install-cloud-hypervisor.sh v53.0 # pin a version
./scripts/install-cloud-hypervisor.sh --no-firmware
./scripts/install-firecracker.sh # latest release, firecracker + jailer
./scripts/install-firecracker.sh v1.16.1 # pin a version
Both scripts resolve the requested (or latest) GitHub release, download the arch-appropriate
binary, verify its SHA-256 digest, and install it to /usr/local/bin (override with
INSTALL_DIR=...). They are safe to re-run β an already-installed matching version is a no-op.
Deploy to a remote hostβ
Easiest (Fabric + FluxVM stack): from this repo with sibling ../fabric:
./scripts/ship sus@HOST # quick redeploy + readiness
./scripts/ship sus@HOST --full # first install
FluxVM-only remote deploy:
scripts/deploy-remote.sh does the above end-to-end over SSH: rsync the source, install system
packages + Cloud Hypervisor/Firecracker, install a Rust toolchain if needed, build, and install the
binary, config, and systemd unit.
./scripts/deploy-remote.sh 10.0.0.5 deploy --key # full deploy, SSH key auth
./scripts/deploy-remote.sh deploy@10.0.0.5 --quick # rsync + build only, skip dep install
./scripts/deploy-remote.sh 10.0.0.5 deploy --verify-only
./scripts/deploy-remote.sh --help
Testing networking and lifecycle end-to-endβ
scripts/test-networking.sh boots real VMs over each supported network mode and proves they're
actually reachable over SSH β not just that the process launched:
- QEMU user-mode NAT + host port forward (no host network changes required).
- TAP + Linux bridge + DHCP (against an existing bridge with a DHCP server on it, e.g.
libvirt's
virbr0or a bridge set up bybootstrap-host.sh). Skipped with a warning if the bridge doesn't exist. - macvtap, against a throwaway
dummy0parent by default so the test never touches a real physical NIC/switch (pass--macvtap-parent eth0to test against a real uplink instead). Since macvtap'sbridgemode can't reach the parent/host directly, the test creates a second, host-side macvtap sibling on the same parent to reach the guest's statically-assigned IP.
All three also assert cleanup: the QEMU process and (for TAP/macvtap) the interface must actually be
gone after fluxctl delete β this is what caught a TAP-interface leak during development (fixed
by making VM shutdown wait for the process to actually exit before releasing its network resources).
sudo ./scripts/test-networking.sh # bridge defaults to vmbr0, macvtap uses dummy0
sudo ./scripts/test-networking.sh --bridge virbr0 # test TAP against libvirt's default network
sudo ./scripts/test-networking.sh --macvtap-parent eth0 # test macvtap against a real uplink
sudo ./scripts/test-networking.sh --image /path/to/base.qcow2 # skip auto-downloading a test image
It downloads an Ubuntu 24.04 cloud image on first run (cached under <state_dir>/images/) unless
--image is given, generates a throwaway SSH keypair, and prints a pass/fail/warn summary.
scripts/test-lifecycle.sh covers the rest of the VM lifecycle the same way: boots a QEMU VM with
the guest agent enabled and network.mode=none, proves exec round-trips real output over vsock
(no network path exists at all), forces a CPU-bound loop into the guest so pausing has something to
verify (an idle guest's VMM process shows ~flat CPU time whether it's paused or just idle β this
avoids that false signal), confirms the VMM's own CPU-time counter actually freezes while paused,
confirms exec works again after resume, confirms stop exits the VMM process, and confirms two
concurrently-created VMs get distinct vsock CIDs. QEMU only β Cloud Hypervisor and Firecracker were
validated manually (see Pause, resume, and exec below) since they need a
Firecracker-compatible uncompressed vmlinux / extracted whole-disk rootfs respectively, more setup
than belongs in an unattended script.
sudo ./scripts/test-lifecycle.sh
sudo ./scripts/test-lifecycle.sh --image /path/to/base.qcow2
Firecracker jailer (chroot, uid/gid isolation, cgroups)β
Lab default is off. Production and untrusted multi-tenant hosts should enable
jailer (Firecracker's own rule: start via jailer only in production). Config
is host-wide (no per-VM flag) β every Firecracker VM either goes through
jailer or none do:
[jailer]
enabled = true
enforce = true # fail closed at serve + launch if enabled is false
jailer_binary = "jailer" # resolved via $PATH unless you give an absolute path
uid = 123 # must be non-root; unique per tenant for a real isolation boundary
gid = 100
chroot_base_dir = "/srv/jailer" # should be on the same filesystem as state_dir (see below)
jailer.enforce = true, or auth.require = true with a non-loopback listen,
makes jailer required: fluxctl serve and Firecracker launches refuse to
proceed until enabled = true. Loopback lab with auth.require alone does
not force jailer.
When jailer is on and the request omits kernel_args, FluxVM uses production
boot args (quiet 8250.nr_uarts=0, no console=ttyS0) so the guest cannot
unbounded-flood host stdout via the 8250 serial (Firecracker prod-host-setup).
Override with an explicit kernel_args if you need a serial console.
firecracker_binary must be an absolute path when jailer is enabled β jailer's --exec-file needs
a real path, not a bare command resolved via $PATH.
FluxVM hardlinks the kernel and rootfs into jailer's chroot (<chroot_base_dir>/<firecracker basename>/<vm-id>/root/) before invoking it β falling back to a real copy if chroot_base_dir is on
a different filesystem than the source files, which is why same-filesystem placement matters (a
multi-GB rootfs copy per VM otherwise). Every subsequent control-plane operation (pause/resume/stop,
vsock exec) is routed through the VM's actual recorded socket paths rather than a path reconstructed
from its workspace directory β necessary because jailing relocates both the Firecracker API socket and
the vsock proxy socket into the chroot, a genuinely different location than the non-jailed case.
Verified on real hardware (scripts/test-firecracker-jailer.sh): the resulting Firecracker process
really runs as the configured unprivileged uid/gid (confirmed via ps, not just "the command didn't
error"); the guest boots and answers exec over vsock through the relocated proxy socket;
pause/resume/stop all work against the relocated API socket; delete cleans up both the normal
workspace and the separate jail chroot tree, leaving no orphaned files or process.
Threat-containment mapping (jailer + virtio rate limiters + Fabric egress) β Track A for the general control plane (not disposable-only): capability-figures.md. Production merge profile: configs/production-hardening.toml.
Auto backend selectionβ
Set "backend": "auto" and the manager picks a concrete backend for you, resolved once at the very
start of create (the resolved value β never "auto" β is what's persisted and returned):
- Firecracker if the request has a
kernel, orfirecracker_kernelis set in the config β the fastest microVM start when a direct-boot kernel is available. - otherwise Cloud Hypervisor if the request has a
kernel/firmware, orcloud_hypervisor_firmwareis set in the config. - otherwise QEMU β the only one of the three that boots from just a disk image, via its own BIOS/UEFI, with no kernel or firmware required.
"backend": "flux-vm" is never chosen by auto β set it explicitly for the agent-sandbox track.
{ "name": "auto-example", "backend": "auto", "image": "/var/lib/fluxvm/images/ubuntu.qcow2", "...": "..." }
Verified on real hardware (scripts/test-auto-backend.sh): all three resolution paths actually boot
the chosen backend and answer over vsock, not just that resolve_backend returns the right enum
value in isolation.
Policy (admission limits)β
[policy] in the config file (see config.example.toml) lets an operator cap what a create
request is allowed to ask for. Every field is optional and defaults to unrestricted β an absent or
empty [policy] table behaves exactly like no policy at all:
[policy]
max_vcpus = 8
max_memory_mib = 16384
max_disk_gib = 100
max_ttl_seconds = 86400 # every request must set ttl_seconds <= this; unbounded VMs are rejected
allowed_backends = ["qemu", "firecracker"]
allowed_image_dirs = ["/var/lib/fluxvm/images"]
Checked once, right after "auto" resolves to a concrete backend and before any disk/network work
starts, so a rejected request fails fast with a specific reason (request vcpus (4) exceeds policy max_vcpus (2), policy requires ttl_seconds to be set..., backend Firecracker is not permitted by policy allowed_backends [Qemu], etc.) rather than a generic 400. allowed_image_dirs is a plain
path-prefix check β good enough to stop a tenant pointing image at an arbitrary host path, not a
symlink-resistant sandboxing boundary. Verified against a real config on real hardware: all five
cases (four rejections, one compliant create that actually boots) behave as documented.
Per-tenant aggregate quotasβ
Everything above is a per-request check β it validates the one incoming CreateVmRequest in
isolation, never looking at what else already exists. [[policy.tenants]] is different: an
aggregate cap, summed across every existing VM a tenant already owns plus the incoming request,
matched against CreateVmRequest.tenant (itself resolved authoritatively by fluxvm-api from
[[auth.tokens]]'s own tenant field or an OIDC claim before fluxvm-scheduler ever sees the
request β see SECURITY.md's "per-token / OIDC tenant is authoritative" note):
[[policy.tenants]]
tenant = "acme"
max_vcpus_total = 32
max_memory_mib_total = 131072
max_vms_total = 20
A tenant with no matching entry is unrestricted by this mechanism (still subject to every
per-request [policy] field above, unchanged). Checked right after the per-request [policy]
check, only when the request actually has a tenant set and at least one [[policy.tenants]]
entry exists at all β a full Store::list() scan is skipped entirely otherwise, so hosts that
don't use per-tenant quotas pay nothing extra per create. This is the aggregate, fleet-wide
counterpart to Kairon's own MachineQuota CRD (a sibling project in the same family) β deliberately
config-file-based here rather than a separate CRD/API object, matching how every other admission
control in this project is already shaped ([policy], [[auth.tokens]]).
REST API rate limitingβ
[policy]/[[policy.tenants]] above cap what a request may ask for; neither caps how often one
caller can ask. Before this existed, fluxvm-api had no request-volume limiting at all β a single
misbehaving script holding one valid token (or hitting an unauthenticated loopback deployment) could
issue an unbounded number of requests, no different in effect from an external DoS, with nothing in
the REST layer itself to push back. auth.rate_limit_rps/auth.rate_limit_burst (both opt-in, must
be set together) close that gap with a small hand-rolled per-caller token bucket
(fluxvm-api::rate_limit, no new dependency):
[auth]
rate_limit_rps = 20
rate_limit_burst = 40
Keyed by the same actor identity the audit log already attributes a request to β a static token's
name, an OIDC subject, an mTLS client-cert CN, or "anonymous-admin" on an unauthenticated
loopback deployment with no credentials configured at all β so a caller's bucket tracks the identity
already established by auth_middleware, not raw connection volume; two callers sharing one token
share one bucket by design, the same way the audit log already attributes them as one actor. Runs
after auth (so every request it sees already carries that identity) and before the per-tenant scope
guard, with its own explicit bypass for GET /healthz/GET /readyz (auth only skips resolving an
identity for those two, it doesn't stop them reaching the layers below it), so a liveness/readiness
probe can never be starved by a caller's own throttling. A throttled request gets
429 Too Many Requests with Retry-After set. Absent by default (both fields unset): no rate
limiting at all, byte-for-byte the behavior before this existed. The admission-side checks above stay
separate and unaffected either way β this only bounds request frequency, never what a single request
is allowed to contain.
Pause, resume, and execβ
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml pause <id>
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml resume <id>
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml exec <id> -- echo hello
exec requires agent.enabled: true in the VM spec (see the JSON contract)
and the guest image to have fluxvm-guest-agent installed and running β build it with cargo build --release -p fluxvm-guest-agent and bake it into an image via build-image's
copy_in/enable_services (see Build an image and
systemd/fluxvm-guest-agent.service).
Guest-agent auth: every agent-enabled VM gets a random shared-secret token (or the one you set in
agent.token) burned into that VM's own disk β never the shared base image β before it boots, at
/etc/fluxvm-guest-agent.token. The agent checks it on every request; eph exec/the REST /agent
route supply it automatically from the VM's own record, so callers never handle it directly. This
stops a process on the host other than fluxvm from opening a raw vsock socket to the VM's CID and
running commands as root β it does not replace REST-layer auth (see api.md),
which answers a different question ("can this caller reach fluxvm's API at all"). A VM created before
this existed, or with no token file baked into its image for another reason, still runs the agent
unauthenticated β check the agent's own startup log line to be sure. Verified on real hardware
(scripts/test-guest-agent-auth.sh): a raw, tokenless (or wrong-token) vsock request is rejected,
the correct token succeeds, and eph exec keeps working unmodified.
stop always tries a graceful VMM-level shutdown first (QMP system_powerdown for QEMU, ch-remote shutdown for Cloud Hypervisor, SendCtrlAltDel for Firecracker β x86_64 only, no ARM equivalent in
Firecracker's API today) and only force-kills the process if it doesn't exit within a grace period.
ping, copy-to, copy-from: CLI parity for the rest of the vsock agent. The REST API has had
POST /v1/vms/{id}/agent/put-file and .../get-file since the guest agent gained PutFile/GetFile,
but for a while exec was the only one of the vsock agent's operations the CLI itself exposed β
copying a file into or out of a VM meant calling the HTTP route by hand, base64-encoding the content
yourself. These three close that gap:
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml ping <id>
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml copy-to <id> ./local-file.txt /etc/app/config.yaml --mode 600
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml copy-from <id> /etc/app/config.yaml ./local-file.txt
machinectl-shaped guest access. Beside console/exec/copy-*, fluxctl also speaks the verbs
operators reach for from systemd-machined:
| machinectl | fluxctl |
|---|---|
list | list |
status / show | status <id> / show <id> (bare status = host panel) |
start / stop | start / stop |
login / shell | console (aliases login, shell; trailing args β vsock exec) |
enable / disable | enable / disable (autostart on fluxctl serve) |
poweroff / reboot | poweroff / reboot (guest agent) |
kill | kill (force VMM; prefer stop for clean) |
terminate | terminate β delete |
bind | bind β virtiofs hotplug share |
copy-to / copy-from | copy-to / copy-from |
set-limit | set-limit / resources |
list-images / image-status | list-images / image-status (show-image) |
clone / rename / remove | clone / rename / remove (catalog images) |
read-only | read-only / read-only --off |
clean | clean |
pull-raw / import-raw / export-raw | same |
pull-tar / import-tar / import-fs / export-tar | rejected β use *-raw (qcow2/raw) |
list-transfers / cancel | empty / no-op (pulls are synchronous) |
edit | rejected β use resources or recreate from spec |
console/login put the local TTY in raw mode so keystrokes reach the guest PTY intact.
ping sends a bare AgentRequest::Ping β the same health check POST /v1/vms/{id}/agent/ping performs
over the REST API β so a caller can confirm the guest agent is up (and, if a token is configured, that
this VM's own token still authenticates) without spending a real exec round trip and whatever
guest-side work that would imply just to find out. This is a distinct channel from qga ping, which
checks the separate QEMU guest-agent (virtio-serial) socket instead β a VM can have either, both, or
neither enabled.
copy-to reads the local file and rejects anything already over the guest agent's own
fluxvm_guest_protocol::MAX_FILE_TRANSFER_BYTES cap (64MB) before spending a base64 encode and a vsock
round trip on content the agent's own put_file would just reject anyway β the same cap get_file
already enforced server-side on the way out, now checked client-side on the way in too. copy-from
restores the guest-reported Unix permission bits on the local copy, not just its bytes: a key or script
copied out of a VM keeps behaving the way its mode implies instead of silently landing at this
process's umask default (--mode on copy-to is the same idea in the other direction β see put_file's
own 0o644 default when it's left unset). Both still go through the same 64MB single-message,
no-chunking transfer PutFile/GetFile already used β bulk data still belongs in a disk image, not
this channel; see MAX_FILE_TRANSFER_BYTES's own doc comment in fluxvm-guest-protocol.
migrate start/status/cancel: CLI parity for live migration. POST /v1/vms/{id}/migration/start,
GET .../migration/status, and POST .../migration/cancel (see
docs/runtime-boundary.md for the full contract) previously had
no CLI equivalent at all β triggering a migration by hand meant a raw HTTP call with a hand-built JSON
body. This closes that gap for the standalone mode docs/runtime-boundary.md already calls out: a
deployment with no Fabric orchestrator driving these routes over HTTP still needs a way to move a VM off
a node by hand.
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml migrate start <id> --destination tcp:10.0.0.9:49152
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml migrate start <id> --destination unix:/run/fluxvm/migrate.sock \
--mode post-copy --bandwidth-mbps 500 --max-downtime-ms 300 --multifd-channels 4
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml migrate status <id>
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml migrate cancel <id>
start requires the VM to already be Running and refuses anything but a tcp:/unix: destination β
the same shared allowlist the REST route validates against, exec: included, so a migration can't be
turned into an arbitrary shell invocation on the source host. Direct-datapath VMs are refused until a
two-host test exists. When the Network Fabric dataplane is attached, start quiesces the VM edge first;
the gate is resumed automatically if start returns an error, if QEMU reports Failed/Cancelled,
when migrate status later sees those terminal phases, or on migrate cancel β so a failed or abandoned
migration does not leave new flows blocked by reason migration-quiesce. --mode takes pre-copy
(default) or post-copy, matching MigrationMode's own wire spelling exactly rather than inventing a
second CLI vocabulary for the same two values. status and cancel are QEMU-only β Cloud Hypervisor's
send-migration is fire-and-forget with no status-polling or cancellation primitive of its own (see
runtime-boundary.md's statusPollable capability field) β calling either against a Cloud Hypervisor VM
returns a clear error rather than hanging. None of this picks a destination host, confirms shared
storage, or does anything Fabric-shaped; it is exactly the three existing REST primitives, reachable
without Fabric or curl.
QEMU migration receivers. The target node reserves host quota and launches QEMU with -incoming defer
via POST /v1/migration/receivers (activate with the returned token), or the matching CLI:
fluxctl migrate receiver create β¦, fluxctl migrate receiver activate --token β¦, and
fluxctl migrate receiver delete. Receivers count against the O(1)
host ledger as untracked capacity until reaped. Default listen is 0.0.0.0 β bind to a management
address or firewall the port in multi-tenant networks. Cloud Hypervisor has no matching receiver API.
Firecracker-specific note: pause/resume were verified correct and fast against Firecracker's own
authoritative GET / state (not CPU-time heuristics β an idle guest and a paused one both show flat
CPU time, which is a false "it's paused" signal either way). exec over vsock works before a VM is
ever paused, but did not survive a pause/resume cycle in testing on this Firecracker version β a
Cloud Hypervisor VM's vsock connection did survive the identical pause/resume/exec sequence using
the same client code, so this looks like a Firecracker vsock characteristic rather than an fluxvm
bug, but it's not something this project has a fix for.
Interactive console: GET /v1/vms/{id}/console?cols=&rows= upgrades to a WebSocket relayed
end-to-end to a real PTY-backed /bin/sh in the guest over the same vsock agent connection as
exec (see fluxvm_vsock_client::open_shell) β real keystrokes, real job control, verified live
against a real QEMU VM (connect, echo a marker string, see it echoed back through the PTY).
Live resize, no reconnect. The console's initial cols/rows used to be the PTY's size for the
whole session β a client whose terminal window changed mid-session (a browser tab resized, an
xterm.js container relayout) had no way to tell the guest. Once the WS upgrade completes, a
binary WS frame is a keystroke (unchanged); a text WS frame is now read as a resize control
message, {"cols":<u16>,"rows":<u16>}, applied to the guest's PTY immediately via a real
ioctl(TIOCSWINSZ) β the shell sees an actual SIGWINCH, not just a cosmetic size report, so
full-screen programs (vim, htop, less) redraw correctly right after. A malformed text frame
(not valid JSON in that shape) is dropped, not treated as an error β it costs the client one missed
resize, not the console session. On the wire between fluxvm-api and the guest agent this rides as
a new fluxvm_guest_protocol::PtyFrame: every byte written to the guest after ShellOpened is now
PtyFrame::Data (keystrokes) or PtyFrame::Resize, while output flowing back stays completely
unframed raw PTY bytes (see PtyFrame's own doc comment for why only one direction needed framing,
and why this uses in-band frames on the existing connection rather than a second control connection
addressed by session ID the way fluxvm-container-protocol's ResizePty does for exec sessions β
the plain guest agent's OpenShell handler has no session registry to address into by design, see
the "Fixed β process isolation" note just below). Verified: a real /bin/sh spawned on a real
PTY, resized mid-session, running stty size and getting back the new size β proves the ioctl lands
on the actual PTY the shell is attached to, not just that the resize frame parses
(fluxvm-guest-agent's open_shell_resize_frame_changes_the_ptys_window_size test). Not
verified: the full path through a real QEMU/Cloud Hypervisor vsock connection and a real browser
WebSocket client sending a text frame β that would need a live VM and a live client, neither
exercised here; the wire format and the guest-side PTY behavior are each independently verified for
real, but not the two ends wired together end to end against real hardware.
Fixed β process isolation, not a kernel-level root cause. For a while, roughly 1-in-3 console
sessions left the guest agent's vsock listener unable to accept any further connections afterward
(exec/console/file-copy calls to the same VM would then fail with a raw Connection reset by peer),
with process/thread tracing showing the listener's accept() thread permanently parked in the
kernel's vsock_accept. Extensive live isolation ruled out every userspace trigger tried β whether
and from which thread child.kill()/.wait()/.try_wait() was called on the spawned shell made no
measurable difference, and a from-scratch reproducer mirroring the real PTY/fork/setsid/relay-thread
structure could not trigger it at all across 40+ trials while the real binary kept failing β pointing
at something below userspace, in the exact AF_VSOCK/vhost_vsock accept path, that was never
pinned down to a specific kernel commit or mechanism.
The actual fix doesn't require knowing that mechanism: OpenShell sessions are no longer handled in a
thread of the guest agent's own process at all. spawn_open_shell_session() double-forks β the
grandchild does the PTY/setsid()/shell/relay work fully detached from the agent's process tree (never
sharing a process, even via a thread, with the vsock listener), while the agent's original process
only reaps the fast-exiting intermediate child and returns straight to accept(). This is exactly how
OpenSSH's sshd and systemd isolate PTY/session-leader work from their own long-lived listeners β see
their session.c/systemd-executor fork-per-session model β for the same underlying reason: signal
disposition and waitpid() are process-wide, so a session leader's lifecycle can affect an unrelated
listener sharing its process in ways a separate process boundary cannot. Verified live: 20/20 console
sessions back-to-back left exec working afterward every time (statistically conclusive against the
prior ~1-in-3 failure rate), including through the real WebSocket console path end-to-end, not just a
raw vsock handshake. zyvor-fabric's FluxVM driver can now safely request agent.enabled: true by
default β see its own docs/guides/vm-drivers/fluxvm.md.
Day-2 VM operationsβ
Every fluxctl VM argument accepts a UUID, an exact VM name, or a unique UUID
prefix (4+ hex chars). Ambiguous names/prefixes are rejected with the matching ids.
| Task | CLI | REST |
|---|---|---|
| Restart (graceful stop, then start) | fluxctl restart <vm> | POST /v1/vms/{id}/restart |
| Rename | fluxctl rename-vm <vm> <new-name> | PATCH /v1/vms/{id} {"name": "..."} |
| Set / remove labels | fluxctl label <vm> env=prod team- | PATCH /v1/vms/{id} {"labels": {"env": "prod", "team": null}} |
| List snapshots | fluxctl snapshot-list <vm> | GET /v1/vms/{id}/snapshots |
| Delete a snapshot | fluxctl snapshot-delete <vm> --tag <t> | DELETE /v1/vms/{id}/snapshots/{tag} |
| Wait for a state | fluxctl wait <vm> --for running|stopped|paused|failed|agent --timeout 120 | poll GET /v1/vms/{id} |
| Events | fluxctl events [--vm <vm>] [--event vm.] [--since RFC3339] [-f] | GET /v1/events?vm=&event=&since=&limit=, GET /v1/events/stream (SSE) |
| Token quota usage | fluxctl quota [--token T | --name N] | GET /v1/quotas/me |
| Shell completions | fluxctl completions bash|zsh|fish | β |
| Filter by labels | fluxctl list -l env=prod,!tmp | GET /v1/vms?label=env%3Dprod |
| Bulk lifecycle | fluxctl start|stop|restart -l <sel>, fluxctl delete -l <sel> --yes | per-VM calls |
| Clone a stopped VM | fluxctl clone-vm <vm> <new-name> | POST /v1/vms/{id}/clone {"name": "..."} |
| Disks | fluxctl disk list|attach|resize|detach <vm> ... | GET/POST /v1/vms/{id}/disks, PATCH/DELETE /v1/vms/{id}/disks/{name} |
| Serial console | fluxctl serial <vm> (Ctrl-] detaches) | websocket GET /v1/vms/{id}/serial |
| Backup root disk | fluxctl backup <vm> [--compress] [--dest PATH] | POST /v1/vms/{id}/backup {"compress": true} |
| Backup root + data disks | fluxctl backup <vm> --all-disks | POST /v1/vms/{id}/backup {"all_disks": true} |
| VM templates | fluxctl vm-template save|list|show|delete, fluxctl vm-template create <tpl> <vm> [--label k=v] | /v1/vm-templates[/{name}], POST /v1/vm-templates/{name}/instantiate |
| Remote contexts | fluxctl context add|use|list|current|unset|delete | β |
| Scheduled snapshots | fluxctl label <vm> fluxvm.io/snapshot-every=6h fluxvm.io/snapshot-keep=7 | same PATCH |
| API description | β | GET /v1/openapi.json (no auth) |
QEMU snapshots are qcow2-internal (qemu-img info -U lists them; delete uses HMP
delvm while running, qemu-img snapshot -d when stopped). Other backends keep
snapshots under instances/<uuid>/snapshots/<tag>/. Tags are [A-Za-z0-9._-]{1,128}.
Events are appended to state_dir/events.jsonl by the daemon and by local
fluxctl invocations, so both fluxctl events and the REST routes see every
source. Tenant-scoped tokens only see events for their own tenant's VMs.
List-style commands take a global -o json|table|wide (--output-format,
FLUXCTL_OUTPUT); JSON stays the default so scripts are unaffected:
fluxctl -o table list
fluxctl list -o wide
fluxctl -o table events --vm web -f
fluxctl console / login / shell now sizes the guest PTY from the local
terminal and forwards resizes (SIGWINCH β PtyFrame::Resize).
Label selectors are a kubectl subset: k=v, k==v, k!=v, k (present),
!k (absent), comma-joined, all terms must match. Bulk delete -l refuses to
run without --yes and lists what it would have deleted.
Clone requires a stopped VM on the default qcow2 storage. The source disk is
flattened into state_dir/images/clones/<name>-<hex>.qcow2, which becomes the
new VM's base image (removed again when the last VM using it is deleted). Data
disks and labels are copied; a Tap MAC is regenerated.
Disks (QEMU). The root disk is root (virtio). Data disks are qcow2 files
in instances/<uuid>/disks/<name>.qcow2, attached as scsi-hd on the boot-time
virtio-scsi controller; the directory is the source of truth, so a disk
hot-added with disk attach comes back on every boot. resize only grows
(block_resize live, qemu-img resize stopped); the guest still has to grow
its partition/filesystem. detach unplugs and deletes the file.
disk attach <vm> <name> --path <file-or-device> (REST: {"name", "path"}
instead of size_gib) attaches an existing qcow2/raw image or block device
(a CSI volume, an RBD map) as a symlink disks/<name>.{qcow2,raw}; block
devices boot with host_device. Files must sit under policy.allowed_image_dirs
when that is set, block devices under /dev. Detach removes only the link and
resize is refused; QEMU's image locking stops two VMs opening the same source.
NIC unplug (QEMU). hotplug nic-unplug <vm> --mac <mac>|--tap <tap> (REST:
POST /v1/vms/{id}/hotplug/nic/unplug) removes an extra NIC from a running
VM and deletes its TAP. The guest must acknowledge PCIe unplug (10 s timeout).
The primary NIC lives on the root bus and can't be removed; direct (Pod-owned)
NICs go with their sandbox. A NIC removed from the middle leaves an empty slot
that the next hotplug reuses, so the remaining NICs keep their PCIe ports.
Serial. QEMU's first serial port is a UNIX socket
(instances/<uuid>/serial.sock) that also tees into console.log, so
/logs keeps working. It needs no guest agent; one client at a time. VMs
started before this change need a restart to get the socket.
Backup writes a standalone (no backing file) qcow2. Running VMs are
captured via a temporary internal snapshot (backup-<utc>, deleted
afterwards), so the copy is crash-consistent. REST backups always land in
state_dir/backups/<name>-<utc>.qcow2; only the local CLI accepts --dest.
With --all-disks the destination is a directory holding root.qcow2 and one
<disk>.qcow2 per data disk; a live VM's disks all come from the same
temporary snapshot, so they are consistent with each other.
With the guest agent enabled (qga.enabled) and answering, a running VM's
filesystems are frozen (guest-fsfreeze-freeze) just for the snapshot and
thawed right after, so the backup is application-consistent (databases see a
clean fsync point). --quiesce auto (default) falls back to crash-consistent
when the agent doesn't answer, required fails instead, never skips it. The
result and a sidecar (<file>.json, or backup.json in a directory backup)
record quiesced, the source VM and the disks.
backups lists them, backup-delete <name> removes one, and
restore-backup <vm> <name> copies one back into a stopped VM in place:
root disk plus every data disk the backup holds (recreated if the VM no
longer has it). Disks attached from an existing image are skipped, and data
disks the backup doesn't hold are left alone. Each disk goes through a temp
file, so a failed copy leaves it untouched.
VM templates are named CreateVmRequest specs in
state_dir/vm-templates.json (separate from sandbox templates at
/v1/templates). Save from a spec file (its name may be omitted) or from an
existing VM (--from-vm, which copies the VM's spec, not its disk β use
clone-vm for that). vm-template create goes through the normal create path,
so policy, tenant and token quotas apply. A template that references a clone
base image keeps that image alive after the clone VM is deleted.
Scheduled snapshots. The daemon's reaper takes an auto-<utc> snapshot of
every running QEMU VM labelled fluxvm.io/snapshot-every (3600, 30m, 6h,
1d; minimum 60s) once the newest auto-* snapshot is older than the interval,
then prunes to fluxvm.io/snapshot-keep (default 7). Manual snapshots are never
pruned.
Remote mode. fluxctl --server http://host:7788 [--server-token T] (or
FLUXVM_URL / FLUXVM_TOKEN) drives a remote daemon over REST for: create,
vm-template, list, get, status <vm>, start, stop, restart,
delete, pause, resume, label, rename-vm, clone-vm, snapshot,
snapshot-list, snapshot-delete, backup, disk, events [-f] (SSE),
serial (websocket), quota, healthz, readyz, wait (not --for agent). VM names/prefixes resolve against the server's VM list. Other commands
exit with an error in remote mode.
Contexts save named endpoints so --server isn't needed every time:
fluxctl context add lab --server http://10.0.0.5:7788 --token "$TOKEN"
fluxctl context use lab # VM verbs now go to lab
fluxctl --context prod list # one-off
fluxctl --context local list # force local mode
fluxctl context unset # back to local by default
The file is $FLUXCTL_CONTEXTS, else $XDG_CONFIG_HOME/fluxctl/contexts.json,
else ~/.config/fluxctl/contexts.json, written mode 0600 (it holds tokens;
context list never prints them). Precedence: --server/FLUXVM_URL, then
--context/FLUXCTL_CONTEXT, then the current context, then local.
fluxctl serve always runs locally.
Resource control (cgroup v2)β
Every VM (all three backends) is migrated into its own fluxvm.slice/{id}.scope cgroup right after
launch, giving real, kernel-enforced control independent of anything a VMM's own API exposes:
curl -sS -X POST http://127.0.0.1:7788/v1/vms/<uuid>/resources \
-H 'content-type: application/json' \
-d '{"cpu_quota_percent": 150, "memory_max_bytes": 536870912, "pids_max": 64}' | jq
curl -sS -X POST http://127.0.0.1:7788/v1/vms/<uuid>/freeze # cgroup-level freeze β works even if the VMM's own API doesn't respond
curl -sS -X POST http://127.0.0.1:7788/v1/vms/<uuid>/thaw
curl -sS http://127.0.0.1:7788/v1/vms/<uuid>/frozen # {"frozen": true|false}
curl -sS http://127.0.0.1:7788/v1/vms/<uuid>/stats # CPU%, memory, disk I/O, read from the cgroup
curl -sS http://127.0.0.1:7788/v1/vms/<uuid>/pressure # PSI: cpu/memory/io some+full, avg10/60/300 + total
# CLI equivalents (freeze/thaw/frozen/resources/stats/pressure):
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml freeze <uuid>
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml frozen <uuid> # {"frozen": true|false}
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml thaw <uuid>
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml resources <uuid> \
--cpu-quota-percent 150 --memory-max-bytes 536870912 --pids-max 64
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml stats <uuid>
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml pressure <uuid>
resources (ResourcePatch) is a partial patch β only the fields you set are touched: cpu_quota_percent
(percentage of one core, e.g. 150 = 1.5 cores), memory_max_bytes, io_weight (1-10000), pids_max,
cpuset_cpus (pin to specific host cores). freeze/thaw act on the cgroup directly via
cgroup.freeze, independent of the VMM's own pause/resume API (see "Pause, resume, and exec" above) β
useful as a control path that still works if a VMM's control socket is unresponsive. Delegation
(cgroup.subtree_control) is set up once at VmManager startup; if that fails (e.g. no cgroup v2, or
insufficient privilege), resource control/metrics are unavailable for that run but VM creation/lifecycle
are otherwise unaffected β a warning is logged, not a hard failure.
Verified on real hardware (scripts/test-cgroup-resources.sh, all through the REST API against a
running fluxctl serve): a launched VM really lands in its own cgroup (confirmed by reading
cgroup.procs directly, not just trusting the recorded path); a memory limit set via resources is
really written to memory.max and reads back correctly; freeze really stops the VMM process (CPU
time frozen with a forced busy-loop running in the guest, same technique used to verify QMP-level
pause) and thaw really resumes it; stats/pressure return real nonzero, cgroup-derived numbers;
delete removes the VM's cgroup directory.
freeze/thaw/frozen: CLI parity for the cgroup-level freezer. POST /v1/vms/{id}/freeze,
POST .../thaw, and GET .../frozen have existed since cgroup v2 resource control landed, but like
resources/stats/pressure above them, triggering one meant a raw curl call β there was no CLI
form at all, unlike pause/resume, the VMM-level operation these are easy to reach for by mistake
instead of. fluxctl freeze <id> calls cgroup.freeze directly via VmManager::freeze, stopping every
process in the VM's fluxvm.slice/{id}.scope at the kernel level β this keeps working even when the
VMM's own control socket is wedged or unresponsive, which is exactly the scenario pause (QMP
stop/ch-remote pause) can't help with, since that goes through the same socket. fluxctl thaw <id>
reverses it, and fluxctl frozen <id> reports the freezer's current state as {"frozen": true|false}
without changing anything, matching the REST route exactly (both are plain GETs/POSTs with no request
body β no new wire types were needed). resources later gained CLI flags; stats/pressure now
have thin read-only CLI forms (fluxctl stats <id> / fluxctl pressure <id>) that print the same
JSON as GET /v1/vms/{id}/stats and GET /v1/vms/{id}/pressure. 7 new
CLI-argument-parsing tests covering all three commands plus a check that they aren't accidentally
aliased to each other or to pause/resume.
resources: CLI parity for the cgroup resource patch. POST /v1/vms/{id}/resources deserved
the real flag design called out above instead of a JSON blob shoved onto the command line: fluxctl resources <id> [--cpu-quota-percent N] [--memory-max-bytes N] [--io-weight N] [--pids-max N] [--cpuset-cpus SPEC] maps one flag onto each ResourcePatch field, and β matching the wire type's
own "only touch what's set" contract exactly β a field is left alone unless its flag is passed; there
is no --clear-cpu-quota or similar, because omitting a flag already means "don't touch this."
Passing none of the five is refused outright at the CLI layer (ResourcePatch's own shape gives clap
no way to express "at least one of these," so the check is a plain bail! in the match arm) rather
than silently issuing a no-op POST with an empty body. --cpuset-cpus takes the exact same range
syntax cpuset.cpus/cpuset.cpus.effective themselves use when read back (fluxvm_cgroup::cpuset's
parse_set/format_set) β "0-3", "0,2,4", "0-1,4-5" β so a value copied straight out of
cpuset.cpus round-trips, but it is deliberately its own independent parser (parse_cpuset_spec in
fluxctl), not a reuse of that one, because it enforces two things the internal parser doesn't need
to: an empty string is rejected rather than accepted as "no CPUs" (the flag is already Option<String>,
so "leave cpuset pinning untouched" is expressed by omitting --cpuset-cpus entirely, not by passing
""), and a reversed range like "5-2" is a hard error instead of silently expanding to an empty
range under plain start..=end and applying an empty cpuset β a typo that would otherwise fail open
into "pin this VM to no CPUs at all" with no error at all. 17 new tests: CLI-argument-parsing coverage
for every flag alone and all five together, the "zero flags parses fine at the clap layer but the
match arm still refuses it" split, a distinctness check against freeze/pause, and dedicated
parse_cpuset_spec coverage (ranges, comma lists, mixed, sort+dedup, whitespace, and all three
rejection cases). Verified building, cargo test -p fluxctl (50/50 passing, including the 17 new),
cargo clippy -p fluxctl --no-deps (clean against this change; the two pre-existing warnings it
reports belong to CatalogCommand's enum size and an unrelated PrivateKeyDer conversion), and
cargo fmt -p fluxctl -- --check on the Linux remote, the same way prior CLI-parity work in this
section was verified β fluxctl doesn't build on macOS (it pulls in fluxvm-network, which uses
Linux-only syscalls). Not verified against a real running VM's cgroup in this pass β set_resources
itself (the code this command calls) was already proven against real memory.max/cgroup.procs files
by scripts/test-cgroup-resources.sh when resources first landed as a REST route; this change adds
no new behavior to that path, only a CLI front end for it.
Warm VM poolsβ
A pool keeps size VMs booted from a template sitting Paused, ready to be handed out on claim in
a fraction of a full create's time instead of a full boot:
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml pool create --spec examples/pool.json
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml pool list
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml pool get my-pool
Pool spec (template is a normal CreateVmRequest β its name/ttl_seconds are ignored for pool
members, which must never expire on their own while sitting idle):
{
"name": "my-pool",
"size": 4,
"template": {
"name": "ignored",
"backend": "qemu",
"image": "/var/lib/fluxvm/images/ubuntu-agent.qcow2",
"vcpus": 2,
"memory_mib": 2048,
"network": {"mode": "none"},
"agent": {"enabled": true, "port": 17777}
}
}
Claim one through REST against a running fluxctl serve daemon β the recommended way, since a
claim's own backfill-the-pool-back-up work runs as a background task inside that long-lived process:
curl -sS -X POST http://127.0.0.1:7788/v1/pools/my-pool/claim \
-H 'content-type: application/json' \
-d '{"name": "job-123", "ttl_seconds": 900}' | jq
fluxctl pool claim <name> also exists on the CLI, but as a one-shot process it exits right
after printing the claimed VM β which can take its own backfill-replenishment task down with it
mid-flight before the process exits. fluxctl pool create avoids this by blocking until the pool is
genuinely full before its own process exits; pool claim deliberately doesn't, to keep a claim fast.
A separately-running fluxctl serve daemon's reaper independently tops up every pool on its own
schedule regardless of which process's claim under-filled it, so pool health converges either way β
but for a claim's own immediate replenishment to be reliable, use REST against a running daemon.
Resize a pool after the fact instead of deleting and recreating it from the same spec just to change one number:
curl -sS -X POST http://127.0.0.1:7788/v1/pools/my-pool/resize \
-H 'content-type: application/json' \
-d '{"size": 8}' | jq
Growing behaves exactly like pool create's own initial fill β the target size is updated
immediately and a background backfill (same code path, same reaper backstop) brings membership up
to it. Shrinking is synchronous: excess ready members are popped and deleted right away, not left
for the reaper. fluxctl pool resize <name> --size N exists on the CLI too, and β like pool create but unlike pool claim β blocks on backfill_pool_sync when growing, so this one-shot
process doesn't take its own background backfill down with it before the pool actually reaches the
requested size.
pool list/pool get (and the equivalent REST responses) report computed occupancy alongside the
stored fields, so you don't have to derive it yourself:
sudo /usr/local/bin/fluxctl --config /etc/fluxvm.toml pool get my-pool
{
"name": "my-pool",
"size": 8,
"members": ["...", "..."],
"claimed_total": 41,
"ready": 6,
"pending": 2
}
ready is members.len() under a clearer name; pending is how many more members are still
needed to reach size (0 once backfill has caught up); claimed_total is a lifetime count of
every member this pool has ever successfully handed out via pool claim β useful for telling
"this pool is sized about right" apart from "this pool has never actually been claimed from" or
"this pool is running dry constantly and should be bigger," none of which the bare size/members
pair said on their own.
Real limits today: resizing has unit and router-level tests (fluxvm-scheduler,
fluxvm-api) but, unlike the rest of this section, has not been exercised against real hardware β
scripts/test-warm-pool.sh doesn't cover it yet. A shrink that fails to delete one excess member
(logged, not fatal) leaves the pool's target size reduced but membership not fully caught up; the
reaper never shrinks a pool on its own, so an under-trimmed pool stays exactly that size until
resized again, rather than silently drifting back up. The ready/pending/claimed_total fields
are pure read-side computation over state pool create/claim/resize already maintain and
already exercise on real hardware above β nothing new to independently verify against a live guest,
but also not separately re-run against real hardware under this name.
Every pool member is verified genuinely ready β not just "a process exists" β before being paused: a
real bug found on real hardware pausing a member immediately after create() returns (before the
guest had even finished booting, let alone started its guest-agent) meant a "warm" member was actually
frozen mid-boot, so resuming it on claim still had to finish booting before exec worked at all,
defeating the point. Backfill now waits for the guest agent to answer a ping before pausing.
Verified on real hardware (scripts/test-warm-pool.sh): a pool backfills to size on its own, a REST
claim is dramatically faster than a plain create (real numbers observed: ~0.2β0.5s vs. ~4β17s), the
claimed VM works immediately (exec succeeds right away), the pool tops itself back up unasked after
each claim, two claims in a row hand out two different VMs, and pool delete cleans up every member
it still owns with no leftover VMs or processes.
Image catalog & signingβ
Reference a named, checksummed image instead of a raw path or URL β resolved transparently by
create before policy/existence checks, so allowed_image_dirs still governs the real resolved file:
{"name": "job-1", "backend": "qemu", "image": "ubuntu-24.04", "...": "..."}
Enable it with [catalog] in the config:
[catalog]
path = "/etc/fluxvm/catalog.json"
# Empty = signatures not required; non-empty = every entry MUST verify against one of these.
# Named (unlike a bare list of keys) so `signed_by` below can report real signer identity,
# not just pass/fail -- the same [[auth.tokens]]-shaped "array of tables with a name" this
# project uses everywhere else it needs a labeled list of credentials.
[[catalog.trusted_signers]]
name = "release-ci"
public_key = "BASE64_ED25519_PUBLIC_KEY"
An image reference that doesn't match any catalog entry's name is treated as a literal path/URL,
exactly like before this existed β the catalog is purely additive.
Signing is a self-contained Ed25519 scheme (not cosign/Sigstore, which need either a local cosign
binary or a live Fulcio/Rekor round trip β neither of which this project can verify end-to-end without
external network-dependent test infrastructure):
fluxvm catalog keygen
# private key (keep secret, use with `catalog sign --key`): ...
# public key -- add as [[catalog.trusted_signers]] with a name: ...
fluxvm catalog sign \
--key <private-key> --name ubuntu-24.04 \
--source https://cloud-images.ubuntu.com/releases/noble/release/ubuntu-24.04-server-cloudimg-amd64.img \
--sha256 <sha256> --distro ubuntu --version 24.04 --arch x86_64 \
--build-pipeline github-actions/build-images.yml --build-run-id 987654321 --build-commit <commit-sha> \
--catalog-file /etc/fluxvm/catalog.json # appends/updates in place; omit to just print the entry
sign also stamps a signed_at (Unix seconds, right when signing runs) onto the entry, covered by
the signature itself β the signed payload now spans name/source/sha256/format/
distro/version/arch/signed_at/build_pipeline/build_run_id/build_commit, not just the
first four. Previously distro/version/arch were present on an entry but excluded from what was
actually signed, so they could be edited in catalog.json after the fact (e.g. relabeling arch to
mislead a platform-matching consumer) without invalidating the signature β closed now. read_only
stays deliberately unsigned, since it's a mutable operational flag toggled via its own REST route
(below), not provenance data β signing it would mean every legitimate toggle silently breaks the
signature.
Build lineage (--build-pipeline/--build-run-id/--build-commit, all optional) records which
CI pipeline, which run within it, and which source commit produced the image's bytes β the real gap
this closes: previously signing only ever vouched for the catalog entry's own fields (name, source,
checksum, ...), with no way to record who built it at all. Read this claim honestly, though: it's
asserted by whoever ran catalog sign, the same posture as signed_by below β there is no
cryptographic attestation chain proving the named CI system actually produced these bytes (that would
need something like Sigstore/in-toto, which this project's signing scheme deliberately avoids for the
same reason it avoids cosign β see above). What signing does give it: once set, build_pipeline/
build_run_id/build_commit are tamper-evident exactly like distro/version/arch β relabeling
any of them in catalog.json after the fact invalidates the signature.
Breaking change for existing signed catalogs: an entry signed before this existed will fail
verification against the new, wider payload β there's no dual-format fallback. Re-run fluxvm catalog sign for every entry after upgrading if trusted_signers is configured.
With trusted_signers set, an unsigned (or wrongly-signed) catalog entry is rejected at create time
β fails closed, no silent fallback to "unsigned is fine." GET /v1/images/catalog lists every entry
with a computed signature_valid, and β new β signed_by: the name of whichever configured
trusted_signers entry's key actually verified the signature (null when unsigned, wrongly signed, or
signatures aren't required at all). This is derived fresh on every call from which key matched, not
something the entry itself claims about its own signer β an entry can't assert its own identity, only a
real key can prove it. Signing itself stays a CLI/offline operation; private keys never touch the API
surface.
Catalog CRUD over REST β add/remove/rename/clone/export entries without hand-editing
catalog.json or going through the CLI's offline sign flow (this is what zyvor-fabric's
FluxVMDriver::ImageDriver uses to replace machinectl's image-management verbs):
# Register a new entry β source can be a local path or an http(s) URL; sha256 is computed
# fresh from what actually lands on disk, not trusted from the caller.
curl -sS -X POST http://127.0.0.1:7788/v1/images/catalog \
-H 'content-type: application/json' \
-d '{"name": "ubuntu-24.04", "source": "/var/lib/fluxvm/images/ubuntu.qcow2", "format": "qcow2"}' | jq
curl -sS -X POST http://127.0.0.1:7788/v1/images/catalog/ubuntu-24.04/clone \
-d '{"target_name": "ubuntu-24.04-staging"}' | jq
curl -sS -X POST http://127.0.0.1:7788/v1/images/catalog/ubuntu-24.04-staging/rename \
-d '{"new_name": "ubuntu-24.04-qa"}' | jq
curl -sS -X POST http://127.0.0.1:7788/v1/images/catalog/ubuntu-24.04/export \
-d '{"path": "/var/lib/fluxvm/exports/ubuntu-24.04.qcow2"}' | jq
curl -sS -X DELETE http://127.0.0.1:7788/v1/images/catalog/ubuntu-24.04-qa
A clone or rename drops any existing signature and signed_at (a signature covers the entry's name,
so it no longer vouches for the new one β the old signed_at timestamp would also be misleading once
detached from a valid signature). All five mutating operations are serialized against each other and
against a fresh catalog.json read on every call β no in-memory cache to go stale.
The same CRUD lives on fluxvm catalog, too β keygen/sign were the only catalog subcommands
this project's CLI had for a while (offline-signing needs to stay CLI/offline; private keys never touch
the API surface), leaving list/add/remove/rename/clone/export/lock unlock/clean reachable only over
REST even though the underlying fluxvm_image::catalog functions backing all of it (add_entry,
remove_entry, ...) always covered them. That meant a one-shot admin task against a catalog.json no
fluxctl serve was currently serving β seeding a fresh host's catalog before the daemon is even up, or a
script that would rather shell out than depend on a running HTTP endpoint β had no CLI path at all. Each
subcommand below is a thin wrapper: it operates on catalog.path from --config/FLUXVM_CONFIG
directly (same one-shot-process model as fluxctl pool claim, see its own doc comment), so it works
identically whether or not fluxctl serve happens to be running against the same state_dir:
fluxvm catalog list
# [{"name": "ubuntu-24.04", "source": "...", "sha256": "...", "signature_valid": null, ...}, ...]
fluxvm catalog add ubuntu-24.04 --source /var/lib/fluxvm/images/ubuntu.qcow2
fluxvm catalog clone ubuntu-24.04 ubuntu-24.04-staging
fluxvm catalog rename ubuntu-24.04-staging ubuntu-24.04-qa
fluxvm catalog export ubuntu-24.04 /var/lib/fluxvm/exports/ubuntu-24.04.qcow2
fluxvm catalog lock ubuntu-24.04 # refuse remove/rename until unlocked
fluxvm catalog unlock ubuntu-24.04
fluxvm catalog remove ubuntu-24.04-qa
fluxvm catalog clean # prune downloads/ orphans no entry's source references any more
lock/unlock are new: the REST route (POST .../read-only) took a {"read_only": bool} body, which
the CLI now exposes as two verbs instead of a boolean flag β clearer on a command line than
--read-only true, and it keeps fluxvm catalog --help self-explanatory without reading the REST docs
first. Every other subcommand mirrors its REST counterpart's semantics exactly (same refuse-on-read-only,
same signature/signed_at clearing on clone/rename), since both ultimately call the identical
fluxvm_image::catalog functions.
Verified on real hardware (scripts/test-image-catalog.sh, 10/10): keygen/sign produce a real
verifiable entry; creating a VM by catalog name actually resolves and boots the underlying image; with
trusted_signers configured, an unsigned entry is rejected while a validly signed one is accepted (both
confirmed by actually trying to boot); a plain literal path still works unchanged; GET /v1/images/catalog correctly reports signature_valid: true/false for the two cases. The CRUD
endpoints above were verified live against a real deployed instance: full add β list β clone β
rename β export (byte-identical file at the destination) β delete round trip, plus the duplicate-name
and not-found error paths.
Storage backendsβ
By default a VM's disk is provisioned the same way it always has been: a
qcow2 copy-on-write overlay for QEMU, a reflinked-or-copied raw file for
Cloud Hypervisor/Firecracker. Setting storage on a create request switches
to one of three alternative provisioning backends instead β
fluxvm_core::model::StorageBackend, implemented in fluxvm_image::storage:
lvm-thinβimagemust be a/dev/<vg>/<lv>path to an existing LVM thin logical volume (in a thin pool). A fresh thin snapshot LV is created per VM (lvcreate --snapshot) and handed to the VMM directly as a raw block device β real copy-on-write at the block layer, and near-instant regardless of image size. Verified end to end on real hardware: create β a genuinely new/dev/<vg>/eph-<id>snapshot LV appears β the guest boots off it and answersexecβdeleteremoves the snapshot LV,stopalone leaves it in place (same as the disk file is left in place for every other backend). Not supported under the Firecracker jailer, since its chroot/hardlink resource-placement model doesn't extend to a shared block device β use direct (non-jailed) Firecracker, QEMU, or Cloud Hypervisor. Real bug found and fixed while testing this: LVM sets a persistent "activation skip" flag on every new thin snapshot by default; without--setactivationskip non thelvcreate, the followinglvchange -ayexits 0 but silently activates nothing, and the VM fails to boot with a "device does not exist" error. There's also a real (if narrow) udev race βlvchange -ayreturns as soon as the kernel dm target is live, before udev has necessarily finished creating the/dev/<vg>/<lv>symlink β so provisioning polls for that symlink for up to 5s rather than trusting the command's exit status alone.nbdβ QEMU only (QEMU has a nativenbd:block client; Cloud Hypervisor and Firecracker don't). The disk is the same disposable qcow2 overlay as the default backend, but it's exported over NBD via aqemu-nbdsubprocess this VM owns (over a UNIX socket, not a TCP port) instead of being opened directly as a local file β the same client/server split real remote/shared NBD storage uses, without needing a separate storage host to prove the mechanism end to end. Verified on real hardware: the exportingqemu-nbdprocess is a real, findable pid; the guest boots over the NBD attachment and answersexec;deletekills the export (stopalone leaves it running, so a laterstartcan reattach). Real bug found and fixed while testing this: injecting the guest-agent token into the disk (viaguestkit, which does its own independent qemu-nbd mount) after this VM's ownqemu-nbd --persistentexport was already running raced its write lock and failed with "Failed to get 'write' lock". Fixed by injecting the token before the export starts, not after. Separately, concurrent creates that each inject a token used to race guestkit's/dev/nbdNpick (same device handed to two mounts β see #104); guestkit β₯1.2.5 serializes allocate+connect withflock.ceph-rbdβrbd clone <pool>/<image>@fluxvm-base ...and QEMU's nativerbd:block driver (QEMU only; Cloud Hypervisor/Firecracker have no built-in Ceph client). Verified end to end against a real, live Rook Ceph cluster (the Atlas storage-control-plane project's lab: Rook v1.20.2- Ceph Squid v19.2.3,
rbd-nvme-prodpool): imported a raw image asrbd-nvme-prod/fluxvm-base, protected anfluxvm-basesnapshot on it, created a VM withstorage=ceph-rbdβrbd cloneproduced a realeph-<id>clone, QEMU booted a real guest straight offrbd:rbd-nvme-prod/eph-<id>:id=admin:conf=...all the way to a login prompt, anddeletereaped the clone (confirmed gone viarbd ls, no leak). Doesn't support automatic guest-agent token injection (guestkitneeds a local file or block device to mount, not an arbitraryrbd:URI) β that combination fails fast with a clear error rather than attempting it.
- Ceph Squid v19.2.3,
ceph-rbd-in-placeβ QEMU opens an existing<pool>/<image>as-is, with no clone and nofluxvm-basesnapshot. The image is owned by whoever created it (typically the Atlas storage control plane, via Kairon'sspec.volumes[].atlas.mode: rbd); FluxVM checks it withrbd infobefore boot and never deletes it. Because every node opens the same image, live migration of these VMs is shared-storage migration with no block copy. Credentials come only from node config ([storage] ceph_user/ceph_conf); pool and image names are limited to[A-Za-z0-9._-], so a request can't smuggle:id=/:conf=options or an@snapshotinto therbd:URI. QEMU only, and no automatic guest-agent token injection (same reason asceph-rbd).
storage defaults to unset (Default) on every create request β nothing
above changes any existing behavior unless a caller opts in.
See scripts/test-storage-backends.sh for the repeatable real-hardware
regression test covering lvm-thin and nbd (it also sets up a
loopback-backed thin pool from scratch if you don't already have one β see
the script's own --help). ceph-rbd isn't in that script β it was
verified manually against the specific external Rook Ceph lab above, which
this repo has no automated way to stand up or tear down; the recipe was:
rbd import a raw image into a pool, rbd snap create + rbd snap protect an fluxvm-base snapshot on it, then create a VM with
"storage":"ceph-rbd","image":"<pool>/<image>".
Distributed node-agentβ
fluxvm-agent is the non-Kubernetes multi-host story β a caller talks to one central endpoint
instead of knowing which host a VM is on, distinct from fluxvm-kube's per-node reconciliation
against a local fluxvm. One binary, two modes:
# Central fleet registry + create/list/delete proxy β one instance for the whole fleet.
fluxvm-agent central --listen 0.0.0.0:7799
# Per-host heartbeat client β one instance per hypervisor host, alongside a local `fluxctl serve`.
fluxvm-agent node --name worker-1 \
--central http://fleet-registry:7799 \
--fluxvm-url http://127.0.0.1:7788 \
--advertise-url http://worker-1.internal:7788 \
--label zone=us-east --label gpu=true
Every --interval-secs (default 10), each node agent reports its name, real capacity (vCPUs off
available_parallelism(), RAM off /proc/meminfo), and current VM count (via its own local
GET /v1/vms) to the central registry. POST /fleet/vms with no "node" field picks the healthy
node with the fewest VMs and proxies the create there; with an explicit "node" it targets that node
directly. GET /fleet/vms aggregates every healthy node's VMs, tagged with which node each came from,
and names any node it couldn't account for in a separate "unreachable_nodes" field (see below) rather
than silently omitting it.
GET /fleet/nodes/{name}/vms narrows that same query to exactly one node, queried directly rather
than filtered out of the aggregate. DELETE /fleet/vms/{node}/{id} proxies to that exact node.
Verified end to end across two real, physically separate hosts (scripts/test-fleet-agent.sh, 11/11
passing): both hosts register with real capacity; an unaddressed create picks the least-loaded host
and produces a real QEMU process confirmed on that exact physical host (and confirmed absent on the
other); a second create lands on the other host once the first host's load is known β real
load-aware placement, not round-robin; the fleet-wide list correctly aggregates and tags VMs from
both hosts; a fleet-proxied delete reaps the right VM on the right host and leaves the other alone.
Real bugs found and fixed while testing this across two actual hosts (bugs that are invisible
running everything on one machine, which is exactly why this got tested on two real, separate hosts
instead of just trusting the code): a node's heartbeat originally reported its own --fluxvm-url
(almost always a loopback address) straight to central β central's proxy calls for a remote node
would then silently hit whatever was listening on central's own localhost instead, with no error at
all. Fixed by splitting --fluxvm-url (what this agent uses to reach its own local fluxvm) from
--advertise-url (what a remote central should use to reach this same fluxvm β must be a real,
externally routable address). Separately, this test script's own cleanup function first tried
sudo pkill -f "target/release/fluxctl --config ..." over SSH β which matched its own command
line (the pattern string is a substring of the pkill invocation's own argv) and SIGTERMed itself
before it ever reached the real target process, leaving the actual fluxctl serve running every
time with no error surfaced. Fixed with the standard [t]arget/... bracket-escape idiom that keeps
pgrep/pkill -f from matching their own invocation.
CLI equivalents. Every /fleet/* route below also has a fluxctl fleet ...
subcommand β previously the only way to drive the central registry was a raw
curl call:
fluxctl fleet --central http://fleet-registry:7799 nodes
fluxctl fleet --central http://fleet-registry:7799 node worker-1
fluxctl fleet --central http://fleet-registry:7799 cordon worker-1
fluxctl fleet --central http://fleet-registry:7799 uncordon worker-1
fluxctl fleet --central http://fleet-registry:7799 node-vms worker-1
fluxctl fleet --central http://fleet-registry:7799 vms
fluxctl fleet --central http://fleet-registry:7799 capacity
fluxctl fleet --central http://fleet-registry:7799 create --spec vm.json [--node worker-1] [--node-selector zone=us-east]
fluxctl fleet --central http://fleet-registry:7799 delete worker-1 <vm-id>
fluxctl fleet --central http://fleet-registry:7799 deregister worker-1
--central also reads CENTRAL_URL; pass --token/FLUXVM_AGENT_TOKEN the
same way fluxvm-agent node does when the registry has a bearer token
configured. Bodies, status codes, and error text are unchanged from the raw
HTTP routes documented throughout this section β this is a thin client, not
a second implementation of any of it.
Auth / TLS / persistence / placement: set --token / FLUXVM_AGENT_TOKEN on both
central and node (Bearer on all /fleet/* except /healthz). Optional
--tls-cert/--tls-key on central. Registry persists to
--state-dir/fleet-nodes.json (survives central restart until heartbeats refresh
last_seen). Unaddressed creates use residual CPU/memory capacity scoring (not
only fewest-VMs).
Cordon a node for planned maintenance without deregistering it, waiting out
HEALTHY_WINDOW_SECS for a heartbeat timeout, or draining/evicting anything
already running on it:
curl -X POST http://fleet-registry:7799/fleet/nodes/worker-1/cordon
curl -X POST http://fleet-registry:7799/fleet/nodes/worker-1/uncordon
(add -H "Authorization: Bearer $FLUXVM_AGENT_TOKEN" when the registry has a
token configured, same as every other /fleet/* route). A cordoned node keeps
heartbeating, keeps reporting real capacity in GET /fleet/nodes (now with a
"cordoned" field alongside "healthy"), and any VM already on it keeps
running and keeps showing up in GET /fleet/vms β cordon only removes the
node from automatic placement's candidate set in a subsequent unaddressed
POST /fleet/vms, the same "stop scheduling here, don't touch what's already
there" semantics as kubectl cordon. It does not block an explicit
POST /fleet/vms {"node": "worker-1", ...} naming that exact node β matching
the same precedent: a Kubernetes Pod with spec.nodeName set bypasses the
scheduler and can still land on a cordoned node, because cordoning was never a
per-node admission check, only a placement-candidate filter. Cordoned state is
per-node-name in the registry, not carried in the node agent's heartbeat body
at all (the node agent has no idea cordoning exists), so it survives every
subsequent heartbeat and a central restart (it's in fleet-nodes.json)
without a node operator having to do anything to keep it in effect, and
without a node agent's own heartbeat ever silently clearing it back off. 9 new
unit tests in fluxvm-agent::central (cordoned-node placement exclusion, the
only-node-cordoned case, uncordon restoring eligibility, cordoning an unknown
node reporting not-found rather than silently succeeding, heartbeat
preserving an existing cordon vs. a brand-new node registering uncordoned,
and both directions of the explicit-"node"-bypasses-cordon behavior).
Check one node's own state without fetching and filtering the whole fleet:
curl http://fleet-registry:7799/fleet/nodes/worker-1
GET /fleet/nodes/{name} returns exactly the same JSON shape a
GET /fleet/nodes list entry already has β healthy, cordoned, labels,
free_vcpus/free_memory_mib, and the rest β for exactly one node, instead
of an operator or script having to pull the whole registry and pick one entry
back out client-side just to answer "is worker-1 cordoned right now?" or "did
worker-1's last heartbeat actually land?". 404 for a name that was never
registered, the same status cordon/uncordon/deregister already use for
an unknown {name} on this same path β unlike GET /fleet/nodes/{name}/vms
below, whose unknown-name case is a bad proxy target (400), not a missing
resource. fluxctl fleet node worker-1 is the CLI equivalent.
See what's actually running on a node before deciding to cordon it (or after, to confirm nothing changed):
curl http://fleet-registry:7799/fleet/nodes/worker-1/vms
This is deliberately a different route from the fleet-wide GET /fleet/vms,
not a documented query-param filter on it: answering "what's on one node"
by filtering the fleet-wide list client-side would still require every
other node in the fleet to also be up and reachable just to answer a
question about one of them. GET /fleet/nodes/{name}/vms proxies straight
to that one node's own GET /v1/vms β it works regardless of the node's
cordoned/healthy state, tags each VM with "node" the same way the
fleet-wide list does, and turns that node being unreachable or erroring
into a real 502 naming the node instead of an empty, unexplained list.
Unknown node name is a 400, same as every other /fleet/* route that
takes a {name} path segment. 5 new tests in fluxvm-agent::central
against a real (not mocked) minimal HTTP server standing in for the node β
happy path with correct tagging, unknown node, an unreachable node, a node
whose own /v1/vms itself errors, and a node with zero VMs.
Know when the fleet-wide list is incomplete. GET /fleet/vms originally
dropped any node it couldn't account for β one excluded up front for a stale
heartbeat, or one whose GET /v1/vms call failed partway through β with
nothing but a server-side tracing::warn! the caller never saw; a fleet
member being down looked identical to it simply having no VMs. The response
now carries a second field naming exactly which nodes that happened to and
why:
curl http://fleet-registry:7799/fleet/vms
# {
# "items": [ ... VMs from every node that did answer, "node"-tagged ... ],
# "unreachable_nodes": [
# { "node": "worker-2", "reason": "unhealthy (stale heartbeat)" },
# { "node": "worker-5", "reason": "unreachable: error sending request..." }
# ]
# }
"unreachable_nodes" is [] on the ordinary all-healthy path β existing
callers that only ever read "items" see no behavior change. A caller that
does care whether the list it just got is the whole fleet's now has a
direct way to check, instead of having to cross-reference GET /fleet/nodes
or grep server logs to rule out a silent gap. Each reason distinguishes a
node excluded before it was even queried (stale heartbeat) from one that
was queried and failed (connection error, a non-2xx from its own
GET /v1/vms, or an unparseable body), matching the same three failure
modes GET /fleet/nodes/{name}/vms already turns into a 502 for the
single-node case β this just extends that surface-don't-hide behavior to
the aggregate instead of requiring a per-node loop to get it. 4 new tests in
fluxvm-agent::central: all-healthy-and-reachable reports an empty
unreachable_nodes; an unreachable node is named without losing the other
node's real VMs; a stale-heartbeat node is named without ever being
contacted; a node that rejects the call with a non-2xx is named with its
own error message included in the reason.
Permanently forget a decommissioned node. Cordoning and a stale
heartbeat both take a node out of automatic placement, but neither ever
actually removes its record β retire a host for good (old hardware
decommissioned, migrated to a different fleet) and it sits in the registry,
excluded from scheduling but still cluttering GET /fleet/nodes, forever.
DELETE /fleet/nodes/{name} removes the record outright:
curl -X DELETE http://fleet-registry:7799/fleet/nodes/worker-1
It refuses with a 409 if that node's heartbeat is still fresh
(healthy(), same HEALTHY_WINDOW_SECS window as everywhere else) rather
than deleting it anyway. That's a deliberate fail-closed choice, not
laziness: apply_register treats a name with no existing record as a
brand-new node and inserts it uncordoned (see its doc comment) β so
deregistering a live node and letting its very next heartbeat re-create
the record would silently drop any cordon an operator had set on it,
undoing the exact maintenance state cordoning exists to hold in place.
Requiring staleness first gives an operator one unambiguous path to
decommission a live host: stop that host's fluxvm-agent node process (or
otherwise let its heartbeat lapse), wait out HEALTHY_WINDOW_SECS, then
deregister β never a race between "delete" and "the next heartbeat wins".
An unknown node name is a 404, same as every other /fleet/* route
keyed by {name}. Deregistering also persists immediately to
--state-dir/fleet-nodes.json, same as cordon/uncordon, so the removal
survives a central restart. 7 new tests in fluxvm-agent::central: a
stale node is removed (both at the pure-function level and through the
HTTP handler); a healthy node is refused and its record left untouched
(both levels); an unknown name reports not-found (both levels); and one
test documents the exact hazard the health check guards against β deleting
a cordoned-but-stale node's record, then feeding it a fresh heartbeat,
recreates it uncordoned.
See the fleet's total capacity without fetching GET /fleet/nodes and
summing every node's free_vcpus/free_memory_mib yourself:
curl http://fleet-registry:7799/fleet/capacity
# {
# "nodes_total": 3, "nodes_healthy": 2, "nodes_cordoned": 1, "nodes_schedulable": 1,
# "vm_count": 4,
# "vcpus_total": 16, "vcpus_used": 8, "vcpus_free": 8,
# "memory_mib_total": 32768, "memory_mib_used": 8192, "memory_mib_free": 24576
# }
Computed entirely from what each node already reported in its own last
heartbeat, sitting in the registry β unlike GET /fleet/vms, this never
proxies a single call to any node's own fluxctl serve, so it's cheap and
never degrades or blocks because one node happens to be slow or briefly
unreachable right now (a stale node is simply excluded from every total,
same as it already is from pick_best_capacity's own candidate set β not
guessed at). nodes_total/nodes_healthy/nodes_cordoned count against
the whole registry regardless of reachability; every other field counts
only healthy nodes. vcpus_total/memory_mib_total sum every healthy
node whether cordoned or not β cordoned hardware still physically exists
and still counts as real fleet capacity, it's simply not accepting new
placements right now β while vcpus_free/memory_mib_free (and
nodes_schedulable) sum only the healthy-and-uncordoned subset: exactly
the same node set an unaddressed POST /fleet/vms itself draws from via
pick_best_capacity, so a nonzero *_free here is a precise answer to
"would an unaddressed create fit right now," not an approximation that
then has to account for cordoning separately. vcpus_used/
memory_mib_used reuse capacity_score's own existing per-VM estimate
(vm_count * DEFAULT_VM_VCPUS/DEFAULT_VM_MEMORY_MIB) rather than asking
any node for each VM's real configured size β an estimate for the same
reason placement's own scoring already is one. This is the natural single
call for a capacity-planning dashboard or an autoscaler deciding whether
the fleet needs another host, instead of it re-deriving these same sums
from GET /fleet/nodes on every poll. 4 new tests in
fluxvm-agent::central: an empty fleet reports all zeros; healthy
uncordoned nodes sum correctly; a stale node contributes to none of the
totals (not just placement); and a cordoned node's hardware counts toward
*_total but is excluded from *_free/nodes_schedulable.
Constrain automatic placement by node label β e.g. keep a create off
every node except ones in a given zone, or ones that actually have a GPU β
without hand-picking an exact --node and losing failover entirely. Each
node agent reports its own operator-set labels on every heartbeat:
fluxvm-agent node --name worker-1 --central http://fleet-registry:7799 \
--label zone=us-east --label gpu=true
and an unaddressed create can require a subset of them via "nodeSelector"
(exact-match key/value, same semantics as a Kubernetes Pod's own
spec.nodeSelector):
curl -X POST http://fleet-registry:7799/fleet/vms \
-d '{"name": "gpu-job", "vcpus": 4, "memory_mib": 8192, "nodeSelector": {"gpu": "true"}}'
# or:
fluxctl fleet --central http://fleet-registry:7799 create --spec vm.json --node-selector gpu=true
pick_best_capacity_excluding only ever considers a node whose labels are a
superset of nodeSelector before scoring residual capacity, so a node with
far more free capacity but the wrong label is never picked over a smaller
one that actually matches; a nodeSelector matching no registered node
fails with a 503 naming the selector, the same "surface, don't hide"
posture the rest of this fleet API already has (no healthy, uncordoned node matches nodeSelector {gpu=true}, not a generic "no nodes"). An
explicit "node" bypasses nodeSelector entirely, exactly like it already
bypasses cordoning β the same single-attempt, no-surprise-landing semantics
resolve_target documents for cordoning apply here too. "nodeSelector" is
stripped from the body before it's forwarded to the target node's own
fluxctl serve, same as "node" already is β neither is part of
CreateVmRequest. Labels are advisory metadata only: an older node agent
that predates --label simply reports none and matches no non-empty
selector, and GET /fleet/nodes includes each node's current labels
alongside its capacity and health. 10 new tests in fluxvm-agent::central
(selector picks the matching lower-capacity node over the non-matching
higher-capacity one, an unmatched selector's error naming, an empty
selector matching everything, both directions of explicit-"node"
bypassing a non-matching selector, nodeSelector extraction/stripping
edge cases) plus 2 more in fluxctl::fleet_client exercising the
--node-selector flag end to end against a real router.
Automatic placement fails over to the next-best node when the first pick
is unreachable. A node's heartbeat only goes stale (and drops out of
placement) after HEALTHY_WINDOW_SECS (30s) with no beat β a node that
crashed, wedged, or dropped off the network moments after its last
heartbeat still looks perfectly healthy to pick_best_capacity and gets
picked anyway. Before this existed, an unaddressed POST /fleet/vms
dispatched to that single best-scoring node with no retry, so the create
failed outright even while other schedulable nodes sat idle. create_vm
now distinguishes a genuinely unreachable node (a connection failure, or
a response that isn't valid JSON) from one that reached its own fluxctl serve and explicitly rejected the request: only the former triggers
failover β the unreachable node is excluded and
pick_best_capacity_excluding tries the next-best remaining candidate,
looping until one accepts the create or every schedulable node has been
tried and found unreachable, at which point the 502 names the last node
tried. A rejection is never retried elsewhere, since every other node
would reject the identical request body identically β only connectivity
failures are worth a second attempt. An explicit "node"/--node target
is unaffected by any of this: it stays a single, non-retried attempt (the
same "bypasses the scheduler" semantics nodeSelector/cordoning already
have above), so a caller who pinned a node gets an honest failure instead
of a surprise landing somewhere else.
State layoutβ
/var/lib/fluxvm/
vms.json
vms.lock
events.jsonl (VM lifecycle/audit events; rotates to events.jsonl.1 at 16 MiB)
autostart/<uuid> (fluxctl enable markers)
network-policy/ (per-VM dataplane policy JSON)
network-groups/ (security groups, CNP store, ipcache.json)
downloads/
images/
clones/ (standalone root images written by clone-vm)
kernels/
templates/ ([sandbox].templates_dir; OCIβtemplate export)
vm-templates.json (VM templates, /v1/vm-templates)
backups/ (fluxctl backup / POST /v1/vms/{id}/backup output)
instances/
<uuid>/
root.qcow2 | root.raw
disks/<name>.qcow2 (QEMU data disks)
seed.img
user-data
meta-data
console.log
serial.sock (QEMU serial chardev; fluxctl serial / websocket)
qmp.sock | ch-api.sock | firecracker.sock | fluxvm.sock
vsock.sock (CH / Firecracker / FluxVm, when agent.enabled)
firecracker.json
snapshot/ (FluxVm memory+disk snapshots)
nbd.sock | nbd.pid (storage=nbd only β see "Storage backends" above)
storage=lvm-thin and storage=ceph-rbd disks live outside this tree entirely β a thin
snapshot LV (/dev/<vg>/eph-<id>) and an RBD clone (rbd:<pool>/eph-<id>:...) respectively,
both torn down by delete via VmRecord.lvm_lv/parsing the rbd: URI, not by deleting
anything under instances/<uuid>/.
vms.lock coordinates vms.json reads/writes across concurrent fluxvm processes (each CLI
invocation is a separate process, not just a separate task inside serve) via an OS-level flock β
without it, two VMs created at the same moment could silently lose one's record, or both get
assigned the same vsock CID.