User guide: NUMA topology, CPU set, and hugepages
Three opt-in spec.resources fields, each a direct passthrough to a real,
already-existing FluxVM (QEMU backend only) capability -- not something
Kairon invents or emulates itself.
Example
apiVersion: kairon.zyvor.dev/v1alpha1
kind: Machine
metadata:
name: db-latency-sensitive
spec:
image: {path: /var/lib/fluxvm/images/db.qcow2}
resources:
cpu: "4"
memory: 8Gi
numaNode: 0
cpuSet: "0-3"
hugepages: true
runtime: {backend: qemu}
powerState: Running
resources.numaNode— binds the guest's virtual NUMA topology to host NUMA node0(FluxVM's ownnuma_node, QEMU-numa node,...).resources.cpuSet— a-numa cpu=...-style CPU expression (FluxVM's owncpuset) describing which vCPUs belong to that virtual NUMA node -- a guest-visible topology hint, not host-level pinning (see "What this is not," below).resources.hugepages— backs the guest's memory with host hugepages (FluxVM's ownhugepages, QEMU-object memory-backend-file,...,mem-path=/dev/hugepages,prealloc=on) instead of regular anonymous memory -- real value for memory-bandwidth- sensitive workloads (DPDK, in-memory databases), at the cost of requiring the host to actually have hugepages configured and free (/proc/sys/vm/nr_hugepagesor a boot-timehugepagesz=/hugepages=kernel parameter) -- Kairon doesn't configure or reserve host hugepages for you, the same "you already needed this infrastructure" posturecsiNode'starget_core_modkernel module requirement already has.
All three also work on a MachineInstanceType (see
machine-instance-types.md) -- bundle them
into a named, reusable shape the same way cpu/memory already are.
Backend requirement
QEMU only. FluxVM only implements NUMA/cpuset/hugepages for its QEMU
backend -- Cloud Hypervisor and Firecracker Machines get nothing from
these fields today. Rather than silently ignoring them the way FluxVM
itself does, Kairon refuses the Machine outright with a clear error
(spec.resources.numaNode/cpuSet/hugepages require the qemu backend) if
any of the three are set and the Machine's spec.runtime.backend resolves
to anything other than qemu -- you get a startup-time error, not a
silently-ineffective setting.
What cpuSet is not -- and where real pinning now lives
resources.cpuSet itself is still not exclusive host-core pinning. It's
a guest-visible NUMA topology hint (what the guest's own OS/scheduler is
told about its vCPUs), not a host-level guarantee that those vCPU threads
run on specific physical cores with nothing else scheduled onto them.
For the real capability telco/NFV workloads usually mean by "dedicated CPU
placement" (KubeVirt's dedicatedCpuPlacement), see
resources.cpuPinning in machine-cpu-pinning.md
instead -- a separate, real, exclusive host-core allocator built on top of
the exact same FluxVM cgroup cpuset.cpus resize API this field used to
say was unreachable. cpuSet and cpuPinning are independent: setting
cpuPinning doesn't imply or set cpuSet (the scheduler's own core
selection isn't NUMA-topology-aware in its first cut, a named limit in
that guide), and combining both doesn't cross-validate that the pinned
cores actually fall within the requested numaNode.
Real limits today (first cut)
- QEMU-only, enforced with a clear error (see above) -- not a silent no-op for other backends.
- No host hugepage reservation/discovery -- an operator's own responsibility, same posture as other host-level prerequisites this project already documents (see SECURITY.md).
resources.cpuSetonly ever describes guest-visible topology, not real exclusive host-core pinning -- seemachine-cpu-pinning.mdfor the field that actually does that.- Creation-time-only, like
spec.resources.cpu/.memorythemselves -- editing these fields on an already-running Machine has no effect (CPU/memory hotplug is a separate, already-existing mechanism -- seemachine-hotplug.md-- and doesn't carry NUMA/ cpuset/hugepages changes either). - No live-migration compatibility checking -- unlike
spec.deviceClaims(seemachine-fencing.md's VFIO preflight section), a Machine using these fields can attempt live migration to a target with a completely different (or absent) NUMA topology; whether that's safe is between you and QEMU, not something Kairon checks.