Service Fabric Phase 6
Shipped on mainβ
BPF schema 4 is unchanged. FluxVM program generation 8 covers connect{4,6} + affinity parity, map-tier ELF selection, and pressure/offload status (value ABI still schema 4).
| Item | Status |
|---|---|
Identity-aware service policy (ipcache β fluxvm_spol / sid4 / sid6) | shipped (v6) |
| Default allow/deny, allow/deny identities, audit-only | shipped (v6) |
| Envoy HTTP/gRPC transparent redirect contract | shipped (v6) |
| Bypass-mark loop prevention | shipped (v6) |
| Fabric multi-node policy fan-out with snapshot/rollback | shipped (v6) |
HA mutation queue (fluxvm_haq) + userspace drain | shipped (v6) |
| Opt-in cgroup/connect4+connect6 acceleration (fail-open to TC/XDP) | shipped |
Adaptive map pressure controller (map_tier, soft/hard %, reload) | shipped |
Compile-time map tier ELFs (-DFLUXVM_MAP_TIER=S|M|L) | shipped |
RSS/offload / XDP mode validation in services/status | shipped |
Perf qualification harness scripts/test-service-fabric-perf.sh | shipped |
Perf / RSS SLO gates (SLO_* env + scripts/test-service-fabric-slo.sh) | shipped |
Universal CI Mpps / EDT / failover gates (SLO_CI=1 defaults) | shipped |
Lab Mpps / CPU ceilings (SLO_LAB=1 / test-service-fabric-lab.sh) | shipped |
RSS under-load PPS gate (scripts/test-service-fabric-rss.sh) | shipped |
| Fabric site_id / route_domain anycast + policy fencing | shipped (minimal) |
Remote ipcache write APIs (POST/DELETE /v1/network/ipcache/remote) for Fabric ClusterMesh-like directory fan-out | shipped (minimal) |
| Full mesh datapath (remote backends) via Fabric Maglev service upsert merge | shipped (lifecycle v2; Fabric-owned) |
| Geneve / VXLAN service tunnels | N/A (L3/anycast + remote backends) |
Map tier objectsβ
./scripts/build-ebpf.sh emits default M objects plus explicit tier ELFs:
fluxvm_service.bpf.o/_xdp/_connectβ default Mfluxvm_service_tier_{S,M,L}.bpf.o(and matching xdp/connect)
Capacities come from bpf/fluxvm_service_maps.bpf.h. Config
[sandbox.dataplane.service] map_tier = "S"|"M"|"L" selects the object under
/usr/lib/fluxvm/bpf/ (override with FLUXVM_SERVICE_BPF_OBJECT /
FLUXVM_SERVICE_XDP_OBJECT / FLUXVM_SERVICE_CONNECT_OBJECT). Changing tier
rewrites pin roots via a .map_tier marker (no schema bump).
SLO harnessβ
scripts/test-service-fabric-perf.sh accepts optional gates. With SLO_CI=1
(set by scripts/test-service-fabric-slo.sh), universal CI defaults apply
(overridable by env):
| Env | CI default | Check |
|---|---|---|
SLO_VIP_P99_MS | 25 when VIP= set (CI); 50 (SLO_LAB=1); else skip | VIP connect p99 β€ threshold |
SLO_PRESSURE_IDLE | 1 | pressure action must not be hard_reload |
SLO_EDT_FAIRNESS | 1 | if any service has max_egress_mbps: require edt_packets present/non-absurd; skip when EDT not configured |
SLO_FAILOVER_LOSS_MS | 500 | when FABRIC_URL+SERVICE set: HA delta fetch RTT β€ threshold; else skip |
SLO_MPPS_MIN | 0.01 (CI) / 0.05 (SLO_LAB=1) | connect-burst Mpps vs VIP; skip if no VIP |
SLO_PKT_MPPS_MIN | unset (CI) / 0.10 (SLO_LAB=1) | packet Mpps from ethtool/service stats during storm; soft-skip if no counters |
SLO_CPU_MAX_PERCENT | unset (CI) / 85 (SLO_LAB=1) | fluxvm process CPU% avg during storm β€ ceiling; soft-skip if no pid |
SLO_CI_SHAPE | 1 | status program_generation + pressure JSON shape (VIP optional) |
SLO_REQUIRE_CHANNELS | (unset) | when NS ifaces configured, status must report RSS channels |
scripts/test-service-fabric-lab.sh sets SLO_LAB=1 (higher Mpps/CPU ceilings +
RSS PPS floor 10k). CI keeps scripts/test-service-fabric-slo.sh (SLO_CI=1).
RSS under-load PPS (scripts/test-service-fabric-rss.sh, also invoked from the
SLO/lab wrappers):
| Env | Default | Check |
|---|---|---|
SLO_MPPS_MIN | 0.01 (CI) / 0.05 (SLO_LAB) | connect-burst Mpps; soft-skips on null-channel/dummy NS unless SLO_MPPS_STRICT=1 |
SLO_MPPS_STRICT | (unset) | enforce Mpps floor even on dummy/null-channel ifaces |
SLO_RSS_PPS_MIN | 1000 (CI) / 10000 (SLO_LAB) | measured iface/service PPS under short VIP storm |
SLO_RSS_STRICT | (unset) | fail when channels null on dummy ifaces; else soft-skip |
RSS_IFACE | from status | override north-south iface |
CI without VIP/API soft-skips load gates; shape gates still run when status is reachable. Lab packet/CPU gates soft-skip when ethtool/stats/pid are unavailable.
Harness scripts:
scripts/test-service-fabric-slo.shβ CI SLO wrapperscripts/test-service-fabric-lab.shβ lab Mpps/CPU ceilingsscripts/test-service-fabric-rss.shβ under-load PPSscripts/test-service-fabric-rss-affinity.shβ multi-queue affinity proof (pin flows)scripts/test-service-fabric-all.shβ runs the battery abovescripts/test-production-readiness.shβ operator prod gate (control + NF health + SF pins; optional VIP SLO)
Fabric WireGuard underlay e2e: fabric/scripts/test-wireguard-service-fabric.sh.
Operator doc: service-fabric.md Β· v6 notes Β· Fabric: ebpf-service-fabric.md.
Remaining candidatesβ
Still open after this tranche:
- Hardware RSS affinity on multi-queue NICs β affinity script shipped; full proofs need channelsβ₯2 + ethtool per-queue stats (soft-skips on dummy
sf-lab0). Full mesh datapath (remote backends / endpoint mesh) is owned by Fabric: catalog + reconcile merges peer Ready and active Draining backends into existing Maglev service upserts (weighted drain handoff + optional VIP match; lifecycle v2). FluxVM needs no new tunnel APIs; Geneve/VXLAN remain N/A.
Ownership remains unchanged: FluxVM owns local packet/runtime mechanics; Fabric owns distributed leases, routing, service discovery, multi-site policy and HA coordination, remote identity directory, and remote backend mesh.