Network Intelligence Guide
Guide to Gryvia's eBPF-based network intelligence: the kernel programs, the per-node collector, the
network-intelligence operator and its CRDs. Read the status box first: the programs and the operator are real code, but
several links between them are not wired yet.
:::caution Status: what works and what does not
Programs. The 30 eBPF programs in ebpf/ are CO-RE (no per-kernel builds; they need a node kernel with BTF, and the
tcx programs Linux 6.6+). All 30 compile and pass the kernel verifier on Linux 7.0 x86_64; on that host the collector
attached 36 of the 83 hooks (the kprobes and tracepoints) and decoded real TCP flows. Not verified: arm64 (compiles,
never loaded), the XDP/TCX/sockops programs (attach is config-gated: ebpf.interface, ebpf.cgroupPath), and
everything GPU-related (NCCL/CUDA uprobes, RDMA, GDS), which needs GPU or RDMA hardware. Attach results per node are at
GET :9090/api/v1/ebpf/status. See ebpf/README.md.
Collector. A privileged, hostNetwork, hostPID DaemonSet in the helm/network-intelligence chart,
disabled by default (ebpf.enabled=false). Its image (gryvia-ebpf-collector) is built in CI but is not among the release
images. Its HTTP listener on :9090 (/metrics, /healthz, /api/v1/{graph,anomalies,gpu/nccl,gpu/memory,fabric,security/alerts,ai/training,ai/pipeline,tuning/tcp,ebpf/status})
is unauthenticated, except /api/v1/flight/diagnose, which needs an HMAC token (-flight-token-file).
Operator. The ten controllers are registered and unit-tested, but most of the data path is a stub:
GryviaFlowPolicy,GryviaAutoPolicyand the auto-mitigation ofGryviaNetworkAnomalycreate CiliumNetworkPolicy objects, so those features need Cilium as the CNI. Nothing was verified on a Cilium cluster.- The Hubble gRPC and Prometheus (PromQL) queries are not implemented.
GryviaTrafficInsightdoes not measure latency or throughput,GryviaTraceSessionmanages its lifecycle and results ConfigMap but does not capture flows,GryviaServiceGraphdoes not discover edges, andGryviaNetworkAnomalysees empty metrics. GryviaSecurityPolicy,GryviaNetworkCost,GryviaTrainingInsightandGryviaInferenceInsightread from the collector over HTTP, but the operator uses a hard-coded base URL,http://gryvia-collector.gryvia-system.svc.cluster.local:9090, and the chart creates no Service with that name (the chart's Services are<release>-collector-metricsin the chart namespace,gryvia-networkby default). Moreover the operator asks for/api/v1/health,/api/v1/network/costsand/api/v1/inference/latency, which the collector does not serve. Until both sides are aligned these resources stay in an "awaiting data" state. Only/api/v1/security/alertsand/api/v1/ai/trainingexist on both sides.
Flow sources. Real flows can also come from Netra, the separate standalone eBPF
network observability product: set apiGateway.netra.url (an address the gateway pod can reach; a *.svc name only works
when Netra runs in the same cluster) and apiGateway.netra.tokenSecret in the gryvia chart, and
GET /api/network/flows returns Netra flows. Netra has its own license; Gryvia only calls its HTTP API. With neither
source the network graph, flows and security pages stay empty.
:::
Table of Contents
Architecture
+--------------------+ +------------------+ +-------------------+ +--------+
| eBPF Programs | --> | Collector | --> | Operator | --> | CRDs |
| (kernel-attached) | | (per-node agent) | | (control plane) | | status |
+--------------------+ +------------------+ +-------------------+ +--------+
| | | |
Kernel hooks Reads maps, folds Reconciles CRDs, User-facing
(kprobe, tracepoint, events, serves HTTP creates Cilium resources
uprobe, XDP, tcx, on :9090 policies
sockops)
eBPF programs run in the kernel on every node and capture socket, syscall, packet and GPU-library events. Their overhead has not been measured.
Collector (collector/) is a per-node DaemonSet that loads the compiled objects with cilium/ebpf, attaches them by
section name, reads the maps and ring buffers, and serves metrics and JSON on :9090.
Operator (operators/network-intelligence, chart helm/network-intelligence, not part of helm/gryvia) runs as a
Deployment and reconciles the ten CRDs below.
CRDs are the user-facing interface. Their status fields are only as good as the data path described in the status box.
eBPF Programs
Gryvia ships 30 eBPF programs (24 original programs in six categories, plus six fabric-signal programs). The collector
loads them and attaches each by section name. Hook types below come from the SEC() annotations in ebpf/*.c.
GPU Programs
| Program | Hook Type | Description |
|---|---|---|
nccl_trace | uprobe (ncclAllReduce, ncclAllGather, ncclBroadcast, ncclReduce, ncclReduceScatter, ncclSend, ncclRecv, group calls) | NCCL collective timing with per-operation histograms. |
gpu_mem_trace | uprobe (cudaMemcpy, cudaMemcpyAsync, cudaMalloc, cudaFree, cudaLaunchKernel, cudaDeviceSynchronize) | Direction-aware GPU memory transfer byte counters and allocation tracking in the CUDA runtime library. |
rdma_trace | kprobe (ib_post_send, ib_post_recv, ib_poll_cq), tracepoint (rdma/rdma_create_qp, rdma/rdma_destroy_qp) | Per-QP RDMA statistics and completion tracking. |
Fabric Signal Programs
Ten programs are listed here. Nine feed the scheduler-facing fabric signals and are observe-only (nothing is dropped or modified); quota_pace is the one exception, is off by default and is described last. Those that use a ring buffer
emit struct fabric_signal on their own fabric_events ring buffer instead of extending the frozen 72-byte gpu_event.
The collector folds the signals per job over a 5-minute window and serves them, with a [0,1] score penalty, at
GET :9090/api/v1/fabric and as gryvia_fabric_* Prometheus gauges. The GryviaFabricSignal CRD (short name gfs)
carries the same fields in its status. With -publish-fabric-status (off by default) the collector patches the status of an existing GryviaFabricSignal whose spec.jobRef names the job; nothing else fills it, and the score is not wired into the scheduler.
| Program | Hook Type | Description |
|---|---|---|
straggler | uprobe (ncclAllReduce) | Flags a rank whose collective takes more than 2x the fastest recent span of the same payload size (and more than 5 ms) on the same node. |
rdma_health | kprobe (ib_post_send, mlx5_ib_post_send, ib_poll_cq, mlx5_ib_poll_cq) | Counts sends per QP and retry-exceeded / RNR-exceeded work completions; emits a signal above thresholds set in the rdma_thresh map. Every probe is optional. |
gds_trace | uprobe (cuFileRead, cuFileWrite), kprobe (nvidia_fs_read) | Counts bytes and times cuFile calls; a call is "direct" only when the nvidia-fs kernel hook is also seen, otherwise it counts as bounce/unknown. Only calls slower than 2 ms are emitted. |
overlap | uprobe/uretprobe (ncclAllReduce, cudaDeviceSynchronize) | Reports a cudaDeviceSynchronize of at least 1 ms that runs nested inside an in-flight ncclAllReduce on the same thread (GPU idle while communicating). Folded as overlapIdleRatio, the share of the window spent in such syncs. |
roce_cnp | XDP | Counts RoCEv2 congestion notification packets (UDP 4791, BTH opcode 0x81) in per-CPU counters; always returns XDP_PASS. Attached only with -iface; the collector polls the counters and folds them as cnpRate (packets per second, job _node/cnp) and gryvia_roce_cnp_packets_total. |
infer_latency | kretprobe (inet_csk_accept), kprobe (tcp_recvmsg) | Time from accept to the first read on connections to the inference ports (-infer-ports, default 8000,8001: vLLM and Triton HTTP; skipped when the list is empty). Folded as inferWaitP99ms, informational only. |
ucx_gloo | uprobe/uretprobe (ucp_tag_send_nb, ucp_tag_send_nbx) | UCX tag-send calls that blocked for at least 5 ms (the span of the posting call, not of the transfer). Skipped when libucp.so is not found (-ucx-lib, -uprobe-pid). Gloo is not probed: its allreduce symbols are C++ mangled, vary by version and are normally linked statically into libtorch_cpu.so, so a fixed probe could never attach. Folded as ucxSlowP99ms, informational only. |
pfc_pause | XDP | Counts 802.1Qbb priority-flow-control pause frames (EtherType 0x8808, opcode 0x0101) in per-CPU counters, with a count per paused priority, plus 802.3x pause frames; always XDP_PASS. Attached only with -iface; folded as pfcRate (frames per second, job _node/pfc) and gryvia_pfc_*_total counters. Only one XDP program can own an interface, so it conflicts with roce_cnp, packet_filter and dns_tracker: the collector attaches the first one and skips the others with a logged reason instead of replacing it. Many NICs handle pause frames in the MAC and never pass them to XDP, in which case the counters stay 0. |
weight_exfil | kprobe (vfs_read, tcp_v4_connect) | Observe only. A read of at least 8 MiB from a model-weight file (by name: .safetensors, .gguf, .ckpt, .onnx, .pt, .pth, .bin, .h5), followed within 30 s by a connect from the same process to a destination outside loopback, RFC1918 and link-local ranges, produces one signal (also logged as a warning). A read alone never does. It is a heuristic (IPv4 only, read() only, name-based) and a model server that legitimately calls an external API after loading weights will trip it. Folded as exfilEvents; it never changes scoreDelta. |
quota_pace | sockops (cgroup v2) | The only program that changes anything, and it is off by default. When a process in a cgroup that has an entry in the pace_rate map opens an outbound TCP connection, the socket's SO_MAX_PACING_RATE is lowered to that rate (never raised, and never below 1 Mbit/s). No entry, no change. Attached only with -quota-pace and -cgroup-path. Entries are lease-gated (Pacer in collector/pkg/fabric): expiry deletes the entry and the collector deletes every entry it wrote when it shuts down. Sockets that are already paced keep their rate until they close; only new connections are unpaced. Alone the flag only attaches the program and paces nothing; the opt-in -quota-pace-sync reconcile of GryviaQuota spec.network.maxEgressMbps is the only thing that grants leases (see Fabric status and quota pacing). |
The fabric score penalty (scoreDelta, capped at 1) adds 0.15 when overlapIdleRatio exceeds 0.3, 0.15 when
cnpRate exceeds 100 packets per second and 0.10 when pfcRate exceeds 1000 pause frames per second, on top of the
straggler, RDMA and GDS terms. exfilEvents, ucxSlowP99ms and inferWaitP99ms are informational. The thresholds are heuristics
that have not been calibrated on real fabrics. FabricPenalty in the ai-operator scheduler package turns scoreDelta
into up to 25 points to subtract, and RescoreNodes applies a per-node map of them, but nothing calls either yet and nothing produces a per-node map.
:::caution Unverified on hardware
These programs compile and pass the kernel verifier (Linux 7.0 x86_64; arm64 compiles only). roce_cnp and
pfc_pause were run on crafted packets with BPF_PROG_TEST_RUN, overlap and ucx_gloo on stand-in libraries,
weight_exfil briefly on the live hooks of a test process, and quota_pace on a throwaway cgroup on that host; none of it has run on real UCX,
PFC or model-serving workloads or on arm64 hardware. The GPU, RDMA
and GDS paths, real RoCE traffic and infer_latency attached to a live server have not been exercised on GPU, RDMA or GPUDirect Storage hardware. On the test host none of
ib_post_send, ib_poll_cq, mlx5_ib_* or nvidia_fs_read exists, and the collector skips those hooks with a log
line (they show as skipped: symbol not found in /api/v1/ebpf/status). ib_post_send and ib_poll_cq are inline
wrappers, so kernel-side hooks only see in-kernel RDMA users; user-space verbs (NCCL) bypass them. The straggler span
is the host-side duration of the NCCL call, which is the enqueue time for asynchronous collectives. Inside pods the
uprobes need the library resolved through /proc/<pid>/root (-uprobe-pid); the per-container mount-namespace
resolver is not built yet. Signals are attributed to a job by pid through the Flight Recorder's resolver (host PID, then
cgroup pod UID, then the node's pod list and its gryvia.io/job label); a pid that cannot be resolved (not in a pod, pod
not on this node, no gryvia.io/job label, stale pod list) is never guessed and stays grouped under _unattributed by
process name, and node-level counters stay under _node/*. Neither is ever published to a GryviaFabricSignal.
Fabric status and quota pacing
Both are opt-in collector features; see docs/fabric-status.md
for the flags, RBAC and safety rules.
| Feature | Flag (chart value) | Default | Changes anything? |
|---|---|---|---|
Publish fabric status into GryviaFabricSignal.status | -publish-fabric-status (ebpf.publishFabricStatus) | off | Writes the status subresource of existing objects every 30 s; needs extra RBAC |
Quota pacing from GryviaQuota | -quota-pace + -cgroup-path + -quota-pace-sync (ebpf.quotaPace.enabled, .sync) | off | Yes: lowers the pacing rate of new outbound TCP connections of the quota's pods on the node |
| Dry run of the above | -quota-pace-dry-run (ebpf.quotaPace.dryRun) | off | No: logs what would be granted or revoked |
spec.network.maxEgressMbps (1 to 32000) on a GryviaQuota is the per-connection egress cap for pods in the quota's
namespaces (SO_MAX_PACING_RATE, not an aggregate limit, TCP only). kube-system, gryvia-system, kube-public,
kube-node-lease and the collector's own namespace are never paced.
:::
Security Programs
| Program | Hook Type | Description |
|---|---|---|
container_escape | raw_tracepoint (sys_enter), kprobe (security_file_open) | Flags suspicious syscalls (unshare, setns and similar) and sensitive file access from containers. |
crypto_detect | kprobe (tcp_v4_connect), tracepoint (sched/sched_process_exec) | Flags outbound connections to mining-pool ports and known miner executables. |
exfil_detect | kprobe (tcp_sendmsg, tcp_v4_connect) | Tracks outbound bytes per (PID, destination) over a window and flags volumes above a threshold. |
privesc_monitor | raw_tracepoint (sys_enter), kprobe (commit_creds) | Flags UID/GID transitions to root, credential changes and capability acquisition. |
driver_fim | kprobe (security_file_open), raw_tracepoint (sys_enter) | Watches write access to NVIDIA driver files and CUDA libraries, and kernel module loading. |
Performance Programs
| Program | Hook Type | Description |
|---|---|---|
tcp_tuning | tracepoint (tcp/tcp_probe) | Collects per-connection cwnd, RTT and receive-window samples for recommendations (GET /api/v1/tuning/tcp). It does not change any TCP parameter. |
connpool_analyze | kprobe / kretprobe (tcp_v4_connect), kprobe (tcp_close) | Tracks connection lifecycles to spot short-lived connections and connection storms. |
numa_path | tracepoint (net/netif_receive_skb), kprobe (__napi_poll) | Correlates the CPU handling each packet with NUMA topology to detect cross-NUMA packet processing. |
sockops_optimize | sockops, sk_msg | Detects same-node connections and redirects their traffic with bpf_msg_redirect_hash, bypassing the TCP stack. Attached only with -cgroup-path. |
Observability Programs
| Program | Hook Type | Description |
|---|---|---|
trace_correlator | tcx/ingress | Extracts W3C traceparent trace and span IDs from incoming HTTP headers. Needs Linux 6.6+ and -iface. |
latency_breakdown | kprobe / kretprobe (udp_sendmsg, tcp_v4_connect, inet_stream_connect, tls_sw_sendmsg, tcp_sendmsg, tcp_recvmsg) | Splits request latency into DNS, TCP handshake, TLS handshake and application phases. |
cost_tracker | tcx/egress, tcx/ingress | Per-pod byte counters classified as same-zone, cross-zone or external. Needs Linux 6.6+ and -iface. |
fingerprint | raw_tracepoint (sys_enter), kprobe (tcp_v4_connect) | Builds per-process behavioural feature vectors (syscall frequency, connection patterns) as an anomaly baseline. |
AI-Specific Programs
| Program | Hook Type | Description |
|---|---|---|
training_pattern | uprobe (ncclAllReduce), kprobe (tcp_sendmsg, tcp_recvmsg) | Builds a rank communication matrix and compute/communication cycle view for training jobs. |
datapipe_bottleneck | tracepoint (block/block_rq_complete), kprobe (tcp_recvmsg), uprobe (cudaLaunchKernel, cudaDeviceSynchronize) | Correlates storage I/O, network ingestion and GPU busy/idle phases. |
gradient_compress | uprobe (ncclAllReduce) | Compares expected and actual bytes per collective to estimate a compression ratio. |
Core Programs
| Program | Hook Type | Description |
|---|---|---|
tcp_trace | kprobe (tcp_v4_connect, inet_csk_accept, tcp_close, tcp_retransmit_skb) | TCP connection lifecycle: establishment latency, bytes, retransmits. This is the program behind the decoded real TCP flows. |
packet_filter | XDP | Filters against a dynamically updatable blocklist in BPF maps, with per-rule hit counters. Attached only with -iface. Nothing in the operator programs the blocklist yet. |
latency_probe | kprobe (tcp_sendmsg, tcp_recvmsg) | Per-connection latency with histogram buckets. |
syscall_monitor | raw_tracepoint (sys_enter) | Tracks which processes call connect, sendto, recvfrom; flags first-time network activity. |
dns_tracker | XDP | Passive DNS query/response correlation, resolution latency, NXDOMAIN/SERVFAIL counts. Attached only with -iface. |
CRDs
All ten kinds are gryvia.io/v1alpha1 and have a registered controller in the network-intelligence operator. The
manifests below are validated against crds/; the fields shown are the complete useful surface of each spec. See
examples/network-intelligence/ for more and reference/crds.md for the generated schema reference.
GryviaFlowPolicy
Intent-based flow policy. The controller translates it into a CiliumNetworkPolicy named ffp-<name> (needs Cilium).
intent is one of low-latency, high-throughput, secure, default.
apiVersion: gryvia.io/v1alpha1
kind: GryviaFlowPolicy
metadata:
name: payment-to-db
namespace: production
spec:
source:
service: payment-service
namespace: production
labels:
app: payment
destination:
service: postgres-primary
namespace: production
port: 5432
labels:
app: postgres
protocol: tcp
action: allow
intent: low-latency
priority: 100
GryviaTrafficInsight
Declares a service and a rolling window to analyse, with the metrics of interest (latency, throughput, drops,
retransmits). The controller requeues every window and would fill latency, throughput and top-talker status, but the
Prometheus/Hubble queries it relies on are not implemented, so no measurements are produced today.
apiVersion: gryvia.io/v1alpha1
kind: GryviaTrafficInsight
metadata:
name: payment-traffic-analysis
namespace: production
spec:
service: payment-service
namespace: production
window: "5m"
metrics:
- latency
- throughput
- drops
- retransmits
GryviaAutoPolicy
Learn / suggest / enforce state machine for generated policies. enforce creates CiliumNetworkPolicy objects from stored
suggestions, gated by approvalRequired. The learning step is a stub (no Hubble connection), so no traffic is learned
today; treat the state handling as the only real part.
apiVersion: gryvia.io/v1alpha1
kind: GryviaAutoPolicy
metadata:
name: production-auto-firewall
namespace: gryvia-system
spec:
mode: learn
learningWindow: "10m"
targetNamespaces:
- production
- staging
excludeServices:
- kube-dns
- metrics-server
approvalRequired: true
GryviaTraceSession
Time-limited debugging session. The controller creates a results ConfigMap, marks the session active and completes it
when duration expires. level is l3, l4 or l7. Flow capture from Hubble is not implemented, so the ConfigMap
holds no captured flows.
apiVersion: gryvia.io/v1alpha1
kind: GryviaTraceSession
metadata:
name: debug-payment-latency
namespace: production
spec:
service: payment-service
namespace: production
duration: "2m"
level: l7
captureHeaders: true
filters:
port: 8080
protocol: tcp
GryviaServiceGraph
Service dependency graph over a set of namespaces, refreshed every refreshInterval. Edge discovery is not implemented;
existing status edges are preserved. The gateway's GET /api/network/flows falls back to these edges when Netra is not
configured.
apiVersion: gryvia.io/v1alpha1
kind: GryviaServiceGraph
metadata:
name: production-graph
namespace: production
spec:
namespaces:
- production
- production-data
refreshInterval: "30s"
includeExternal: true
depth: 5
GryviaNetworkAnomaly
Threshold rules on a service. Metrics come from Prometheus in design; today the metric source is empty, so rules do not
fire. What is real: the webhook call and, with autoMitigate, a temporary deny CiliumNetworkPolicy for critical/high
anomalies that expires after 15 minutes.
apiVersion: gryvia.io/v1alpha1
kind: GryviaNetworkAnomaly
metadata:
name: payment-anomaly-detector
namespace: production
spec:
targetService: payment-service
detectionRules:
- metric: latency
operator: gt
threshold: 100
window: "5m"
- metric: drops
operator: gt
threshold: 100
window: "1m"
alertWebhook: "https://alerts.example.com/network-anomaly"
autoMitigate: true
GryviaSecurityPolicy
Selects which detections (type: escape, mining, exfiltration, privesc, driver_fim; sensitivity: low,
medium, high) apply to which namespaces. The controller polls the collector's /api/v1/security/alerts, counts alerts
per type into status.detectionCounts (which gryvia security alerts reads), sends the webhook, and with autoBlock
creates temporary CiliumNetworkPolicy blocks. It is subject to the collector URL mismatch described in the status box.
apiVersion: gryvia.io/v1alpha1
kind: GryviaSecurityPolicy
metadata:
name: gpu-cluster-security
namespace: gryvia-system
spec:
targetNamespaces:
- ml-research
detectionRules:
- type: mining
enabled: true
sensitivity: high
- type: escape
enabled: true
sensitivity: high
- type: exfiltration
enabled: true
sensitivity: medium
- type: privesc
enabled: true
- type: driver_fim
enabled: true
autoBlock: false
alertWebhook: "https://alerts.example.com/security"
GryviaNetworkCost
Per-namespace network cost reports from same-zone, cross-zone and external byte counters, priced with costPerGB and
attributed with costCenters. Reports are appended to status.reports every reportingInterval (default 1h). The
byte counters would come from the collector's cost_tracker program, but the operator asks for
/api/v1/network/costs, which the collector does not serve, so reports are empty today. The rates below are
placeholders, not real prices.
apiVersion: gryvia.io/v1alpha1
kind: GryviaNetworkCost
metadata:
name: monthly-network-costs
namespace: gryvia-system
spec:
targetNamespaces:
- ml-research
reportingInterval: "1h"
costPerGB:
sameZone: 0.0
crossZone: 0.01
internetEgress: 0.09
costCenters:
- namespace: ml-research
team: ml-research
costCenter: cc-1001
GryviaTrainingInsight
NCCL analysis for one GryviaAIJob. The controller reads /api/v1/ai/training?job=<targetJob> from the collector (an
endpoint both sides have), then fills status.rankStats, stragglers, commPattern, commComputeRatio and a
bottleneck verdict; phase is AwaitingData until ranks are reported. The straggler and bottleneck rules are simple
heuristics on that data. Nothing here has run against a real NCCL job.
apiVersion: gryvia.io/v1alpha1
kind: GryviaTrainingInsight
metadata:
name: llm-training-insight
namespace: ml-research
spec:
targetJob: llm-distributed-training
analysisWindow: "5m"
metrics:
- collective_timing
- straggler_detection
- communication_ratio
- pattern_analysis
GryviaInferenceInsight
Latency breakdown for an inference service (status.latencyBreakdown: DNS, TCP connect, TLS handshake, GPU queue, GPU
execution, postprocess; p50/p95/p99TotalNs; bottleneck). The controller asks the collector for
/api/v1/inference/latency, which the collector does not serve, so the status stays empty today. The infer_latency
program measures only accept-to-first-read wait, not this full breakdown.
apiVersion: gryvia.io/v1alpha1
kind: GryviaInferenceInsight
metadata:
name: llm-serving-insight
namespace: ml-production
spec:
targetService: llama-serving
analysisWindow: "5m"
GryviaFabricSignal
The CRD exists for scheduler-facing fabric signals (see the fabric programs above), but no controller fills it and it is not part of the operator's ten kinds.
CLI Commands
The CLI reads Kubernetes objects with your kubeconfig; it does not call the collector or the gateway. That determines what each command can show:
| Command | Reads | Note |
|---|---|---|
gryvia network flows, graph, trace --follow | GryviaFlow objects (labelled gryvia.io/service) | There is no GryviaFlow CRD in crds/ and nothing creates such objects, so these print a warning or an empty result on a stock install. |
gryvia network trace | creates a GryviaTraceSession | See the trace caveats above. |
gryvia network policy list/suggest/apply | GryviaFlowPolicy / GryviaAutoPolicy | Works on whatever the operator stores. |
gryvia network anomalies, status | GryviaNetworkAnomaly and related objects | Empty until anomalies are produced. |
gryvia security alerts/status/policy | GryviaSecurityPolicy (status.detectionCounts) | Depends on the collector data path. |
gryvia gpu nccl, training | GryviaTrainingInsight | Needs an insight object for the job. |
gryvia gpu memory, rdma | GryviaGpuNode / placeholder | Print a hint or placeholder values; they do not query the collector. |
For real flows in a browser, use the dashboard with Netra configured (GET /api/network/flows). See the
CLI guide for every flag.
Network Tracing
gryvia network trace training-worker --duration 5m --trace-namespace ml-research
gryvia network trace training-worker --level l4 --duration 2m --trace-namespace ml-research
gryvia network trace training-worker --follow --trace-namespace ml-research
kubectl get gryviatracesessions -n ml-research
Flow Analysis and Service Graph
gryvia network flows --flow-namespace ml-research
gryvia network flows --service training-worker --flow-namespace ml-research --last 1h
gryvia network flows --flow-namespace ml-research --output json
gryvia network graph --graph-namespace ml-research
gryvia network graph --graph-namespace ml-research --format json > graph.json
Policy Management
gryvia network policy list --policy-namespace ml-research
gryvia network policy suggest --policy-namespace ml-research
gryvia network policy apply suggestion-name --policy-namespace ml-research
# Flow policies and security policies are ordinary manifests
kubectl apply -f flow-policy.yaml
To evaluate a policy without enforcing it, set a GryviaAutoPolicy's mode to learn or suggest instead of enforce.
Anomalies and Security
gryvia network anomalies --anomaly-namespace ml-research
gryvia network anomalies --severity critical --anomaly-namespace ml-research
gryvia network anomalies --service training-worker --anomaly-namespace ml-research
gryvia security alerts --security-namespace ml-research
gryvia security alerts --severity critical --security-namespace ml-research
gryvia security alerts --alert-type mining --security-namespace ml-research
gryvia security status
gryvia security policy list --security-namespace ml-research
GPU Network Analysis
gryvia gpu nccl --job llm-distributed-training
gryvia gpu memory --node gpu-node-01
gryvia gpu rdma --node gpu-node-01
gryvia gpu training --job llm-distributed-training
kubectl get gryviatraininginsights -n ml-research
Deployment
Prerequisites
- A node kernel with BTF (
/sys/kernel/btf/vmlinux,CONFIG_DEBUG_INFO_BTF=y; standard on Ubuntu 22.04+, RHEL 9, Debian 12+); Linux 6.6+ for thetcxprograms. - A privileged DaemonSet is acceptable in your cluster (the collector runs as root,
hostNetwork,hostPID). - Cilium as the CNI, for the CiliumNetworkPolicy-based features (FlowPolicy, AutoPolicy enforce, anomaly and security blocking).
- NVIDIA drivers and the NCCL/CUDA libraries visible to the collector for the GPU uprobes (
ebpf.ncclLib,ebpf.cudaLib).
Install the Operator and Collector
The chart is helm/network-intelligence; it is separate from helm/gryvia. The operator is always installed; the
collector DaemonSet only with ebpf.enabled=true:
helm install network-intelligence ./helm/network-intelligence \
--set ebpf.enabled=true \
--set ebpf.interface=eth0
Relevant values (see helm/network-intelligence/values.yaml): operator.*, collector.*, ebpf.enabled,
ebpf.interface (XDP/tcx attach), ebpf.cgroupPath (sockops), ebpf.ncclLib / ebpf.cudaLib,
ebpf.flightTokenSecret (Flight Recorder token), prometheus.serviceMonitor.*, namespace.name (default
gryvia-network). The collector image is not published with releases; build it yourself (collector/Dockerfile) and set
collector.image.repository and collector.image.tag.
Verify
kubectl get pods -n gryvia-network
kubectl logs -n gryvia-network -l app.kubernetes.io/component=collector --tail 100
# per-node attach results (from a pod that can reach the node, or kubectl port-forward)
curl -s http://<node-ip>:9090/api/v1/ebpf/status
Best Practices
- Start with the core programs (
tcp_trace,latency_probe,dns_tracker); the GPU programs need the NCCL/CUDA library paths and hardware. - Keep
ebpf.enabled=falseon clusters where a privileged, hostPID DaemonSet is not acceptable, and use Netra for flows instead. - Keep
autoBlockandautoMitigateoff until you have confirmed on a test cluster that the resulting CiliumNetworkPolicy objects match your intent. - Use
learnorsuggestmodes for GryviaAutoPolicy and review suggestions before accepting them, especially in multi-tenant clusters. - Use GryviaTraceSession for targeted debugging rather than cluster-wide capture (once flow capture is implemented).
- Treat all thresholds in the programs (2x straggler ratio, 0.3 overlap idle ratio, 100 CNP per second) as uncalibrated heuristics.
Support
- Issues: https://github.com/zyvorai/gryvia/issues
- Discussions: https://github.com/zyvorai/gryvia/discussions