USER GUIDE
Prometheus metrics
Purpose¶
Scrape hook execution, event throughput, routing-rule counts, and netlink health for fleet observability — the same HTTP server as the REST API on :9090.
When to use it¶
- Alert when
netevd_script_failures_totalrises after a hook deploy - Dashboard event rates per interface/backend
- Detect silent daemon failure (uptime gauge flatlines)
How to get there¶
- Enable:
metrics.enabled: truein/etc/netevd/netevd.yaml - Endpoint:
http://<host>:9090/metricson the API server (Prometheus text exposition) metrics.port(default 9091) is stored in the config and logged at startup; the daemon serves/metricsonapi.port
Operate from CLI¶
- Confirm metrics are enabled:
grep -A3 '^metrics:' /etc/netevd/netevd.yaml
curl -sf http://127.0.0.1:9090/metrics | head -20
- Inspect key series locally:
curl -sf http://127.0.0.1:9090/metrics | grep -E '^netevd_(info|events_total|script_|routing_rules|ebpf_)'
- Prometheus scrape config (replace
<host>with your edge/management target):
scrape_configs:
- job_name: netevd
scrape_interval: 30s
static_configs:
- targets: ['<host>:9090']
labels:
role: network-events
- Example alert rules (hook failures):
groups:
- name: netevd
rules:
- alert: NetevdHookFailures
expr: increase(netevd_script_failures_total[5m]) > 0
for: 2m
labels:
severity: warning
- alert: NetevdDown
expr: absent(netevd_info)
for: 5m
labels:
severity: critical
- Remote check from your workstation:
curl -sf http://<host>:9090/metrics | grep netevd_uptime_seconds
-
Empty / fail: Connection refused →
metrics.enabled: falseor firewall; series missing → daemon just started or metrics init failed (seejournalctl -u netevd). -
Success:
netevd_info{version="…"}present; counters increment when you bounce a link and routable hooks run.