Skip to content

FEATURES

🏥 Health Monitoring & Uptime Tracking

Track workload health over time with historical records, uptime calculations, and timeline views.

📑 Table of Contents


🔎 Overview

aether's health monitoring system records periodic health check observations for each workload, building a historical timeline that enables:

  • 📈 Uptime percentage calculations (ratio of ready checks to total checks)
  • 🔄 Restart tracking across time
  • 🕐 Health timeline showing state transitions
  • 📋 Summary statistics per workload
  • 🗂️ Bounded storage with automatic pruning of old records

Health data is stored in ~/.aether/health.json and is updated whenever you run status checks, background watches, or orchestration health checks.


📊 Health Records

Each health check produces a HealthRecord with the following fields:

Field Type Description
timestamp string (RFC 3339) When the health check occurred (e.g. 2026-01-15T10:00:00Z)
workload string Name of the workload
runtime RuntimeKind Runtime hosting the workload: podman, kubernetes, kubevirt
state InstanceState Instance state at check time: pending, running, stopped, failed, unknown
ready bool Whether the workload was healthy/ready at check time
restart_count u32 Cumulative restart count reported by the runtime
latency_ms Option\<f64> Health check probe latency in milliseconds (null if not measured)

Example Record (JSON)

{
  "timestamp": "2026-01-15T10:05:00Z",
  "workload": "api-service",
  "runtime": "kubernetes",
  "state": "running",
  "ready": true,
  "restart_count": 0,
  "latency_ms": 12.5
}

Instance States

State Icon Description
running 🟢 Workload is running normally
pending 🟡 Workload is starting or transitioning
stopped ⚪ Workload was intentionally stopped
failed 🔴 Workload has failed
unknown ❓ State cannot be determined

📈 Uptime Calculation

Uptime is calculated as the percentage of health checks where ready == true:

uptime_percent = (ready_checks / total_checks) * 100.0

Examples

Scenario Ready Checks Total Checks Uptime
Always healthy 100 100 100.00%
3 of 4 healthy 3 4 75.00%
Two-thirds healthy 2 3 66.67%
Never healthy 0 50 0.00%
No records 0 0 0.00%

Key Behaviors

  • Uptime is calculated per workload -- records from other workloads are excluded
  • If no records exist for a workload, uptime returns 0.0
  • The calculation spans all retained records (up to max_records)
  • Multiple workloads are fully isolated: good uptime on "web" is not affected by "api" failures

🕐 Timeline View

The timeline view shows the most recent health records for a workload in chronological order (oldest first).

aether health api-service                  # Default: last 20 records
aether health api-service --last 10        # Last 10 records
aether health api-service --last 50        # Last 50 records

Example Timeline Output

╭────────────────────────┬───────────┬─────────┬───────┬──────────┬────────────╮
│ Timestamp              │ Runtime   │ State   │ Ready │ Restarts │ Latency    │
├────────────────────────┼───────────┼─────────┼───────┼──────────┼────────────┤
│ 2026-01-15T10:00:00Z   │ ☸️ kube   │ running │ ✅    │ 0        │ 11.2ms     │
│ 2026-01-15T10:05:00Z   │ ☸️ kube   │ running │ ✅    │ 0        │ 12.5ms     │
│ 2026-01-15T10:10:00Z   │ ☸️ kube   │ failed  │ ❌    │ 1        │ --         │
│ 2026-01-15T10:15:00Z   │ ☸️ kube   │ running │ ✅    │ 1        │ 15.0ms     │
╰────────────────────────┴───────────┴─────────┴───────┴──────────┴────────────╯
  Uptime: 75.00%  |  3 / 4 checks ready  |  1 restart(s)

Timeline Behavior

Situation Behavior
--last N with fewer than N records Returns all available records
--last 0 Returns empty timeline
No records for workload Returns empty timeline
Records from other workloads Filtered out (only target workload shown)

📋 Summary View

Get a high-level summary of a workload's health without individual records.

aether health api-service --summary

Summary Fields

Field Type Description
workload string Workload name
total_checks usize Total number of health checks recorded
ready_checks usize Number of checks where the workload was ready
uptime_percent f64 Uptime percentage (0.0 -- 100.0)
last_restart_count u32 Most recent restart count from the runtime
last_state string String representation of the most recent instance state

Example Summary Output

╭────────────────────┬────────────────╮
│ Property           │ Value          │
├────────────────────┼────────────────┤
│ Workload           │ api-service    │
│ Total Checks       │ 156            │
│ Ready Checks       │ 152            │
│ Uptime             │ 97.44%         │
│ Last Restart Count │ 2              │
│ Last State         │ running        │
╰────────────────────┴────────────────╯

Edge Cases

Scenario total_checks ready_checks uptime_percent last_state
No records 0 0 0.0 "unknown"
All ready N N 100.0 "running"
Never ready N 0 0.0 "failed"

🐳 Podman Native Health Checks

When deploying to Podman, Aether automatically maps workload health probes to native Podman health check flags. This enables container-level health monitoring without external tooling.

Probe Type Mapping

Probe Type Podman --health-cmd
httpGet curl -sf http://localhost:{port}{path} \|\| exit 1
tcpSocket `bash -c '