FEATURES
🏥 Health Monitoring & Uptime Tracking
Track workload health over time with historical records, uptime calculations, and timeline views.
📑 Table of Contents¶
- Overview
- Health Records
- Uptime Calculation
- Timeline View
- Summary View
- Integration with Status Command
- Background Health Checks
- Health in the TUI Dashboard
- REST API Endpoint
- Data Storage and Retention
- Cross-References
🔎 Overview¶
aether's health monitoring system records periodic health check observations for each workload, building a historical timeline that enables:
- 📈 Uptime percentage calculations (ratio of ready checks to total checks)
- 🔄 Restart tracking across time
- 🕐 Health timeline showing state transitions
- 📋 Summary statistics per workload
- 🗂️ Bounded storage with automatic pruning of old records
Health data is stored in ~/.aether/health.json and is updated whenever you run status checks, background watches, or orchestration health checks.
📊 Health Records¶
Each health check produces a HealthRecord with the following fields:
| Field | Type | Description |
|---|---|---|
timestamp |
string (RFC 3339) | When the health check occurred (e.g. 2026-01-15T10:00:00Z) |
workload |
string | Name of the workload |
runtime |
RuntimeKind | Runtime hosting the workload: podman, kubernetes, kubevirt |
state |
InstanceState | Instance state at check time: pending, running, stopped, failed, unknown |
ready |
bool | Whether the workload was healthy/ready at check time |
restart_count |
u32 | Cumulative restart count reported by the runtime |
latency_ms |
Option\<f64> | Health check probe latency in milliseconds (null if not measured) |
Example Record (JSON)¶
{
"timestamp": "2026-01-15T10:05:00Z",
"workload": "api-service",
"runtime": "kubernetes",
"state": "running",
"ready": true,
"restart_count": 0,
"latency_ms": 12.5
}
Instance States¶
| State | Icon | Description |
|---|---|---|
running |
🟢 | Workload is running normally |
pending |
🟡 | Workload is starting or transitioning |
stopped |
⚪ | Workload was intentionally stopped |
failed |
🔴 | Workload has failed |
unknown |
❓ | State cannot be determined |
📈 Uptime Calculation¶
Uptime is calculated as the percentage of health checks where ready == true:
uptime_percent = (ready_checks / total_checks) * 100.0
Examples¶
| Scenario | Ready Checks | Total Checks | Uptime |
|---|---|---|---|
| Always healthy | 100 | 100 | 100.00% |
| 3 of 4 healthy | 3 | 4 | 75.00% |
| Two-thirds healthy | 2 | 3 | 66.67% |
| Never healthy | 0 | 50 | 0.00% |
| No records | 0 | 0 | 0.00% |
Key Behaviors¶
- Uptime is calculated per workload -- records from other workloads are excluded
- If no records exist for a workload, uptime returns
0.0 - The calculation spans all retained records (up to
max_records) - Multiple workloads are fully isolated: good uptime on "web" is not affected by "api" failures
🕐 Timeline View¶
The timeline view shows the most recent health records for a workload in chronological order (oldest first).
aether health api-service # Default: last 20 records
aether health api-service --last 10 # Last 10 records
aether health api-service --last 50 # Last 50 records
Example Timeline Output¶
╭────────────────────────┬───────────┬─────────┬───────┬──────────┬────────────╮
│ Timestamp │ Runtime │ State │ Ready │ Restarts │ Latency │
├────────────────────────┼───────────┼─────────┼───────┼──────────┼────────────┤
│ 2026-01-15T10:00:00Z │ ☸️ kube │ running │ ✅ │ 0 │ 11.2ms │
│ 2026-01-15T10:05:00Z │ ☸️ kube │ running │ ✅ │ 0 │ 12.5ms │
│ 2026-01-15T10:10:00Z │ ☸️ kube │ failed │ ❌ │ 1 │ -- │
│ 2026-01-15T10:15:00Z │ ☸️ kube │ running │ ✅ │ 1 │ 15.0ms │
╰────────────────────────┴───────────┴─────────┴───────┴──────────┴────────────╯
Uptime: 75.00% | 3 / 4 checks ready | 1 restart(s)
Timeline Behavior¶
| Situation | Behavior |
|---|---|
--last N with fewer than N records |
Returns all available records |
--last 0 |
Returns empty timeline |
| No records for workload | Returns empty timeline |
| Records from other workloads | Filtered out (only target workload shown) |
📋 Summary View¶
Get a high-level summary of a workload's health without individual records.
aether health api-service --summary
Summary Fields¶
| Field | Type | Description |
|---|---|---|
workload |
string | Workload name |
total_checks |
usize | Total number of health checks recorded |
ready_checks |
usize | Number of checks where the workload was ready |
uptime_percent |
f64 | Uptime percentage (0.0 -- 100.0) |
last_restart_count |
u32 | Most recent restart count from the runtime |
last_state |
string | String representation of the most recent instance state |
Example Summary Output¶
╭────────────────────┬────────────────╮
│ Property │ Value │
├────────────────────┼────────────────┤
│ Workload │ api-service │
│ Total Checks │ 156 │
│ Ready Checks │ 152 │
│ Uptime │ 97.44% │
│ Last Restart Count │ 2 │
│ Last State │ running │
╰────────────────────┴────────────────╯
Edge Cases¶
| Scenario | total_checks | ready_checks | uptime_percent | last_state |
|---|---|---|---|---|
| No records | 0 | 0 | 0.0 | "unknown" |
| All ready | N | N | 100.0 | "running" |
| Never ready | N | 0 | 0.0 | "failed" |
🐳 Podman Native Health Checks¶
When deploying to Podman, Aether automatically maps workload health probes to native Podman health check flags. This enables container-level health monitoring without external tooling.
Probe Type Mapping¶
| Probe Type | Podman --health-cmd |
|---|---|
httpGet |
curl -sf http://localhost:{port}{path} \|\| exit 1 |
tcpSocket |
`bash -c ' |