Credential sources (Vault, files)
Status. All three steps are built: the
filesource, thevaultsource with token, AppRole and Kubernetes auth, the status output and the chart values. See vault/README.md. They follow the recommendations below. Nothing has been run against a real Vault; the behaviour is tested against a mock.
Why
Every credential secret lives in the runtime's process environment. A descriptor names a variable ("env": "STRIPE_KEY") and the runtime reads it
with std::env::var on each use. That is simple, and it has four costs:
- Custody. The secret sits in a unit file, an env file or a Kubernetes manifest, so it is copied to wherever the runtime is configured, and it
is readable in
/proc/<pid>/environby anyone who can read the process. - Rotation. Changing a key means editing the env and restarting the runtime, which ends running sessions.
- No central audit or revocation. A company that keeps its secrets in Vault, or in Kubernetes Secrets mounted by a CSI driver, cannot see or cut off Keep's use of them from the place that owns them.
- The Helm chart (
charts/zyvor-keep) has no good answer: secrets go throughruntime.extraEnv, which puts them in the pod spec.
This note is about where the secret comes from. It does not change who may use it: host, method, path, port, user and approval checks stay exactly as they are.
What exists today
Checked against the code (agent-runtime/src/credentials.rs unless stated):
- A
CredentialDescriptorhasenv(the variable name) andkind:provider(default, injects the env secret),fabric(optionalFABRIC_AI_API_KEY) oroauth-refresh(client id, client secret and refresh token, each from an env variable, and a token minted and cached in memory by a background task). - Descriptors come from the file named by
ZYVOR_AGENT_CREDENTIALS_FILE, read once at startup (CredentialVault::load). No API route writes a descriptor. That is worth keeping: a network caller can never change where a secret is read from. authorize_resolvechecks the descriptor's host, method, path prefixes, ports and allowed users first, and only then callsresolve_for, which reads the secret. So a request that fails policy never touches a secret.resolve_foris synchronous and does no I/O beyondstd::env::var. It has three callers:egress.rs(proxy_inner, async),browser.rs(browser_fill_secret, async) andmodel_call.rs(authorize, a plainfn).- The unwrap ceremony (
ZYVOR_AGENT_VAULT_UNWRAP_REQUIRED=1) gates injection behind an operator unlock. It sits in front of resolution (egress.rs,model_call.rs) and is independent of where the value comes from. docs/keep/vault/README.mdis honest about the limit, and this note keeps that honesty: the secret is in host process memory, so the host operator can read it. A credential source changes where the secret is kept and who can rotate or revoke it, not whether the operator can read it while Keep is using it.oauth-refreshalready shows the pattern for a secret that expires: atokensmap behindArc<RwLock<..>>, avalid_until, and a failed refresh fails closed.- There is no
vaultbinary and no Docker on the machines I test on, so a real Vault cannot be run here.
Principles
- Same checks, different source. Authorisation stays in front; a secret is fetched only after a request passes it.
- Fail closed. If a secret cannot be fetched, the request is refused. A cached value may be used until its own TTL runs out and not a moment longer; there is no "serve stale during an outage".
- Never in a log, an error or an audit row. Errors name the credential and the class of failure ("the secret store refused the request"), never a response body, which can echo paths or policy.
- Operator-only configuration. A
sourceis read from the credentials file at startup likeenvis. Nothing an agent or a user token can call may create or change one. envkeeps working, unchanged. A descriptor withenvand nosourcebehaves exactly as it does now.
Proposal
A descriptor may carry a source instead of env:
{
"stripe": {
"host": "api.stripe.com", "header": "authorization", "prefix": "Bearer ",
"allowed_methods": ["POST"], "requires_approval": ["POST"],
"source": { "kind": "file", "path": "/run/secrets/stripe-key" }
},
"github": {
"host": "api.github.com", "header": "authorization", "prefix": "Bearer ",
"source": {
"kind": "vault", "addr": "https://vault.internal:8200", "mount": "secret",
"path": "keep/github", "field": "token", "ttl_seconds": 300,
"auth": { "method": "kubernetes", "role": "keep" }
}
}
}
Two source kinds, in this order:
file. Read the secret from a file, trimmed, cached forttl_seconds(default 60). This one small source covers Vault Agent (which renders secrets to files), Kubernetes Secrets mounted as volumes, the CSI Secrets Store driver, systemdLoadCredential=, andsops/agetooling. It needs no client, no network and no new dependency, and rotation works by replacing the file.vault. HashiCorp Vault KV v2 over its HTTP API:GET /v1/{mount}/data/{path}withX-Vault-Token, the value atdata.data.{field}. Authentication bytoken(from a file or variable, for Vault Agent),approle(role_idandsecret_id) orkubernetes(the pod's service account JWT and a role). A 403 triggers one re-login and one retry, no more.
Resolution changes shape. A new async CredentialVault::resolve_secret is used for sourced credentials; env, fabric and oauth credentials keep
the synchronous path. authorize_resolve becomes async (its three callers are already async or one small step from it: model_call::authorize
becomes async). The cache is a RwLock<HashMap<name, {value, valid_until}>> like the OAuth token map, with one fetch in flight per
credential, so a burst of requests after expiry makes one call to the store, not many.
Guards on vault: addr must be https (plain http only to loopback, the rule the OAuth token endpoint already follows); redirects are
never followed; the response is size-capped and the request has a timeout; mount, path and field are restricted to a plain character set with
no ..; an optional CA file for private PKI. On file: the path must be absolute, no .., must be a regular file, and a file readable by group or
others is refused at startup with a clear message (it is a secret).
Visibility, without values. GET /v1/vault/status (and keepctl vault-status) gains, per credential, the source kind, the last successful
fetch time, and the last failure class. It never returns a value, and neither do the journal or the cockpit.
What I would refuse to build
- Any API that writes a source, or lets a session, agent or user token choose a path. That would let a caller point a credential at a secret it should not reach.
- Templated paths from request data (
keep/users/{user_id}/token) in this step. It is useful for per-person credentials, but it puts untrusted input into a path on the secret store and needs its own design. - Serving a stale secret while the store is down. A revoked secret must stop working when its TTL ends.
- Reading the store from inside the cell, or handing a Vault token to a cell. The host fetches; the cell never sees a token or a path.
- Fetching every secret at startup and holding it forever. The point is rotation and revocation.
- A claim that this protects a secret from the host operator. It does not, and the docs will keep saying so.
- A bundled Vault client SDK. The KV v2 read and the three login calls are a handful of HTTP requests; a dependency is not worth it.
How I would verify it
- Unit tests with a mock Vault (an HTTP server in the test): KV v2 read and field extraction, each login method, a 403 causing exactly one
re-login and retry, TTL expiry, one fetch per burst, a refusal on a 5xx, a timeout, an oversized body, a redirect, a malformed response, and
httpto a non-loopback host. Tests forfile: trimming, rotation after the TTL, a missing file, a group-readable file refused,..refused. - A test through the real broker path: a sourced secret is injected on an allowed request, a request that fails the descriptor's checks never causes a fetch (the mock records zero calls), and no secret value appears in the journal, an error or the cockpit.
- The existing credential tests and
keep-e2erun unchanged, which is the test thatenvstill works. - What I cannot verify here: a real Vault. The mock proves the requests and the failure handling, not that a particular Vault version or auth backend accepts them. Someone with a Vault should run the documented dev-server steps once before this is trusted; I would write them down.
Size
Three PRs, each reviewable on its own:
sourcein the descriptor, the async resolution path, the cache, and thefilesource, with docs. About 500 lines plus tests. Theauthorizechange inmodel_call.rsis the only edit outsidecredentials.rsand its callers.- The
vaultsource with token auth. - AppRole and Kubernetes auth, the status output, and chart values (mount the service account or a Secret volume; no secrets in
extraEnv).
oauth-refresh credentials (client secret and refresh token) are not in these three; see decision 3.
Decisions for you
- Order. Recommendation:
filefirst, thenvault, becausefilealone covers Kubernetes Secrets, the CSI driver and Vault Agent, and is the smaller step to review. The alternative isvaultfirst if a native client is what you actually need. - Vault auth methods. Recommendation: token (for Vault Agent), AppRole, and Kubernetes, in that order. Tell me if you only use one, and I will build only that.
- OAuth credentials. Should the client secret and refresh token of
oauth-refreshcredentials also be readable from a source? Recommendation: not in the first three PRs. Per-person refresh tokens are stored by the runtime itself, so the shared client secret is the only piece that would move, and it can wait. - TTL defaults. Recommendation: 60 s for
file, 300 s forvault, both capped at one hour, and no stale serving. A shorter TTL revokes faster and calls the store more. - Startup checks. Should a sourced credential be fetched once at startup so a wrong path or a denied policy fails at boot, not on the first request? Recommendation: yes, log a warning but do not refuse to start, since the store may come up after Keep.
- Namespaces and other stores. Do you need Vault Enterprise namespaces (
X-Vault-Namespace), or AWS or GCP secret managers directly? Recommendation: anamespacefield is cheap and I would include it; cloud managers are better reached through Vault or a mounted file than through their SDKs.