Skip to main content

Operations: upgrade, uninstall, backup, troubleshooting

Upgrade​

Helm installs CRDs on the first install only and never upgrades them, so apply the CRDs first:

kubectl apply --server-side --force-conflicts -f crds/ # from the release you are upgrading to
helm upgrade gryvia oci://ghcr.io/zyvorai/charts/gryvia -n gryvia-system --reuse-values --version <new>

Gryvia is alpha and gryvia.io/v1alpha1 may change between releases; read the changelog before upgrading. Two changes are breaking: the Fabric* → Gryvia* kind rename and the move from gryvia.io/v1 to gryvia.io/v1alpha1: objects created under the old kinds (FabricAIJob, ...) or the old v1 version are not converted. Export them, change kind: and apiVersion: and re-apply, then delete the old CRDs (scripts/rename-kinds.sh and scripts/rename-version.sh document the mapping).

The self-signed certificate (gryvia-tls) and the API key Secret are kept across upgrades. Open dashboard tabs reload themselves once when a new version is deployed.

Uninstall​

helm uninstall gryvia -n gryvia-system

CRDs, your custom resources and the gryvia-tls Secret are kept on purpose. To remove everything, including data:

kubectl get crd -o name | grep '\.gryvia\.io$' | xargs kubectl delete # deletes every Gryvia object
kubectl delete namespace gryvia-system

Back up​

Gryvia keeps its state in Kubernetes custom resources. Export them all:

kinds=$(kubectl get crd -o name | grep '\.gryvia\.io$' | sed 's#.*/##' | paste -sd, -)
kubectl get "$kinds" -A -o yaml > gryvia-backup.yaml

Restore with kubectl apply -f gryvia-backup.yaml (status fields are recomputed by the operators). Also back up the gryvia-api-key Secret if you generated a random key.

Troubleshooting​

SymptomWhat to check
Pods stuck in ImagePullBackOffThe images are pulled from global.imageRegistry (default ghcr.io/zyvorai). Check the tag exists, or set imagePullSecrets for a private mirror.
Dashboard shows "Wrong username or password."Sign in as admin with the API key. Read a generated key with kubectl -n gryvia-system get secret gryvia-api-key -o jsonpath='{.data.GRYVIA_API_KEY}' | base64 -d.
"Too many sign-in attempts"Sign-in is limited to 10 per minute per client address. Wait a minute.
Browser certificate warningExpected with the default self-signed certificate. See Authentication and TLS to use your own.
A page says "Gryvia was updated"A new version was deployed while the tab was open. Reload.
Nodes show phase Failed with "node not found"The GryviaGpuNode names a Kubernetes node that does not exist (the kind demo data is fictional). Register real nodes.
GPU or network pages say no source is connectedThey need the collectors and operators that feed them (DCGM exporter, network-intelligence, security operator). The page names what is missing.
A job stays Pendingkubectl -n <ns> describe gryviaaijob <name> shows the scheduling reason, for example no node with the requested GPU type or a quota limit.
Operators restart or lose leadershipkubectl -n gryvia-system logs deploy/gryvia-gpu-operator (and -ai-, -quota-). Leader election uses leases in gryvia-system.

Collect logs for a bug report with kubectl -n gryvia-system logs -l app=gryvia-api-gateway --tail=200 and the operator deployments above; remove secrets before sharing.