Skip to main content

INTEGRATIONS

External System Integrations​

How to use Gryvia together with other ML and platform tools.

:::caution There are no built-in connectors Gryvia has no integration controller and no integrations: field on GryviaAIJob. It does not talk to Weights & Biases, Neptune, MLflow, HuggingFace, Airflow, Datadog, Slack, LDAP and so on by itself, and it does not log anything to those services for you. Earlier versions of this page described such features; they never existed. What is real:

  • A GryviaAIJob runs your container. Anything your training code does (log to W&B, push to the Hub) works if the container has network access and the credentials, which you pass with spec.env and Kubernetes Secrets.
  • External tools can create and read jobs through the gateway REST API, the CLI (which uses your kubeconfig), the Python SDK (REST) or the Go SDK (Kubernetes API).
  • Metrics are exposed for Prometheus (DCGM exporter, operator metrics) and can be scraped by whatever you run.
  • Sign-in through an OIDC identity provider is implemented in the gateway (Authentication and TLS).

The recipes below are ordinary Kubernetes usage. None has been tested against the third-party service it mentions. :::

CRDs such as GryviaJobHook, GryviaDataset, GryviaWorkflow, GryviaModelRegistry, GryviaBudget and GryviaSLA exist in crds/, but no controller is wired for them yet; objects of those kinds are stored and nothing acts on them. See reference/crds.md for the list of kinds that do have a controller.

Experiment Tracking​

Weights & Biases​

Put the key in a Secret and hand it to your training process as an environment variable. The logging itself is done by your script (wandb.init(...)); Gryvia does not add GPU metrics, cost or job metadata to your W&B run.

kubectl create secret generic wandb-api-key --from-literal=api-key=YOUR_API_KEY
apiVersion: gryvia.io/v1alpha1
kind: GryviaAIJob
metadata:
name: training-with-wandb
spec:
type: training
gpus: 1
gpuType: any
image: nvcr.io/nvidia/pytorch:24.01-py3
command: ["python", "train.py", "--wandb-project=my-project"]
env:
- name: WANDB_API_KEY
valueFrom:
secretKeyRef:
name: wandb-api-key
key: api-key
- name: WANDB_ENTITY
value: my-team

Neptune, MLflow and TensorBoard​

The same pattern applies: pass the token or tracking URI as spec.env (for example NEPTUNE_API_TOKEN, MLFLOW_TRACKING_URI) and log from your code. Gryvia does not create a TensorBoard Service; run TensorBoard yourself against a shared volume (spec.storage gives a job a PVC mounted at /data, see the platform setup guide) and expose it with your own Deployment and Service.

Model Registries​

HuggingFace Hub and other registries​

Do the upload at the end of your training script, or as a follow-up Kubernetes Job you create yourself, with the token from a Secret (kubectl create secret generic hf-token --from-literal=token=YOUR_HF_TOKEN). Post-completion automation is the purpose of the GryviaJobHook CRD, but it has no controller today. This is its schema, shown as a design sketch (it validates against the CRD, and nothing will run it):

apiVersion: gryvia.io/v1alpha1
kind: GryviaJobHook
metadata:
name: push-to-huggingface
spec:
trigger: post-completion
action:
type: k8sJob
k8sJob:
image: python:3.11
command: ["python", "push_model.py"]

GryviaModelRegistry is likewise schema-only. There is no SageMaker or MLflow registry integration.

Data Platforms​

GryviaDataset (CRD only; the storage operator contains a dataset reconciler that is not registered in main.go) does not fetch, version or mount data. Use whatever you already use (DVC, Pachyderm, object storage) inside the job container or an init container, with credentials from Secrets, and a PVC or CSI volume for shared data (see Storage and Network Operators).

Workflow Orchestration​

Any orchestrator that can run kubectl or the Gryvia CLI in a container with cluster credentials can create jobs. There is no published Gryvia CLI container image; the release provides CLI binaries (gryvia-<tag>-<os>-<arch> on the GitHub release), so build a small image that contains the binary and mount a kubeconfig or use an in-cluster service account.

Airflow​

from airflow import DAG
from airflow.providers.cncf.kubernetes.operators.pod import KubernetesPodOperator

with DAG('ml_training', schedule='@daily') as dag:
train = KubernetesPodOperator(
task_id='train_model',
namespace='default',
name='submit-job',
image='registry.example.com/your-org/gryvia-cli:latest', # your own image containing the gryvia binary
cmds=['gryvia'],
arguments=['submit', '--file', '/jobs/job.yaml', '--wait'],
)

Kubeflow Pipelines and Argo work the same way: run gryvia submit --file job.yaml --wait, or kubectl apply, or call POST /api/jobs on the gateway from a step.

Monitoring and Observability​

Gryvia does not push metrics to Datadog, New Relic or Grafana Cloud, and it does not define gryvia_gpu_utilization, gryvia_job_duration or gryvia_cost_total Prometheus metrics. What exists:

  • The DCGM exporter (port 9400) on GPU nodes and the operators' controller-runtime metrics endpoints, scrapeable by any Prometheus. helm/observability wraps kube-prometheus-stack; monitoring/ has a ServiceMonitor, alert rules and dashboards as templates.
  • Gateway routes that read Prometheus when apiGateway.prometheusUrl is set (/api/metrics/gpu and cost history). The gateway itself has no /metrics endpoint.
  • To forward metrics to a SaaS, use that vendor's Prometheus scraper or remote-write against the Prometheus you run.

CI/CD Integration​

The CLI talks to the Kubernetes API with a kubeconfig, so a CI runner needs the gryvia binary and a kubeconfig for a service account that may create GryviaAIJob objects.

GitHub Actions​

name: Train Model
on:
push:
branches: [main]

jobs:
train:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4

- name: Install gryvia
run: |
# download the release binary for linux-amd64 and put it on PATH (see the GitHub release page)
echo "install gryvia here"

- name: Submit training job
run: gryvia submit --file job.yaml --wait
env:
KUBECONFIG: ${{ github.workspace }}/kubeconfig # written from a secret in an earlier step

- name: Job status
run: gryvia status training-job

GPU type and count are set in job.yaml (spec.gpuType, spec.gpus); gryvia submit has no --gpu-type or --gpu-count flag. --wait waits for completion and --logs follows logs. Adapt the same commands for GitLab CI or Jenkins (gryvia submit --file job.yaml --wait, gryvia logs training-job).

Notifications​

Webhooks that exist today​

The network-intelligence CRDs accept a webhook URL that the operator calls: alertWebhook on GryviaNetworkAnomaly and GryviaSecurityPolicy (see Network Intelligence; those controllers need Cilium and a working collector path). Point it at a Slack incoming webhook or any HTTP receiver. There is no job-completion notification, no Slack bot and no /gryvia slash command.

Job hook (design sketch)​

GryviaJobHook describes a webhook on job completion; nothing reconciles it today:

apiVersion: gryvia.io/v1alpha1
kind: GryviaJobHook
metadata:
name: slack-notifications
spec:
trigger: post-completion
action:
type: webhook
webhook:
url: https://hooks.slack.com/services/YOUR/WEBHOOK/URL
method: POST
body: '{"text": "Job completed"}'

Cost Management​

Cost estimates come from the quota operator (GryviaCostPredictor, GryviaUsageRecord) and the SKU catalog: gryvia cost, gryvia usage, gryvia invoice, gryvia catalog, and the gateway's /api/metrics/costs and /api/usage/* routes. There is no CloudHealth or Kubecost connector. To export usage, use GET /api/usage/export (see the API reference).

Authentication​

OIDC sign-in is implemented in the API gateway; LDAP is not supported. Configure OIDC through the gateway environment and the chart values apiGateway.oidc.*, and map users to tenants with GryviaTenant. See Authentication and TLS and GPU as a Service. Payments and a real identity provider have not been exercised end to end.

Backup and DR​

State lives in custom resources; see Operations for export, restore and tools/backup-restore.sh. A generic Velero schedule is a possible addition (not tested with Gryvia). Many Gryvia kinds are cluster-scoped, so include cluster resources:

apiVersion: velero.io/v1
kind: Schedule
metadata:
name: gryvia-backup
namespace: velero
spec:
schedule: "0 2 * * *"
template:
includedNamespaces:
- gryvia-system
includeClusterResources: true
includedResources:
- gryviaaijobs.gryvia.io
- gryviaquotas.gryvia.io
- gryviatenants.gryvia.io
- persistentvolumeclaims

There is no GryviaJobHook-based checkpoint sync; use your own sidecar or script (aws s3 sync and similar) to copy checkpoints, or the checkpoint features of your framework.

Security Scanning​

Image scanning is independent of Gryvia. For example, a plain Kubernetes CronJob running Trivy against an image you use:

apiVersion: batch/v1
kind: CronJob
metadata:
name: image-scanning
spec:
schedule: "0 */6 * * *"
jobTemplate:
spec:
template:
spec:
restartPolicy: Never
containers:
- name: trivy
image: aquasec/trivy:0.58.0
args: ["image", "--severity=HIGH,CRITICAL", "nvcr.io/nvidia/pytorch:24.01-py3"]

Release images are signed with cosign in the release workflow.

SDKs​

Python SDK​

An async REST client for the gateway (jobs, nodes, quotas, costs, metrics). Install from source (cd sdk/python && pip install -e .); publication to PyPI is not verified. It has unit tests against mocked HTTP only. It does not cover SKUs, tenants, usage, invoices or Flight Recorder, and Jobs.stream_logs does not work against the current gateway (see sdk/python/README.md).

import asyncio
from gryvia import Gryvia

async def main():
async with Gryvia(api_url="https://gryvia.example.com", token="your-api-token") as tr:
job = await tr.jobs.create({
"apiVersion": "gryvia.io/v1alpha1",
"kind": "GryviaAIJob",
"metadata": {"name": "my-training"},
"spec": {
"type": "training",
"image": "nvcr.io/nvidia/pytorch:24.01-py3",
"gpus": 4,
"gpuType": "A100-80G",
"command": ["torchrun", "--nproc_per_node=4", "train.py"],
},
})
final = await tr.jobs.wait_for_completion("my-training", timeout=7200)
print(final.status.phase)

asyncio.run(main())

Go SDK​

github.com/zyvorai/gryvia/sdk/go is a typed controller-runtime client for the CRDs (it uses the Kubernetes API, not the gateway). It wraps create/get/list/delete for jobs and several other kinds (whether a controller acts on those kinds is a separate question, see above).

import (
sdk "github.com/zyvorai/gryvia/sdk/go"
"sigs.k8s.io/controller-runtime/pkg/client/config"
)

cfg, _ := config.GetConfig()
c, _ := sdk.NewClient(cfg)

job := &sdk.GryviaAIJob{}
job.Name = "my-training"
job.Namespace = "ml-training"
job.Spec.Type = "training"
job.Spec.Image = "nvcr.io/nvidia/pytorch:24.01-py3"
job.Spec.GPUs = 8

err := c.CreateJob(ctx, job)

Check sdk/go/client.go and the CRD types in operators/ai-operator/api/v1 for exact field names.

REST API​

Gateway routes live under /api/ (there is no /api/v1 prefix) and take Authorization: Bearer <token>; the API reference lists every route and its access rule. Creating a job:

curl -k -X POST https://<dashboard-host>/api/jobs \
-H "Authorization: Bearer $GRYVIA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"apiVersion": "gryvia.io/v1alpha1",
"kind": "GryviaAIJob",
"metadata": {"name": "training-job"},
"spec": {
"type": "training",
"image": "nvcr.io/nvidia/pytorch:24.01-py3",
"gpus": 8,
"gpuType": "A100-80G",
"command": ["python", "train.py"]
}
}'

curl -k https://<dashboard-host>/api/jobs/training-job -H "Authorization: Bearer $GRYVIA_API_KEY"
curl -k "https://<dashboard-host>/api/jobs/training-job/logs?tail=100" -H "Authorization: Bearer $GRYVIA_API_KEY"

The gateway sets the namespace itself. There is no GraphQL API.

Best Practices​

Secret management​

Use Kubernetes Secrets for API keys and tokens, and reference them with secretKeyRef:

kubectl create secret generic wandb-api-key --from-literal=api-key=YOUR_API_KEY
kubectl create secret generic hf-token --from-literal=token=YOUR_HF_TOKEN

Least privilege for automation​

Give CI and orchestrators a dedicated service account with only the rights to create and read gryvia.io jobs in the namespaces they need, or an OIDC tenant identity when going through the gateway, rather than the admin API key.

Troubleshooting​

# Platform and component status
gryvia status
kubectl get pods -n gryvia-system

# Why a job did not start
gryvia get job <name>
kubectl describe gryviaaijob <name>
kubectl logs -n gryvia-system deploy/gryvia-ai-operator

# A secret referenced by a job
kubectl get secret wandb-api-key

Support​


See ADVANCED_FEATURES.md for more information