CRD reference
Gryvia defines 50 custom resources in the gryvia.io API group (version v1alpha1, alpha: fields may change between releases). The last column names the operator that registers a controller for the kind; none means no controller in operators/*/main.go reconciles the kind: the CRD, gateway routes or dashboard pages may exist, and other components may read it (for example GryviaGpuSku), but nothing acts on its spec. Full schemas are in crds/; kubectl explain <kind>.spec works on any cluster with Gryvia installed.
| Kind | Resource | Scope | Controller | Required spec fields | Other spec fields |
|---|---|---|---|---|---|
GryviaAIJob | gryviaaijobs | Namespaced | ai-operator | gpus, image, type | affinity, args, command, distributed, env, gpuType, imagePullPolicy, model, … |
GryviaAudit | gryviaaudits | Cluster | none | none | anomalyDetection, compliance, events, reporting, retention, scope |
GryviaAutoPolicy | gryviaautopolicies | Namespaced | network-intelligence | none | approvalRequired, excludeServices, learningWindow, mode, targetNamespaces |
GryviaAutoScaler | gryviaautoscalers | Cluster | none | gpuType, maxNodes, queueRef | advanced, costControls, minNodes, nodeProvider, scaleDownPolicy, scaleUpPolicy |
GryviaAutoTuner | gryviaautotuners | Namespaced | none | jobTemplate, maxTrials, objective, parameterSpace, searchAlgorithm | ashaConfig, earlyStoppingRounds, parallelism |
GryviaBenchmark | gryviabenchmarks | Namespaced | none | target, type | baseline, custom, gpuMemory, ioThroughput, mlperf, nccl, schedule |
GryviaBudget | gryviabudgets | Cluster | none | period, scope | alerts, enforcement, limits, priority, rollover |
GryviaChargeback | gryviachargebacks | Cluster | none | none | allocationModel, costCenters, mode, period, pricing, reports |
GryviaCheckpointGuard | gryviacheckpointguards | Namespaced | ai-operator | checkpointPolicy, jobSelector | monitoring, restore, validation |
GryviaCostPredictor | gryviacostpredictors | Cluster | quota-operator | none | alternatives, historicalData, integration, models, pricing |
GryviaDRTest | gryviadrtests | Cluster | none | type | approvalRequired, backupRestore, chaosEngineering, dataIntegrity, failover, fullDrill, notifications, rpoRto, … |
GryviaDataset | gryviadatasets | Cluster | none | source | access, cache, description, license, statistics, tags, type, version, … |
GryviaFabricSignal | gryviafabricsignals | Namespaced | none | none | jobRef, observeOnly |
GryviaFederation | gryviafederations | Cluster | none | none | clusters, costManagement, distribution, failover, loadBalancing, resourceSharing |
GryviaFlowPolicy | gryviaflowpolicies | Namespaced | network-intelligence | none | action, destination, intent, priority, protocol, source |
GryviaGPUSharingPolicy | gryviagpusharingpolicies | Cluster | none | strategy | fractionalGPU, mig, nodeSelector, priority, qos, tenantQuotas, timeSlicing |
GryviaGpuMemoryOptimizer | gryviagpumemoryoptimizers | Cluster | gpu-operator | scope | inferencePacking, oomPrevention, rightSizing |
GryviaGpuNode | gryviagpunodes | Cluster | gpu-operator | gpuCount, gpuType, nodeName | bandwidth, computeCapability, drivers, healthCheck, interconnect, labels, memoryGB, rdma, … |
GryviaGpuSku | gryviagpuskus | Cluster | none | gpuType, hourlyRate | currency, description, enabled, gpusPerUnit, spotDiscount |
GryviaHealthCheck | gryviahealthchecks | Cluster | none | target | checks, onFailure, remediation, schedule |
GryviaInferenceInsight | gryviainferenceinsights | Namespaced | network-intelligence | none | analysisWindow, targetService |
GryviaInferenceService | gryviainferenceservices | Namespaced | none | backend, modelRef | args, autoscaling, canary, gpuCount, gpuType, healthCheck, image, replicas, … |
GryviaJobHook | gryviajobhooks | Cluster | none | action, trigger | condition, failurePolicy, retry, selector |
GryviaLiveExperiment | gryvialiveexperiments | Namespaced | ai-operator | comparison, jobs | description, notifications, strategy |
GryviaMetric | gryviametrics | Cluster | none | name, source | description, labels, retention, thresholds, type, unit, visualization |
GryviaModelLineage | gryviamodellineages | Namespaced | ai-operator | model | compliance, provenance |
GryviaModelRegistry | gryviamodelregistries | Namespaced | none | artifacts, modelName, version | autoServe, description, metadata, servingConfig, source, stage |
GryviaNetwork | gryvianetworks | Cluster | network-operator | networkType | mtu, nodeSelector, rdma, sriov, targetNamespace |
GryviaNetworkAnomaly | gryvianetworkanomalies | Namespaced | network-intelligence | none | alertWebhook, autoMitigate, detectionRules, targetService |
GryviaNetworkCost | gryvianetworkcosts | Namespaced | network-intelligence | none | costCenters, costPerGB, reportingInterval, targetNamespaces |
GryviaNodeFabric | gryvianodefabrics | Cluster | none | measuredAt, nodeName, scoreDelta | expiresAt, reasons, ttlSeconds |
GryviaPriority | gryviapriorities | Cluster | none | value | description, preemptionPolicy, quotaOverride, sla |
GryviaQuota | gryviaquotas | Cluster | quota-operator | gpuQuota, namespaces, team | budget, network, priority |
GryviaQuotaPolicy | gryviaquotapolicies | Cluster | none | none | alerts, allocation, enforcement, hierarchy, limits, scope, timeBased |
GryviaReservation | gryviareservations | Cluster | none | owner, resources, schedule | billing, guarantees, notifications |
GryviaRetryPolicy | gryviaretrypolicies | Cluster | none | none | backoff, budget, circuitBreaker, maxRetries, noRetryOn, resourceAdjustment, retryOn |
GryviaSLA | gryviaslas | Cluster | none | scope, tier | availability, failureHandling, monitoring, performance, resources, support |
GryviaSecurityPolicy | gryviasecuritypolicies | Namespaced | network-intelligence | none | alertWebhook, autoBlock, detectionRules, targetNamespaces |
GryviaServiceGraph | gryviaservicegraphs | Namespaced | network-intelligence | none | depth, includeExternal, namespaces, refreshInterval |
GryviaStorage | gryviastorages | Cluster | storage-operator | backend, capacity, endpoint | credentials, iops, mountOptions, performance, protocol, quotas, rdma, storageClass, … |
GryviaTemplate | gryviatemplates | Cluster | none | category, defaults | description, parameters, tags |
GryviaTenant | gryviatenants | Cluster | quota-operator | none | allowedSkus, billing, description, displayName, governance, jobDefaults, members, networkPolicy, … |
GryviaTraceSession | gryviatracesessions | Namespaced | network-intelligence | none | captureHeaders, duration, filters, level, namespace, service |
GryviaTrafficInsight | gryviatrafficinsights | Namespaced | network-intelligence | none | metrics, namespace, service, window |
GryviaTrainingInsight | gryviatraininginsights | Namespaced | network-intelligence | none | analysisWindow, metrics, targetJob |
GryviaTrainingProfiler | gryviatrainingprofilers | Namespaced | ai-operator | target | analysis, metrics, output |
GryviaTrainingTimeMachine | gryviatrainingtimemachines | Namespaced | ai-operator | sourceJob | forks, retention, timeline |
GryviaUsageRecord | gryviausagerecords | Namespaced | quota-operator | cost, final, gpuHours, gpus, job, jobUID, rate, start, tenant | currency, end, gpuType, sku |
GryviaWorkflow | gryviaworkflows | Namespaced | none | steps | parameters |
GryviaWorkspace | gryviaworkspaces | Namespaced | none | type | cpuLimit, cpuRequest, env, gpuCount, gpuType, idleTimeoutMinutes, image, maxLifetimeHours, … |