Skip to main content

AI Workloads at a glance

The short version; the full reference is ai-workloads.md and the walkthrough is Tutorial 15. Back to the README.

AI Workloads (Beta)​

OpenAI-compatible inference on the same daemon — no separate AI control plane.

MaturitySingle-cluster Beta · multi-site HA store stays Preview · not GA
Console/app/ai — Models, Deployments, Endpoints, API keys, Nodes
CLIfabricctl ai model | profile | deploy | endpoint | key | gpus | node | capacity
Gateway/api/ai/openai/{endpoint}/v1/chat/completions
Lab without NVIDIASet FLUXVM_AI_JANUS_URL — Zyvor Janus is the virtual upstream
Real GPUsFluxVM inventory + VFIO VM + runtime image (FLUXVM_AI_IMAGE)
Runtimesvllm, tensorrt-llm, triton, llama.cpp, tei by default · deny/allow via env
MIGJanus records always · PCI via FLUXVM_AI_PCI_MIG=1
fabricctl ai model add demo-qwen --source hf://Qwen/Qwen3-8B
fabricctl ai profile add demo-24g --runtime vllm --gpu 1 --vram 24 --cpu 8 --memory 32
fabricctl ai deploy demo-qwen --profile demo-24g --replicas 1
fabricctl ai endpoint expose demo-qwen --openai-compatible
fabricctl ai key create demo-key --endpoint demo-qwen-openai
# → POST $FABRIC_URL/api/ai/openai/demo-qwen-openai/v1/chat/completions

Guides: Tutorial 15 — how to use · Full reference · Website tutorial · Lab smoke: ./scripts/smoke-ai-janus-lab.sh