Inference Stack
The reference for how AI infrastructure actually fits together
Architectures
Catalog
Where it runs
About
Back to the interactive map
Orchestration
Kubernetes + NIM
nvidia
NVIDIA Inference Microservices deployed as containerized K8s workloads
Official docs
Used by
vLLM
TensorRT-LLM
ONNX Runtime
HF Diffusers
Leads to
CUDA