Inference Stack
The reference for how AI infrastructure actually fits together
Architectures
Catalog
Where it runs
About
Back to the interactive map
Hardware Abstraction
CUDA
nvidia
NVIDIA's parallel computing platform and programming model, CUDA 12.x
Official docs
Used by
Triton Inference Server
Ray Serve
Kubernetes + NIM
SGLang Runtime
Leads to
H100
B200
L4
A10G