Inference Stack
The reference for how AI infrastructure actually fits together
Architectures
Catalog
Where it runs
About
Back to the interactive map
Orchestration
Triton Inference Server
nvidia
Production-grade model server with dynamic batching and multi-model support
Official docs
Used by
vLLM
TensorRT-LLM
ONNX Runtime
Faster Whisper
Leads to
CUDA
ROCm
Intel oneAPI
AWS Neuron SDK