Back to the interactive map
Inference Frameworks

vLLM

PagedAttention-based LLM serving engine with continuous batching, speculative decoding, and an OpenAI-compatible API