Back to the interactive map
Inference Workload

RAG

Retrieval Augmented Generation — embed a query, retrieve relevant chunks from a vector store, then generate a grounded answer.