vLLM
VerifiedOpen sourceHigh-throughput LLM serving engine
Best for github · inference · serving
4.8
Toolora score — editorial, not public reviewsToolora score94/100
Strengths
- PagedAttention memory management
- OpenAI-compatible server
Skip if
you want hosted support, not a repo to run
vLLM is an open-source library for fast LLM inference and serving. PagedAttention and continuous batching make it a default choice for production model APIs.
Key features
- PagedAttention memory management
- OpenAI-compatible server
- Tensor and pipeline parallelism
- Wide model support
Why it is on Toolora
We list vLLM for people who need github · inference · serving. Skip it if you want hosted support, not a repo to run.
Closest alternatives on Toolora: llama.cpp, Ollama, Groq.
Try instead
Comparellama.cpp
★ 92kCommunityRun LLMs locally with pure C/C++
AI Tools4.9
Free / open source
Open-source option with 92k GitHub stars.
Open-source option with 165k GitHub stars.
Also built for inference.
Featured in guides
vLLM
Free / open source