vLLM
A serious inference server for serving a model to many people at once. Not a chatbot — the thing a chatbot runs on.
- Self-hosted
- Active
- Allowed adult content
- Open source
At a glance
- Where it runs
- Self-hostedYou run it on your own machine. Nothing leaves it unless you send it.
- Adult content
- AllowedAdult content is permitted within the published rules.
- Trains on your content
- NoYour content is not used to train models.
- Status
- ActiveMaintained and working today.
- Account needed
- No
- Licence
- Apache-2.0
- Cost
- Free and open source.
- What it keeps
- Whatever the operator configures. The software itself keeps nothing.
What it takes to run
- Graphics memory
- 16 GB VRAM, 24 GB recommended
- System memory
- 16 GB, 64 GB recommended
- Disk
- about 50 GB
- Runs on
- Linux
Find what fits your machine on the tools page — pick your hardware from the first menu.
vLLM is what you use when a local model has to answer more than one conversation at a time. Its scheduling and memory handling give it throughput the single-user tools do not attempt, which is why hosted providers and self-hosted communities alike run it underneath.
It wants a proper GPU and unquantised or GPU-quantised weights. It is not the tool for a laptop.
Worth knowing before you start
- Linux and NVIDIA in practice Other platforms are supported to varying degrees but are not the well-trodden path.
- No interface of its own It serves an API; you bring the front-end.
What it does
- High-throughput batched serving
- OpenAI-compatible API
- Tensor parallelism across GPUs
Models it runs
Signs of life
Checked automatically. These are the only figures on this page a machine wrote, and they say when they were taken.
- Last answered
- Yes, 3 hours ago Its website responded when we asked.
- Stars on GitHub
- 91,871 22,247 forks.
- Last commit
- 4 hours ago2026-09-16 02:02 UTC
- Latest release
- v0.29.0
Categories
Something here wrong or out of date? Tell us — this page is only worth having if it is right.