vLLM

A serious inference server for serving a model to many people at once. Not a chatbot — the thing a chatbot runs on.

At a glance

Where it runs
Self-hostedYou run it on your own machine. Nothing leaves it unless you send it.
Adult content
AllowedAdult content is permitted within the published rules.
Trains on your content
NoYour content is not used to train models.
Status
ActiveMaintained and working today.
Account needed
No
Licence
Apache-2.0
Cost
Free and open source.
What it keeps
Whatever the operator configures. The software itself keeps nothing.

What it takes to run

Graphics memory
16 GB VRAM, 24 GB recommended
System memory
16 GB, 64 GB recommended
Disk
about 50 GB
Runs on
Linux

Find what fits your machine on the tools page — pick your hardware from the first menu.

vLLM is what you use when a local model has to answer more than one conversation at a time. Its scheduling and memory handling give it throughput the single-user tools do not attempt, which is why hosted providers and self-hosted communities alike run it underneath.

It wants a proper GPU and unquantised or GPU-quantised weights. It is not the tool for a laptop.

Worth knowing before you start

What it does

Models it runs

Signs of life

Checked automatically. These are the only figures on this page a machine wrote, and they say when they were taken.

Last answered
Yes, 3 hours ago Its website responded when we asked.
Stars on GitHub
91,871 22,247 forks.
Last commit
4 hours ago2026-09-16 02:02 UTC
Latest release
v0.29.0

Categories

Something here wrong or out of date? Tell us — this page is only worth having if it is right.