vLLM vs llama.cpp vs Text Generation Inference
A side-by-side look at how actively vLLM, llama.cpp and Text Generation Inference are developed on GitHub — popularity, contributor base, commit activity and how quickly issues and pull requests move.
What the numbers say
- llama.cpp is the most-starred of the 3, with 130k stars — about 11.9× Text Generation Inference's 11k.
- Over the 52 weeks to the snapshot, vLLM recorded the most commits (12k), against 8 for Text Generation Inference.
- vLLM has closed or merged the most pull requests in its lifetime: 34k.
- Relative to its popularity, llama.cpp carries the lightest open-issue load: 6.7 open issues per 1,000 stars, versus 27.4 for vLLM.
- Text Generation Inference had gone 193 days without a push at the time of the snapshot.
- Text Generation Inference is archived on GitHub — read-only, no longer accepting changes.
Head-to-head
Snapshot taken 30 September 2026.
| Metric | vLLM Active | llama.cpp Active | Text Generation Inference Archived |
|---|---|---|---|
| Stars | 93k | 130k | 11k |
| Forks | 23k | 24k | 1.3k |
| Contributors | 3.6k | 2.1k | 148 |
| Commits, last 52 weeks | 12k | 4.6k | 8 |
| Open issues | 2.5k | 871 | 285 |
| Open pull requests | 5.8k | 1.7k | 39 |
| Closed / merged PRs | 34k | 14k | 1.7k |
| Watchers | 598 | 841 | 103 |
| Last push | 30 Sep 2026 | 30 Sep 2026 | 21 Mar 2026 |
| Latest release | v0.30.0 (Sep 2026) | v0.5.0 (Sep 2026) | v3.3.7 (Dec 2025) |
| Main language | Python | C++ | Python |
| Licence | Apache-2.0 | MIT | Apache-2.0 |
| Created | Feb 2023 | Mar 2023 | Oct 2022 |
Bold marks the highest value in a row. “Active / Slowing / Dormant” is a simple heuristic from how recently code was pushed and open issues per star — a starting point, not a verdict. Stars measure popularity, not quality.
Open these in the live comparer →