vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

Archived This repository is archived: read-only and no longer maintained.

Fork A fork of vllm-project/vllm; the README may describe the upstream project.

Overview

Easy, fast, and cheap LLM serving for everyone

Ray Summit CPF is Open (June 4th to June 20th)!

There will be a track for vLLM at the Ray Summit (09/30-10/02, SF) this year! If you have cool projects related to vLLM or LLM inference, we would love to see your proposals. This will be a great chance for everyone in the community to get together and learn. Please submit your proposal here

vLLM is a fast and easy-to-use library for LLM inference and serving.

vLLM is flexible and easy to use with:

vLLM seamlessly supports most popular open-source models on HuggingFace, including:

Find the full list of supported models here.

Install vLLM with pip or from source:

Visit our documentation to learn more.

We welcome and value any contributions and collaborations. Please check out CONTRIBUTING.md for how to get involved.

If you use vLLM for your research, please cite our paper:

From the project’s README on GitHub.

At a glance

RepositoryKentucky-Open-Science/vllm
Research areaForks of other projects
Primary languagePython
LanguagesPython 82%, Cuda 13.4%, C++ 3.3%, Shell 0.7%, CMake 0.5%, Dockerfile 0.1%
LicenseApache-2.0
Stars / forks0 / 0
Open issues and pull requests7
Created2024-07-01
Last push2025-12-09
Default branchmain
Homepagehttps://docs.vllm.ai
Forked fromvllm-project/vllm — A high-throughput and memory-efficient inference and serving engine for LLMs

What the README covers

Top contributors

Get the code

git clone https://github.com/Kentucky-Open-Science/vllm.git