We read every piece of feedback, and take your input very seriously.
To see all available qualifiers, see our documentation.
There was an error while loading. Please reload this page.
A high-throughput and memory-efficient inference and serving engine for LLMs
Python 88.1k 20.2k
A framework for efficient model inference with omni-modality models
Python 5.8k 1.4k
Common recipes to run vLLM
JavaScript 947 354
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Python 3.6k 601
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
Python 688 177
A programmable Mixture-of-Models router for heterogeneous LLM inference
Go 5.1k 797
Community maintained hardware plugin for vLLM on Intel Gaudi
Community maintained hardware plugin for vLLM on Ascend
Cost-efficient and pluggable Infrastructure components for GenAI inference
TPU inference for vLLM, with unified JAX and PyTorch support.
vLLM plugin for attention-ffn disaggregation support
Community maintained hardware plugin for vLLM on Apple Silicon
An LLM post-training framework with vLLM for RL Scaling
Loading…