Skip to content

Pinned Loading

  1. vllm vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 91.6k 22.1k

  2. vllm-omni vllm-omni Public

    A framework for efficient model inference with omni-modality models

    Python 6.8k 1.7k

  3. recipes recipes Public

    Common recipes to run vLLM

    JavaScript 1k 414

  4. llm-compressor llm-compressor Public

    Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM

    Python 3.8k 660

  5. speculators speculators Public

    A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM

    Python 826 218

  6. semantic-router semantic-router Public

    A programmable Mixture-of-Models router for heterogeneous LLM inference

    Go 5.7k 936

Repositories

Showing 10 of 49 repositories
  • semantic-router Public

    A programmable Mixture-of-Models router for heterogeneous LLM inference

    vllm-project/semantic-router's past year of commit activity
    Go 5,747 Apache-2.0 936 372 (4 issues need help) 161 Updated Sep 12, 2026
  • ci-infra Public

    This repo hosts code for vLLM CI & Performance Benchmark infrastructure.

    vllm-project/ci-infra's past year of commit activity
    Python 54 Apache-2.0 83 0 72 Updated Sep 12, 2026
  • tpu-inference Public

    TPU inference for vLLM, with unified JAX and PyTorch support.

    vllm-project/tpu-inference's past year of commit activity
    Python 430 Apache-2.0 310 93 (2 issues need help) 397 Updated Sep 12, 2026
  • vllm-project/vllm-dashboard's past year of commit activity
    TypeScript 13 13 2 12 Updated Sep 12, 2026
  • vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    vllm-project/vllm's past year of commit activity
    Python 91,556 Apache-2.0 22,095 2,396 (32 issues need help) 5,000+ Updated Sep 12, 2026
  • vllm-ascend Public

    Community maintained hardware plugin for vLLM on Ascend

    vllm-project/vllm-ascend's past year of commit activity
    C++ 2,805 Apache-2.0 2,237 1,432 (77 issues need help) 1,753 Updated Sep 12, 2026
  • DeepSelect Public Forked from deepseek-ai/DeepSelect

    DeepSelect: TopK kernels for DeepSeek Sparse Attention (DSA) and Samplers

    vllm-project/DeepSelect's past year of commit activity
    Cuda 0 MIT 16 0 1 Updated Sep 12, 2026
  • vime Public

    An LLM post-training framework with vLLM for RL Scaling

    vllm-project/vime's past year of commit activity
    Python 458 Apache-2.0 91 39 27 Updated Sep 12, 2026
  • llm-compressor Public

    Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM

    vllm-project/llm-compressor's past year of commit activity
    Python 3,776 Apache-2.0 660 42 (4 issues need help) 74 Updated Sep 12, 2026
  • DeepGEMM Public Forked from deepseek-ai/DeepGEMM

    DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling

    vllm-project/DeepGEMM's past year of commit activity
    Cuda 1 MIT 1,258 0 0 Updated Sep 12, 2026