Skip to content
@Infini-AI-Lab

Infini-AI-Lab

Next Generation AI algorithms and systems

Popular repositories Loading

  1. Sequoia Sequoia Public

    scalable and robust tree-based speculative decoding algorithm

    Python 376 37

  2. TriForce TriForce Public

    [COLM 2024] TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

    Python 281 21

  3. MagicPIG MagicPIG Public

    [ICLR2025 Spotlight] MagicPIG: LSH Sampling for Efficient LLM Generation

    Python 256 21

  4. MagicDec MagicDec Public

    [ICLR2025] Breaking Throughput-Latency Trade-off for Long Sequences with Speculative Decoding

    Python 158 13

  5. MonarchRT MonarchRT Public

    Python 145 13

  6. UMbreLLa UMbreLLa Public

    LLM Inference on consumer devices

    Python 132 15

Repositories

Showing 10 of 44 repositories
  • vortex_torch Public

    Vortex: Programmable Sparse Attention for Agents as Algorithm Designers

    Infini-AI-Lab/vortex_torch's past year of commit activity
    Python 68 Apache-2.0 14 0 6 Updated Aug 22, 2026
  • flashrt-blog Public
    Infini-AI-Lab/flashrt-blog's past year of commit activity
    HTML 0 0 0 0 Updated Aug 10, 2026
  • aibattle Public
    Infini-AI-Lab/aibattle's past year of commit activity
    HTML 4 MIT 0 0 0 Updated Aug 5, 2026
  • astraflow Public

    Dataflow-Oriented Reinforcement Learning for (Multi-)Agentic LLMs

    Infini-AI-Lab/astraflow's past year of commit activity
    Python 103 Apache-2.0 17 0 3 Updated Jul 26, 2026
  • FlashRT Public
    Infini-AI-Lab/FlashRT's past year of commit activity
    Python 5 0 0 0 Updated Jul 23, 2026
  • Sparrow Public
    Infini-AI-Lab/Sparrow's past year of commit activity
    Python 17 2 0 0 Updated Jun 15, 2026
  • Infini-AI-Lab/sparrow_project_release's past year of commit activity
    HTML 1 0 0 0 Updated Jun 10, 2026
  • GRESO Public
    Infini-AI-Lab/GRESO's past year of commit activity
    Python 83 7 1 0 Updated Jun 8, 2026
  • Infini-AI-Lab/vortex_inference's past year of commit activity
    0 0 0 0 Updated Jun 4, 2026
  • vllm Public Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Infini-AI-Lab/vllm's past year of commit activity
    Python 0 Apache-2.0 22,317 0 0 Updated May 16, 2026