Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
-
Updated
Oct 3, 2026 - Python
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
A transparent, in-container GPU resource controller that enforces memory and compute limits by intercepting CUDA calls without application or driver changes.
Easily manage virtual GPU resource pools in Kubernetes clusters.
SLURM-native software GPU slicing for NVIDIA clusters using memory limits and compute time-slicing.
Boosting GPU utilization for LLM serving via dynamic spatial-temporal prefill & decode orchestration
Node-level enforcement and metrics for fractional GPU allocations made by KAI Scheduler.
Distributed, peer-to-peer, decentralized network for LLM inference — peers seed & leech completions, metered by BitTorrent-protocol-style upload/download ratios. OpenAI-compatible · Rust · libp2p.
Hands-on GPU/HPC infrastructure operations: K8s GPU scheduling, HAMi sharing, Slurm, observability & vLLM inference. Learn it free on a laptop; validate on one cheap GPU.
TensorFusion landing page and product docs
GPU memory isolation for KAI-Scheduler GPU-sharing workloads, powered by HAMi-core.
Native Steam Gaming Mode + Sunshine in a Proxmox VE LXC. Use an AMD iGPU or dGPU for headless gaming without dedicating it to a VM, while keeping the GPU available for other workloads. If this helped you, consider giving the repo a ⭐.
A zero-config, decentralized local compute mesh. Pool RAM, GPU VRAM, and storage across macOS, Linux, Windows, and Raspberry Pis to run distributed AI models and sandboxed tasks locally-with zero cloud billing.
RunSnack: one link, instant terminal on your GPU‑ready Docker host. Sandboxed, read‑only root, dropped caps, no accounts, no cloud. Peer‑to‑peer, ephemeral, free. Share a shell in seconds — not minutes. No tracking, no logs, just your hardware, your session.
Tricks for GPU sharing pods on Kubernetes without any use of middleware like HAMi or DRA
Containerized System Architecture for Robotics — companion repository for the CSAR paper
Private P2P GPU sharing for trusted circles — WireGuard tunnels, Firecracker microVMs, local credits. No blockchain. MIT.
Self-hosted, multi-user AI filmmaking workspace with GPU reservations, Codex/Claude tmux sessions, private models, and MiniMax H3 workflows.
GPUs unite using secure and private crypto transactions to distribute compute to decentralized nodes.
Share a GPU, borrow a GPU. Contribute spare capacity so anyone can run their local AI on it — free, no account, no weights to download.
Distributed peer-to-peer LLM inference network. Volunteer your GPU, earn AI credits, run any open-source model for free. Anonymous, encrypted, unstoppable.
To associate your repository with the gpu-sharing topic, visit your repo's landing page and select "manage topics."