Skip to content
#

tensor-parallelism

Here are 79 public repositories matching this topic...

DeepSeek-V4-Flash DSpark speculative decoding on 2x DGX Spark (GB10/sm_121), 1M context — our fp8 recipe + independent reproduction and cross-build benchmarks of the NVFP4-KV build. Honest, apples-to-apples measurements.

  • Updated Jul 5, 2026
  • Python

Running large LLMs on pre-Ampere NVIDIA hardware — Tesla V100 (sm_70), RTX 2080 Ti (sm_75), CMP 170HX. Measured benchmarks, vLLM forks, and the hardware side: NVLink on SXM2 carrier boards, driver traps, cooling, used-kit acceptance.

  • Updated Aug 19, 2026
dual-radeon-vllm

Multi-GPU tensor-parallel vLLM on AMD Radeon RX 7900 XT / XTX / GRE (7900XT, 7900XTX, RX7900XT, gfx1100, RDNA3, ROCm): root cause and fix for the RCCL hostcall / PCIe atomics (AtomicOps) crash "NCCL error: unhandled cuda error" / "operation cannot be performed in the present state", Proxmox VFIO passthrough; LLM inference benchmarks on 13 machines

  • Updated Sep 11, 2026
  • Python

GPU Memory Calculator for LLM Training - Calculate GPU memory requirements for training Large Language Models with support for multiple training engines including PyTorch DDP, DeepSpeed ZeRO, Megatron-LM, and FSDP.

  • Updated Jan 26, 2026
  • Python

Add this topic to your repo

To associate your repository with the tensor-parallelism topic, visit your repo's landing page and select "manage topics."

Learn more