Skip to content
#

gpu-profiling

Here are 50 public repositories matching this topic...

Cross-platform .NET performance engineering skill for coding agents, covering CPU, memory, GC, benchmarking, concurrency, startup, native profiling, GPU rendering, and production diagnostics on macOS, Windows, and Linux.

  • Updated Jul 28, 2026
  • Python

Profine automatically profiles and optimizes PyTorch training jobs on real GPUs, delivering measurable speedups and lower GPU costs before teams waste days tuning configs by hand.

  • Updated May 20, 2026
  • Python

Communication cost modeling for tensor parallel LLM inference with TP vs PP vs hybrid comparison, VRAM analysis, pipeline bubble modeling, regime detection, and cost-efficiency. Shows TP dominates on NVLink, PP has 47% bubble at 8 GPUs, and LLaMA-70B needs 8× A100 or 2× H100 for VRAM.

  • Updated Jul 15, 2026
  • Python

Unified benchmarking and profiling framework for the JAX scientific ML ecosystem. Timing, GPU/energy monitoring, FLOPS counting, roofline analysis, statistical testing, regression detection, and CI integration.

  • Updated Sep 9, 2026
  • Python

Add this topic to your repo

To associate your repository with the gpu-profiling topic, visit your repo's landing page and select "manage topics."

Learn more