E8-lattice codebook quantization for LLM weights — 2–8 bpw, mixed-precision, fused CUDA tensor-core inference on vLLM + HF
-
Updated
Aug 12, 2026 - Python
E8-lattice codebook quantization for LLM weights — 2–8 bpw, mixed-precision, fused CUDA tensor-core inference on vLLM + HF
💎 LTS Industrial Standard: PDF/Word optimization with scientific Trellis Mimic engine. Features Turbo Parallel processing & Global Camouflage (Ricoh/Fujitsu/Canon 2025 profiles). Embedded Python 3.12, zero-install, no admin needed. Ultimate document privacy for Windows LTSC/Enterprise.
Add a description, image, and links to the trellis-quantization topic page so that developers can more easily learn about it.
To associate your repository with the trellis-quantization topic, visit your repo's landing page and select "manage topics."