Author
Pier-Jean Malandrino
Recent research
- AI & ComputingOpen access
Unfolding the Leech Lattice: Fused Multi-Shell Decoding and VRAM Layouts for 2-Bit LLM Weights
Leech lattice vector quantization gives the best reported quality at 2 bits per weight, but the CUDA kernel published alongside it decodes only one shell of the lattice. The decoder you actually need at that rate covers a union of shells: 301 classes, and a 47-bit index that name...