RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention
Anthony. Lui, Mohamed. Elsaied, N. P. Savani
Abstract
Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneous pressures: resident weight matrices, key-value (KV) cache state that grows linearly with context, and dozens of expert sublayers that must be paged on demand. We present RotaryQuant, a three-axis compression system that addresses all three. Mixed-precision weight quantization assigns bit-widths by architectural role: 4-bit for dense layers, 2-bit for routed experts, and 8-bit for the shared expert whose high activation kurtosis resists aggressive compression. LRU expert offloading pages non-resident experts to disk under genuine memory pressure. The novel axis is IsoQuant, a KV cache compression method that applies a Walsh--Hadamard transform followed by block-diagonal SO(4) rotations to isotropize activation distributions before 3-bit scalar quantization, requiring O(d d) operations and 256 stored parameters per head versus O(d2) and 16,384 for dense rotation methods. A fused four-kernel Metal GPU pipeline performs attention directly on packed 3-bit tensors without materializing full-precision KV state---a different execution model, not just a quantization scheme. The combined system fits Gemma 4-26B-A4B and Qwen3-30B-A3B within a 16\,GB budget and Nemotron-H 120B within 32\,GB, running interactively at 9--19 tok/s with near-zero perplexity degradation (ΔPPL ≤ +0.0012) and 100\% retrieval accuracy at 32K context.
Create a lesson
Related papers
A Metaheuristic Optimization Framework for Discrete Optimization under Strict Time Limits
Umut Çalıkyılmaz, Nitin Nayak, Sven Groppe
Benchmarking Tabular Foundation Models as Surrogates in Expensive Evolutionary Optimization
Lu Han, Jin Wang, Yuchen Li et al.
A Spatiotemporal Extension of the Neuromorphic DBSCAN Implementation
Charles P. Rizzo, James S. Plank
Machine Zygote: Causal Biparental Heredity Before Learning in a Germline--Soma Artificial Agent
Lyes Saad Saoud
Bio-Inspired Palette Evolution in Indirectly Encoded Substrates: Timescale Compatibility Shapes Activation Function Discovery
Romain Claret, Michael O'Neill, Paul Cotofrei et al.
LLMDE: A Large Language Model-Driven Differential Evolution Algorithm for Portfolio Optimization
Rong Chai, Vaclav Snasel, Xiaopeng Wang et al.