MeshKV: A Network-on-Chip KV Cache Fabric for Scalable Transformer Decoding Accelerators
Dong Liu, Yanxuan Yu
Abstract
Autoregressive transformer decoding is constrained by irregular key-value (KV) cache movement on tiled accelerators. Prior compression and DRAM-placement systems still concentrate traffic on centralized memory paths that bottleneck long-context serving. We present MeshKV, a KV cache fabric that moves blocks as packetized flows over a lightweight NoC. It co-designs (i) TaKV affine striping to spread homes and cut hotspot load, (ii) Mare multicast with verified duplicate suppression, and (iii) Pad, which overlaps prefetch, tile multiply, and streaming softmax behind credit-aligned FIFOs. Together they convert bisection back-pressure into useful KV transfer. On our 8x8 FPGA implementation with LLaMA-2-7B and Mistral-7B at 8K-32K, MeshKV reduces interconnect traffic by up to 58%, improves KV bandwidth utilization by 2.1x, and delivers up to 1.9x multi-stream throughput.
Create a lesson
Related papers
Evaluating Positive Feedback Adiabatic Logic in 16nm FinFET with a Realistic Power-Clock
Franciszek Łukowski, Maciej Pyrzowski, Aida Todri-Sanial
Evaluation of Power-Clock Waveforms for Positive Feedback Adiabatic Logic in 16 nm FinFET Technology
Maciej Szymon Pyrzowski, Franciszek Łukowski, Aida Todri-Sanial
High-frequency Multispeculative Multiply-Accumulation Unit for Fused Posit Arithmetic
Mario Alonso, Miguel Ángel Sacristán, Guillermo Botella et al.
MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration
Yuan Liao, Jae-sun Seo
Epic: Efficient Programming Paradigm for In-Storage Computing
Yuyue Wang, Zhenyu Zhang, Glenn Reinman et al.
Locus: A Framework for Exploring and Optimizing Point Addition Hardware for Zero-Knowledge Proofs
Gaurav Kuwar, Alhad Daftardar, Jianqiao Mo et al.