High-frequency Multispeculative Multiply-Accumulation Unit for Fused Posit Arithmetic
Mario Alonso, Miguel Ángel Sacristán, Guillermo Botella, Alberto A. Del Barrio
Abstract
Posit arithmetic offers a compelling alternative to the IEEE 754 floating-point standard, providing enhanced accuracy. Its fused multiply-accumulate operations avoid intermediate rounding, ensuring exact numerical reproducibility through the quire, a wide fixed-point accumulator spanning the format's full dynamic range to prevent precision loss and overflow during long accumulations. However, integrating such large accumulators incurs significant area and power overheads. This paper presents an optimized, high-frequency Multispeculative PositMAC architecture for 32- and 64-bits Posit. First, the pipeline is restructured to balance the different stages. Second, high-speed multiplication topologies are evaluated, showing that a Booth-4 scheme with Kogge-Stone adders meets a stringent 0.5ns target (2Ghz). Finally, the wide monolithic quire accumulator is replaced with a Multispeculative Adder, diminishing area up to 19.8\% while reducing energy consumption by more than 50\% when compared to the baseline. Compared to other state-of-the-art designs, our proposal achieves the highest operating frequency and reduces cycle time by up to 79.0\% with respect to 64-bit quire-enabled alternatives. This performance is attained without increasing resource overhead, as the design remains strictly smaller in area and achieves lower per-cycle energy consumption than all quire-capable counterparts.
Create a lesson
Related papers
Evaluating Positive Feedback Adiabatic Logic in 16nm FinFET with a Realistic Power-Clock
Franciszek Łukowski, Maciej Pyrzowski, Aida Todri-Sanial
Evaluation of Power-Clock Waveforms for Positive Feedback Adiabatic Logic in 16 nm FinFET Technology
Maciej Szymon Pyrzowski, Franciszek Łukowski, Aida Todri-Sanial
MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration
Yuan Liao, Jae-sun Seo
MeshKV: A Network-on-Chip KV Cache Fabric for Scalable Transformer Decoding Accelerators
Dong Liu, Yanxuan Yu
Epic: Efficient Programming Paradigm for In-Storage Computing
Yuyue Wang, Zhenyu Zhang, Glenn Reinman et al.
Locus: A Framework for Exploring and Optimizing Point Addition Hardware for Zero-Knowledge Proofs
Gaurav Kuwar, Alhad Daftardar, Jianqiao Mo et al.