Concurrent Processing Memory
Chengpu Wang
Abstract
A theoretical memory that embeds limited, application-specific processing power and nearest-neighbor connectivity at every storage element is proposed. Such a memory performs parallel computation within itself to solve generic array problems, while remaining pin- and function-compatible with conventional random-access memory. The applicability of this in-memory, finest-grain, massive-SIMD approach is examined in detail through a family of increasingly capable devices---content movable, searchable, value-comparable, and computable memory. For an array of N items, the approach reduces the instruction-cycle count of universal operations (insertion, deletion, and match finding) to 1, of local operations (filtering and template matching) to the operation footprint, and of global operations (summation and finding minimum/maximum) to N. It eliminates most data-processing traffic on the system bus, yet remains general-purpose, easy to program, backward-compatible with existing bus-sharing architectures and operating systems, and practical to implement along a clear road map.
Create a lesson
Related papers
Replication-Aware Placement of Functions and Data in the Edge-Cloud Continuum
Dario d'Abate, Matteo Cenzato, Matteo Briscini et al.
Ermes: a Stateful Serverless Platform for the Edge-to-Cloud Continuum
Matteo Cenzato, Dario d'Abate, Arianna Dragoni et al.
Fluid Notarization: Verifiable Evolution of Concurrently Edited Structured Documents
Amos Brocco, Giuliano Gremlich, Roberto Guidi
Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It
Yipeng Liu, Yingqiang Zhang, Feifei Li et al.
A Distributed Computing Framework for Satellite Swarms
Ezra Fielding, Clement Demazure, Guthemberg Silvestre et al.
Vigil: Accountable Liveness against Selective Silence
Jiawei Cheng, Huiping Sun, Rui Zhou et al.