Skip to content

Concurrent Processing Memory

Chengpu Wang

cs.DCarXiv:cs/0608061

Abstract

A theoretical memory that embeds limited, application-specific processing power and nearest-neighbor connectivity at every storage element is proposed. Such a memory performs parallel computation within itself to solve generic array problems, while remaining pin- and function-compatible with conventional random-access memory. The applicability of this in-memory, finest-grain, massive-SIMD approach is examined in detail through a family of increasingly capable devices---content movable, searchable, value-comparable, and computable memory. For an array of N items, the approach reduces the instruction-cycle count of universal operations (insertion, deletion, and match finding) to 1, of local operations (filtering and template matching) to the operation footprint, and of global operations (summation and finding minimum/maximum) to N. It eliminates most data-processing traffic on the system bus, yet remains general-purpose, easy to program, backward-compatible with existing bus-sharing architectures and operating systems, and practical to implement along a clear road map.

Create a lesson