Where Should Experience Live? Hierarchical Hebbian Memory for Continual Vision Transformers
Mohammed Yusuf Mujawar, Noorbakhsh Amiri Golilarz
Abstract
Vision Transformers provide strong visual representations but typically rely on slowly updated parameters, limiting their ability to organize newly acquired information across different memory timescales. This work proposes Hierarchical Hebbian Memory, a three-level memory architecture composed of rapid Working Memory, persistent Routed Episodic Memory, and slower Semantic Memory. A learned controller regulates memory contribution, read and write routing, plasticity, retention, and consolidation. A causal read-before-write lifecycle ensures that the current outcome cannot influence the prediction it supervises. The architecture is evaluated on Omniglot 5-way 1-shot recognition and CORe50 continual object recognition. With Swin-Tiny, the hierarchical model reaches 97.39\% accuracy on Omniglot and 95.37\% final accuracy on CORe50 when combined with experience replay. Learned multi-bank retrieval reaches 47.50\% delayed-association accuracy, compared with 24.17\% for a single persistent bank and 25.00\% without memory. After intervening distractors, Episodic Memory retains approximately 0.96 cosine similarity with stored associations, while Working Memory falls to approximately 0.05. These results show that Hebbian association and learned memory routing can jointly organize online visual experience across rapid, persistent, and consolidated memory timescales within Vision Transformers.
Create a lesson
Related papers
Instance-Guided Report Anchoring for Text-Free 3D Abnormality Segmentation in Chest CT
Zhenyu Bu, Haoyan Ding, Chushu Shen et al.
Unmasking Face Embeddings: Reading, Rendering and Naming with Foundation Models
Fizza Rubab, Yiying Tong, Arun Ross
SlideMix: Enhancing Whole Slide Image Analysis via Multimodal Shuffling
Chad Wong, Sicheng Chen, Tianyi Zhang et al.
FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos
Maya Moriya, Sigal Raab, Yael Vinker et al.
Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
Vida Adeli, Soroush Mehraban, Jacob Rommann et al.
TRUST: Threshold-Recalibrated Uncertainty-Safe Training for Certified Dismissal in Breast Cancer Screening
Parham Hajishafiezahramini, Matthew Hamilton, Edward Kendall et al.