Pooling Helps, Learned Weighting Hurts In-Context: Decomposing Group Attention
Michael Fore, James Mason Inder, Mrishika Nair, Praneetha Vaddamanu, Sharlina Keshava
Abstract
Group attention, introduced by the time series forecasting model Chronos-2, attends over the variates of a group at a fixed patch index and serves both multivariate (MV) and in-context learning (ICL) forecasting. Rather than evaluating this cross-variate attention design as a whole, we ask which part of the mechanism earns the benefit and probe its applicability to both MV and ICL regimes. By editing the attention matrix α at inference we separate the two pathways a head comprises: V/O, which projects a weighted summary of the group, and Q/K, which decides the weights. Uniform pooling (V/O without any Q/K weighting) is positive on 18 of our 20 sensor-network configurations, while the learned weighting (Q/K) splits by group type: its contribution is positive or negligible for MV, but materially degrades 8 of the 10 sensor-network ICL configurations, leaving 4 of them worse than univariate inference. By isolating the impact of different layers, we find that uniforming α in the first block alone improves every ICL configuration we test.
Create a lesson
Related papers
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
Jichao Jiang, Cristian McGee, El Houcine Bergou et al.
FERPO: Forward Entropy-Regularized Policy Optimization
Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv
Cost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control
Akshay Balsubramani
The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
Shuo Xing, Zilin Dai, Chengyuan Qian et al.
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
Cristian McGee, El Houcine Bergou, Aritra Dutta
Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
Jason X. Liu, Sebastian Ibarraran, Frank Hu et al.