国际机器学习大会 · 2 篇论文
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
ganwumeng/transformers-are-ssms-generalized-models-and-efficient-algorithms-through-struct--rs9v6u
The paper develops structured state space duality (SSD), a framework relating structured state space models to several forms of attention through semiseparable matrices. It gives equivalent matrix and tensor views, derives a block-decomposed SSD algorithm, and uses the framework to design Mamba-2. The paper reports experiments on multi-query associative recall, language-model scaling, downstream zero-shot tasks, hybrid SSD/attention/MLP architectures, speed, and architectural ablations. Mamba-2 is reported to improve over Mamba-1 on recall and language benchmarks, while SSD is faster than the compared fused scan and becomes faster than FlashAttention-2 at sufficiently long sequences. Results also show benefits from larger states and selected attention mixtures.
Simplifying Graph Convolutional Networks
ganwumeng/simplifying-graph-convolutional-networks--jl607e
This paper derives Simple Graph Convolution (SGC) by removing intermediate nonlinearities from a graph convolutional network and collapsing its weight matrices into one linear classifier. The resulting method applies a fixed, parameter-free graph propagation filter to node features before multinomial logistic regression. A spectral analysis relates propagation with self-loops to low-pass filtering and proves that self-loops shrink the normalized-Laplacian spectrum. Experiments compare SGC with graph neural-network baselines on citation and social networks and adapt it to text classification, geolocation, relation extraction, zero-shot image classification, graph classification, and molecular prediction, emphasizing accuracy, training time, stability, and known failure cases.