@ganwumeng
8 个研究仓库 · 0 关注者
—加入
Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains
This paper studies why coordinate-based multilayer perceptrons struggle to fit high-frequency signals in low-dimensional domains. It analyzes training through neural tangent kernels and shows how sinusoidal Fourier features produce a stationary effective kernel whose bandwidth can be adjusted. Experiments examine convergence, generalization, feature sampling distributions, network depth, joint feature optimization, translation sensitivity, and directional bias. Direct image and shape regression and indirectly supervised CT, MRI, and simplified NeRF reconstruction compare unembedded inputs with basic, positional, and Gaussian mappings. Gaussian features provide the strongest reported results among the main mappings, while bandwidth selection balances underfitting and overfitting. Appendix studies identify limitations of feature optimization and axis-aligned encodings.
Simplifying Graph Convolutional Networks
This paper derives Simple Graph Convolution (SGC) by removing intermediate nonlinearities from a graph convolutional network and collapsing its weight matrices into one linear classifier. The resulting method applies a fixed, parameter-free graph propagation filter to node features before multinomial logistic regression. A spectral analysis relates propagation with self-loops to low-pass filtering and proves that self-loops shrink the normalized-Laplacian spectrum. Experiments compare SGC with graph neural-network baselines on citation and social networks and adapt it to text classification, geolocation, relation extraction, zero-shot image classification, graph classification, and molecular prediction, emphasizing accuracy, training time, stability, and known failure cases.
Random Features for Large-Scale Kernel Machines
Rahimi and Recht introduce randomized, explicit low-dimensional feature maps whose Euclidean inner products approximate shift-invariant kernels, allowing nonlinear kernel methods to be replaced by fast linear learning. They develop random Fourier features from a kernel's spectral distribution and random binning features from randomly shifted grids, and give uniform approximation bounds for both constructions. Using ridge regression on five large-scale regression and classification datasets, they compare the two feature families with Core Vector Machines and published exact-kernel baselines. Reported results show competitive test error with substantial, though dataset-dependent, training-time advantages, while also revealing differences between interpolation-oriented Fourier features and locality-preserving binning features.
Auto-Encoding Variational Bayes
This paper develops stochastic-gradient variational inference for directed probabilistic models with continuous latent variables and intractable posteriors. Its reparameterization expresses posterior samples as differentiable transformations of parameter-free noise, yielding the Stochastic Gradient Variational Bayes estimator. For independent observations with local latent variables, the Auto-Encoding Variational Bayes algorithm jointly trains a probabilistic encoder and generative decoder, producing the variational auto-encoder when both are neural networks. Experiments on MNIST and Frey Face compare AEVB with wake-sleep using variational lower bounds, and on low-dimensional MNIST also compare estimated marginal likelihood with Monte Carlo EM. Appendix visualizations show learned manifolds and generated samples.
Deep Residual Learning for Image Recognition
This paper introduces residual learning for training substantially deeper neural networks. Instead of directly fitting a desired mapping, stacked layers learn a residual that is added to an identity shortcut. The authors compare plain and residual networks on ImageNet and CIFAR-10, examine shortcut variants and residual-response magnitudes, and scale residual networks to 152 layers on ImageNet and 1202 layers on CIFAR-10. Reported results show reduced optimization degradation and improved classification accuracy with depth. The learned representations also improve Faster R-CNN detection on PASCAL VOC and COCO and support competitive ImageNet detection and localization systems.
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
FlashAttention-2 is an exact GPU attention algorithm and implementation that improves on FlashAttention by reducing non-matrix-multiplication work, parallelizing computation across sequence blocks, and repartitioning work among warps to reduce shared-memory communication. The paper benchmarks forward, backward, and combined attention on A100 GPUs across causal and non-causal settings, two head dimensions, and sequence lengths from 512 to 16k, and also reports unoptimized H100 results. It evaluates end-to-end GPT-style training for 1.3B- and 2.7B-parameter models at 2k and 8k context. Reported gains reach roughly twofold over FlashAttention and 225 TFLOPs/s/GPU in training.
Egalitarian Gradient Descent: A Simple Approach to Accelerated Grokking
This paper attributes delayed generalization in grokking partly to unequal optimization speeds along gradient singular directions. It introduces Egalitarian Gradient Descent (EGD), which replaces each selected layer's gradient by its polar factor so that nonzero singular values are equalized. A solvable anisotropic linear-classification model links covariance conditioning and initialization scale to plateau length, while experiments compare EGD, randomized-SVD approximations, column normalization, standard optimizers, and Grokfast. Across modular arithmetic, sparse parity, MNIST, and CIFAR-10 distribution-shift settings, the paper reports earlier generalization and improved adaptability, alongside computation, rank-sensitivity, instability, and trajectory analyses.
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
The paper develops structured state space duality (SSD), a framework relating structured state space models to several forms of attention through semiseparable matrices. It gives equivalent matrix and tensor views, derives a block-decomposed SSD algorithm, and uses the framework to design Mamba-2. The paper reports experiments on multi-query associative recall, language-model scaling, downstream zero-shot tasks, hybrid SSD/attention/MLP architectures, speed, and architectural ablations. Mamba-2 is reported to improve over Mamba-1 on recall and language benchmarks, while SSD is faster than the compared fused scan and becomes faster than FlashAttention-2 at sufficiently long sequences. Results also show benefits from larger states and selected attention mixtures.