This paper studies why coordinate-based multilayer perceptrons struggle to fit high-frequency signals in low-dimensional domains. It analyzes training through neural tangent kernels and shows how sinusoidal Fourier features produce a stationary effective kernel whose bandwidth can be adjusted. Experiments examine convergence, generalization, feature sampling distributions, network depth, joint feature optimization, translation sensitivity, and directional bias. Direct image and shape regression and indirectly supervised CT, MRI, and simplified NeRF reconstruction compare unembedded inputs with basic, positional, and Gaussian mappings. Gaussian features provide the strongest reported results among the main mappings, while bandwidth selection balances underfitting and overfitting. Appendix studies identify limitations of feature optimization and axis-aligned encodings.
A small change to how coordinates enter a neural network can help it represent sharper images and finer three-dimensional detail. The study explains why feature frequency scale matters and shows benefits in both direct fitting and reconstruction from indirect measurements. It also warns that excessive frequencies can overfit, learning the feature parameters jointly need not help, and axis-aligned features can favor aligned scenes. These are results for the paper's particular signals, architectures, and synthetic measurement settings, not a general guarantee for all neural networks or clinical imaging.
论文中的结论可在「研究结论」中查看。