Branch Scaling Manifests as Implicit Architectural Regularization for Improving Generalization in Overparameterized ResNets
cs.LG, math.ST, stat.TH
Submitted: 2024-03-07
Updated: 2026-08-30
Project page: https://yao-lab.github.io/publications/earlystop.pdf
Terminology
Sources
- Language Models are Few-Shot Learners
- Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA
- On the Impacts of the Random Initialization in the Neural Tangent Kernel Theory
- On Lazy Training in Differentiable Programming
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? -- A Neural Tangent Kernel Perspective
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks
- RAM: Recover Any 3D Human Motion in-the-Wild
- Generalization Ability of Wide Neural Networks on $\mathbb{R}$
- Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning
- Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning
- Proximal Policy Optimization Algorithms
- Divergence of Empirical Neural Tangent Kernel in Classification Problems
- Going deeper with Image Transformers
- Transformers without Normalization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks