A type of generalization error induced by initialization in deep neural networks
Yaoyu Zhang, Zhi-Qin John Xu, Tao Luo, Zheng Ma
Introduction
The wide application of deep learning makes it increasingly urgent to establish quantitative theoretical understanding of the learning and generalization behaviors of deep neural networks (DNNs). In this work, we study theoretically the problem of how initialization and loss function quantitatively affect these behaviors of DNNs. Our study focuses on the regression problem, which plays a key role in many applications, e.g., simulation of physical systems