DNN Training Stages Understanding
Recent works show that DNN training undergoes different stages, each showing different effects depending on the hyperparameter setting, which therefore warrants detailed explanation. Below, I aim to analyze and share a deep understanding of DNN training, especially from the following three perspectives:
- On the optimization and generalization perspective
- On the frequency domain perspective
- What happens during the early phase of DNN training
On the Optimization and Generalization Perspective
The connection between optimization and generalization of deep neural networks (DNN) is not fully understood. For instance, using a large initial learning rate often improves generalization, which can come at the expense of the initial training loss reduction. In this context, four works aimed at understanding the connection between optimization and generalization are discussed below.
- The Two Regimes of Deep Network Training: The learning rate schedule has a major impact on the performance of deep learning models. Instead of relying on a heuristic choice, this paper aims to understand the effects of different learning rate schedules and thereby develop a more principled way to select one. Specifically, two regimes are discussed:
On the other hand, the two regimes inherit the same reaction to momentum from the generalization perspective.
Building upon the aforementioned two regimes, this paper proposes a new training scheme consisting of two stages: 1) using the large-step regime to target good generalization; 2) using the small-step regime coupled with large momentum to target good optimization. Also, they show an ablation study of the transition epoch between the first stage and the second stage, benchmarked against the aforementioned heuristic three-step learning rate schedule.
-
Towards Explaining the Regularization Effect of Initial Large Learning Rate in Training Neural Networks: This paper shares the same motivation as the previous two-regime work, aiming to theoretically explain the effectiveness of an initial large learning rate and the annealing scheme. Its unique contribution is that it provides a concrete proof for the two-layer fully-connected network case.
-
Stiffness: A New Perspective on Generalization in Neural Networks: This paper investigates neural network training and generalization using the concept of stiffness. Specifically, it measures how stiff a network is by looking at how a small gradient step on one example affects the loss on another example. Given a data pair $(X, y)$, suppose the corresponding loss gradient can be represented as $\vec{g} = \nabla \mathbf{L} (f(X), y)$; we can then discuss the mutual influence between two independent data pairs, as shown in Fig. 4.
and then formulates the discrete (sign) or continuous (cos) stiffness metrics:
- The Break-Even Point on Optimization Trajectories of Deep Neural Networks: This paper investigates how the hyperparameters of SGD used in the early phase of training affect the rest of the optimization trajectory. Before talking about the concrete analysis, we need to keep in mind two concepts:
Based on the break-even point observation, this paper proposes two conjectures to investigate the effects of different hyperparameters: 1. Along the SGD trajectory, the maximum attained values of $\lambda_H^1$ and $\lambda_K^1$ are smaller for a larger learning rate or a smaller batch size. 2. Along the SGD trajectory, the maximum attained values of $\lambda_H^\* / \lambda_H^1$ and $\lambda_K^\* / \lambda_K^1$ are larger for a larger learning rate or a smaller batch size.
On the Frequency Domain Perspective
Understanding the training process of Deep Neural Networks (DNN) is a fundamental problem in the area of deep learning. Here are the papers analyzing DNN training from the frequency domain perspective. The concept of “frequency” is central to understanding the papers below. In this context, “frequency” refers to response frequency, not image (or input) frequency, as explained below.
- Training Behavior of Deep Neural Network in Frequency Domain: This paper analyzes the network training from the frequency perspective, aiming to claim the F-Principle: DNNs often fit target functions from low to high frequencies during the training process.
By examining the relative error of certain selected key frequency components (marked by black squares), one can clearly observe that DNNs of both structures, for both datasets, tend to capture the training data in order from low to high frequencies, as stated by the F-Principle.
- On the Spectral Bias of Neural Networks: This paper shares the same motivation and claim as the F-Principle.
What happens during the early phase of DNN training
Similar to humans and animals, deep artificial neural networks exhibit critical periods, which correspond exactly to the early phase of training. A lot of phenomena have been discovered during the early phase of network training. For example, sparse, trainable sub-networks emerge, gradient descent moves into a small subspace, and the network undergoes a critical period. Two recent works are briefly introduced below.
- Critical Learning Periods in Deep Networks: Researchers have documented critical periods affecting a range of species and systems; as machine learning researchers, it is natural to ask whether neural network training also experiences such critical periods. If so, when is the critical period? This paper answers that question using a deficit ablation study.
Further, to explore whether the critical period occurs in the early training phase, they conduct another ablation study on the deficit's starting epoch. The decrease in final performance can be used to measure the sensitivity to the deficit; the most sensitive epochs correspond to the early, rapid training phase. Afterwards, the network is largely unaffected by the temporary deficit.
- The Early Phase of Neural Network Training: Since the early stage of training is critical, this paper investigates it further, aiming to provide a unified framework for understanding the changes that DNNs undergo during this early phase of training.
Among them, the most attractive phenomenon is that during 500-2000 iterations (2-10 epochs; 1/80-1/16 of the training stages), rewinding starts to be highly effective. Building upon the [Lottery Ticket Hypothesis (LTH)](https://arxiv.org/abs/1803.03635), something important happens during the early phase of training, such that when rewinding the network, one should rewind to these early phases instead of the initial phase. As demonstrated by Fig. 13, rewinding variants perform better than lottery initialization.
Then, they probe what is more important for the early phase of training: the signs of the weights or the magnitudes of the weights? By conducting ablation studies on weight signs and weight magnitudes from initialization or the early phase, this paper finds that both signs and magnitudes are important for handling highly sparse scenarios. Also, they probe whether the weights in the early phase can be sampled from a distribution by shuffling the weights globally or locally and then testing their performance in highly sparse scenarios. They find that the weights do not exhibit any clear distributional structure; thus, so far, the early phase of training remains the only way to obtain a good initialization for the retraining phase.