AdaLoss: A Computationally-Efficient and Provably Convergent Adaptive Gradient Method

Xiaoxia Wu; Simon Du; Yuege Xie; Rachel Ward

AdaLoss: A Computationally-Efficient and Provably Convergent Adaptive Gradient Method

Xiaoxia Wu, Simon Du, Yuege Xie, Rachel Ward

[AAAI-22] Main Track

Keywords
Poster Session 5 @ Blue 2, Poster Session 10 @ Blue 2, Poster Session 5, Poster Session 10

Download Paper

Enter the Virtual Venue

Abstract: We propose a computationally-friendly adaptive learning rate schedule, ``AdaLoss", which directly uses the information of the loss function to adjust the stepsize in gradient descent methods. We prove that this schedule enjoys linear convergence in linear regression.

Moreover, we extend the to the non-convex regime, in the context of two-layer over-parameterized neural networks. If the width is sufficiently large (polynomially), then AdaLoss converges robustly \emph{to the global minimum} in polynomial time. We numerically verify the theoretical results and extend the scope of the numerical experiments by considering applications in LSTM models for text clarification and policy gradients for control problems.

Introduction Video

Sessions where this paper appears

Timezone

Poster Session 5

Sat, February 26 12:45 AM - 2:30 AM (+00:00)

Blue 2

Add to Calendar
Apple
Google
iCal File
Microsoft 365
Outlook.com
Yahoo

Poster Session 5
Poster Session 10

Sun, February 27 4:45 PM - 6:30 PM (+00:00)

Blue 2

Add to Calendar
Apple
Google
iCal File
Microsoft 365
Outlook.com
Yahoo

Poster Session 10