Skip to results
MLSift
← Feed
routineTheory & OptimizationScalar Linear Network2607.07884

Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks

Yedi Zhang, Peter E. Latham, Leena Chennuru Vankadara, Andrew Saxe

cs.LG

Abstract

In this short note we consider the gradient descent dynamics of deep scalar linear networks, $f(x) = \prod_{l=1}^L w_l x$, which enjoy exact time-course solutions for any integer depth. We show that even in this minimal model, the optimal depth-wise learning rate scaling depends on data, whereas data-agnostic scaling rules fail to transfer across depths. Under the data-dependent optimal scaling, the learning dynamics is independent of data and weakly dependent on depth, resulting in a constant linear convergence rate across all depths including infinity. We further show similar data-dependent effects in deep scalar linear networks with residual connections.

Topics

Classified with taxonomy v2 on Sat, 5 Sept 2026.

The PDF is 1–3 MB. Open it in your browser's viewer, or load it here.

Open PDF