All labs

Gradient Descent: Step Size, Conditioning, and Momentum

Watch a convex bowl defeat gradient descent — the trouble is the condition number and the step size, not the landscape.

Ch 13 13.3
Animator
Optimization
18 mindifficulty 3/5

Controls

Optimizer
0.090

Stability needs η < 2/L = 0.100

20

Ratio of the largest to the smallest curvature.

0.900
60
-8
6
Non-quadratic valley

A Rosenbrock-style curved ravine instead of the bowl.

The path over the landscape

Ellipses are level sets; the gradient is perpendicular to each of them.

-10-50510-10-50510θ₁θ₂
trajectorystartminimumlevel sets

Loss per iteration

Linear scale, clipped at the start value.

01002003000102030405060iterationloss

Log loss per iteration

Straight line means geometric convergence.

-3-2-101230102030405060iterationlog₁₀ loss
Final loss
3.89e-4
start 392
Distance to optimum
0.02790
Iterations to 10⁻⁴ of start
36
within 60 steps
Stability bound 2/L
0.1000
η is inside the bound
Predicted rate (κ−1)/(κ+1)
0.9048
per-step contraction for optimal η
Status
converging
f(θ) = ½(θ₁² + κθ₂²)