by Haytham ElFadeel - hfadeelm@gmail.com September 6, 2026
Every deep learning practitioner carries around a set tradition, today we open up one of them ‘Weight Norm and Weight Decay’. This essay explore what does weight decay actually do, and it relationship with learning rate, Grokking, effective learning rate, Muon, Spectral spheres and Maximal Update Parametrization.
1. Why Would Magnitude Matter?
The classical reasons are:
- Smaller weights generalize better (Classic L2 regularization, Krogh & Hertz, 1991.)
- Weight decay is basically required for grokking: the delayed jump from memorization to generalization on algorithmic tasks (Power et al., 2022). Turn it off and the network tend to not grok / generalize and mostly memorizes.
- Some people even describe a "Goldilocks zone" of weight norm where generalization happens (Liu et al., 2023).
We rarely ask why. Almost every modern architecture is full of normalization layers (LayerNorm, RMSNorm, QK-Norm). A linear layer followed by a normalization is scale-invariant:
.
Multiply the weights by 10 or by 0.1 and the network computes exactly the same function. So in what sense can "small weights" be better? For these layers, a weight-norm penalty can't regularize the function at all, because there is nothing to regularize. So why we care about Weight decay and weight magnitude. Van Laarhoven (2017) pointed this out years ago, and Zhang et al. (2019) gave the answer this blog post builds on. Weight decay still matters a lot in these networks, but through the optimization dynamics, not through the function class.
So here's the question behind the rest of this post:
If the norm is a free choice, why does controlling it matter so much?
Spoiler alert / TLDR: The magnitude of the weights doesn't set what the network computes. It sets how big a step the optimizer takes in function space. Without weight decay, the weights grow and that effective step size quietly shrinks until the network stops learning new features. Weight decay, Muon, spectral spheres and μP are all ways of controlling it.
2. Weight Decay 101: Penalty vs. Constraint
2.1 SGD: weight decay is an L2 penalty
For plain SGD weight decay is a penalty. Take the regularized loss:
.
One SGD step on gives
which is exactly "shrink the weights by , then take a gradient step”. So L2 regularization and weight decay are the same thing here. The optimizer accepts a slightly higher in exchange for a smaller . This is a penalty.
2.2 Adam and other adaptive optimizers breaks the equivalence
If you add to the gradient (a "coupled" L2 penalty), it gets divided by along with everything else. Coordinates with a history of large gradients then get less regularization. AdamW fixed this by decoupling the decay from the adaptive update:
The same template now covers Lion, Muon, and pretty much every modern optimizer.
2.3 Adaptive optimizers turn the penalty into a constraint
With AdamW, Lion, and Muon weight decay acts like a constraint rather than a penalty.
SGD's push grows with the gradient. If the data really wants a weight to be large, the gradient on that weight stays large, and SGD pushes as hard as it needs to. Weight decay pulls back in proportion to the weight, the two meet at a compromise, and a strong enough signal can always win.
Adaptive optimizers push with a capped strength. Adam divides each gradient by its own running size, so no coordinate moves much more than per step, whether its gradient is tiny or huge. When the gradient is mostly noise, the momentum averages out and Adam's step is much smaller. Lion takes the sign of the momentum, so its step is always exactly , and Muon fixes the step size for a whole matrix at once. Either way, the optimizer's push now has a ceiling.
Weight decay doesn't push with a fixed strength. It pulls each weight back by , a pull that keeps growing as the weight grows. When the weight is small, the optimizer wins easily. But at some point the weight decay pull catches up with the hardest push the optimizer can ever make, and that point is . Beyond it, decay wins every time, no matter how badly the loss wants the weight to grow. That's the wall.
So where does AdamW end up? Each weight lands in one of two places:
- Pressed against the wall, at exactly , if the loss keeps pushing it outward, step after step.
- Somewhere inside, where the loss is roughly balanced. Inside the wall, decay is still pulling, and what happens next depends on whether the loss pushes back. If shrinking the weight raises the loss, the gradient resists. Because Adam normalizes gradient size, even a small but consistent gradient gets a sizable step. If shrinking the weight costs nothing weight decay keeps shrinking it.
Optimizer | How hard it can push | What weight decay becomes |
SGD | grows with the gradient | a penalty: minimize |
AdamW, Lion | fixed, per coordinate | a box: every weight stays within |
Muon | fixed, per matrix (spectral norm) | a cap/spectral ball: the largest singular value stays below |
3. Weight norm and the effective learning rate
What the effective learning rate is
For a scale-invariant layer, that depends on how big the weights already are. The layer only cares about the direction of its weight matrix, not its length. So the only part of a step that matters is how much it turns that direction. And how much a fixed-size step turns the weights depends on their size.
So the learning rate you set doesn't tell the whole story. What matters is the step size relative to the size of the weights. That relative step size is the effective learning rate (ELR). If the weights double in size, the ELR halves (for Adam), even though you never touched the learning rate. And if the weights keep growing during training, the ELR keeps falling.
Scale invariance, , gives us two facts:
- The gradient is always perpendicular to the weights: .
- Bigger weights get proportionally smaller gradients: .
Using the second fact, an SGD step of size on weights of norm changes their direction exactly as much as a step of size: would on weights of norm 1.
Adam normalizes away the gradient's scale, so one factor cancels and . Either way, the angle the weights turn in one step is roughly .
Why the weights grow on their own
Using the first fact. Every gradient step is perpendicular to the weights, so it moves them sideways. Moving sideways from a point on a circle always takes you slightly outside it. Without weight decay:
The norm can only grow, so the ELR can only shrink. Even with a constant learning rate, training is running a decaying learning-rate schedule that nobody chose. Over time the network makes smaller and smaller changes to its features, and learning slows down more than it should.
How weight decay keeps the ELR from collapsing
Weight decay adds a force that pulls the weights back toward zero. Now two things act on the norm: gradient steps push it outward, and decay pulls it inward. After a while they balance and the norm settles at a steady value.
What matters is what the ELR does at that balance point. For SGD, the norm settles where , with the gradient size at unit norm. The angle the weights turn per step becomes
The gradient size and the starting norm cancel out. Once the norm has settled, the effective step size depends only on the product of the learning rate and the weight-decay coefficient. The norm ends up wherever it needs to be for that to hold. Kosson et al. (2024) call this rotational equilibrium. It takes roughly steps to reach, and AdamW behaves the same way.
We can summarize our story so far as:
Weight decay matters because it keeps the effective learning rate from collapsing as the weights grow, not because small weights are better in themselves.
Two recent papers test this directly: one on grokking and reinforcement learning, the other on LLM pretraining.
Evidence from grokking and reinforcement learning
Lyle et al. (2025) start from the observation that grokking and primacy bias share a symptom. Primacy bias is the RL problem where features learned from early data get in the way of learning later. In both cases the network is stuck with features that don't generalize and has stopped learning new ones. If the ELR controls how much the network can change its features, raising it should help in both.
To separate the norm from the ELR, they train a one-layer transformer on modular addition in two ways. Once normally, and once with the weights rescaled back to their original norm every 100 steps. In the second setup, weight decay can't make the weights smaller. If small weights were what caused grokking, it should stop happening. It doesn't. The rescaled networks grok faster and more consistently, while the normal ones need strong weight decay to do the same. Learning rate matters more than weight decay, and a higher value of one can make up for a lower value of the other. That's what you'd expect if the ELR, not the norm, is what matters.
There's one exception, when LayerNorm is added, the attention inputs become large and the attention softmax becomes too sharp to train well. At that point a large ELR alone no longer leads to grokking. Shrinking the LayerNorm gains with weight decay fixes it. So the size of the weights still matters in the parts of the network that aren't scale-invariant.
Because the ELR controls feature learning, grokking can be triggered on demand. Train for hundreds of thousands of steps with a learning rate that's too small, then raise it, and the model groks shortly after. This works cleanly with Normalize-and-Project (NaP): after each step, every weight matrix is rescaled back to its original norm, so the ELR is exactly the learning rate you set.
The same approach helps on harder problems:
- Warm-started image classification (train on part of the data first, then on all of it): keeping the norm fixed and raising the learning rate again removes the usual drop in test accuracy.
- Atari, DQN and Rainbow at a high replay ratio: cycling the learning rate early in training and then lowering it gives consistent gains. A high ELR lets the network replace outdated features, and a low ELR afterwards lets it fine-tune.
Evidence from LLM pretraining
LLM transformers aren't exactly scale-invariant: residual connections, embeddings and the output layer all depend on scale. So it's fair to ask whether the ELR still explains training at that scale.
Liu et al. (2026) test this directly. They train runs with very different learning-rate and norm schedules (rising, falling, oscillating), set up so that every run has the same ELR, , at every step.
The loss curves come out almost identical. They ran 26 comparisons, covering:
- AdamW, Muon and Signum
- dense and mixture-of-experts models from 100M to 1B parameters
- three datasets
The average difference in loss is a few thousandths. Changing only the random seed moves the loss by one to two hundredths. So two runs with very different learning rates and norms, but the same ELR, end up closer than two runs that differ only in their seed.
They then look at weight decay directly. Take an AdamW run with and a run without weight decay. At the same learning rate, their loss curves differ. But give the no-decay run a learning-rate schedule adjusted so its ELR matches the other run's, and it reproduces the weight-decay run's loss curve almost exactly. The same holds for Hyperball (Wen et al., 2026), which keeps each matrix at a fixed Frobenius norm. In both cases, norm control affects the loss through the ELR.
The ELR also explains delayed acceleration: runs with weight decay often have higher loss for most of training and only pull ahead near the end. Weight decay keeps the ELR higher for longer. A higher ELR makes more progress, but it also adds more noise, and that noise hides the progress until the learning rate decays at the end of training. The benefit is earned early and only shows up late.
The authors get an even lower final loss by making the norm grow faster near the end of training, because that makes the ELR drop more sharply. A larger norm giving a better result doesn't fit the idea that smaller weights are better. What helps is the ELR schedule that norm control produces, not any particular norm. These results are about the loss curve; they don't directly say anything about learned representations or downstream accuracy.
What this means for "smaller weights generalize better"
Weight decay does two things at once: it keeps the weights small, and it keeps the ELR from shrinking. The second is what helps, because it keeps the network learning new features for longer. Small weights tend to appear alongside good generalization because both come from weight decay, not because one causes the other. That's why the correlation is so reliable even though the explanation stayed unclear for so long.
The exception is the parts of the network that aren't scale-invariant. There, the size of the weights changes the output directly:
- how sharp the attention and output softmax are
- how much each block adds to the residual stream
- the normalization gains
- the embeddings
4. Cautious Weight Decay: Stop Fighting the Optimizer
So far weight decay has two faces. It's a useful ELR thermostat, and for adaptive optimizers it's also a constraint that quietly changes the problem you're solving. Can we keep the first and drop the second? Chen et al. (2026) propose a one-line answer.
Look at a single coordinate of the update, . Either the optimizer is already moving toward zero, in which case weight decay pushes the same way, or the optimizer wants to grow and decay pulls it back. In the second case the two forces play tug-of-war, and that tug-of-war is exactly what pins AdamW's weights to the wall: the fixed point is where the optimizer's push outward and decay's pull inward cancel out.
Cautious Weight Decay (CWD) just fix this. It applies decay only where the update and the weight agree in sign:
Once CWD reaches a minima one, the only force left is weight decay where it doesn't hurt the loss, so the weights slide along the valley floor, shrinking wherever they can. Weight decay stops being a tax or a wall and becomes a tie-breaker: among solutions that fit equally well, prefer the smaller ones.
In practice there are gains, consistent, and free. On 338M to 2B models, CWD lowers the final loss for AdamW, Lion and Muon with the baseline's hyperparameters unchanged, and the best doesn't move. On OLMo-1B (100B tokens) it lowers validation loss from 2.65 to 2.56 with AdamW and from 2.51 to 2.42 with Muon.
How does this fit the ELR story? Standard decay multiplies the whole matrix by , which shrinks the norm without touching the direction, so on a scale-invariant layer it acts only through the ELR. CWD shrinks some coordinates and not others, so it also rotates the weights, and it ends training with a norm between AdamW's and no decay's. Whether its gain comes from the rotation or from a different ELR schedule is still open.
5. μP: the theory underneath
Sections 3-4 were about how the update-to-weight ratio behaves over training. μP is the theory that says how that ratio should behave as models get wider.
The problem
Train the same Transformer at increasing widths with Adam and the default setup, and the best learning rate moves as the model grows, which makes tuning big models painful. In the Maximal Update Parametrization (μP), the best learning rate stays put, so you can tune a small proxy and copy the settings to the big model (Yang et al., 2022). Their headline result: settings tuned on a 40M-parameter proxy beat the published 6.7B GPT-3, at a tuning cost of 7% of one pretraining run.
Loss vs. learning rate at different widths. In the standard parametrization (left) the best learning rate shifts with width; in μP (right) it stays put. (Image source: Yang et al. 2022)
The core idea
The whole theory rests on one asymmetry. A sum of independent, zero-mean terms grows like (the central limit theorem). A sum of correlated terms grows like (the law of large numbers).
At initialization the weights are independent of the input, so each output is the first kind of sum, and entries of size keep it order one (standard fan-in init). But the update is built from the input, since a linear layer's gradient is , so is the second kind of sum and grows like . To keep it order one, the update's entries need to be about . The standard setup gets the first part right and the second wrong, which is why logits blow up with width after only a few steps.
What μP asks for
μP is defined by three requirements that should hold throughout training as width grows: every activation stays order one, the output stays bounded, and every layer is updated as much as possible without blowing up, so that all layers keep learning features instead of freezing. For Adam this comes down to a short recipe (standard-parametrization values in parentheses):
Input weights | Hidden weights | Output weights | |
Init. variance | |||
Adam LR |
For Transformers, attention logits are also scaled by instead of , because queries and keys become correlated during training. Learning rate, schedule, momentum and initialization scale all transfer across widths, but the paper lists weight decay among the settings that don't. Hold on to that.
The spectral version
Yang, Simon & Bernstein (2023) restated μP as a single condition on spectral norms, which is the form that connects to everything else here. Asking each layer to keep its activations, and their changes, at an order-one RMS size comes down to
The updates obey the same law because they're low-rank and aligned with the incoming activations, so their spectral norm really is their effect on the output. Divide one condition by the other and you get . In other words, μP is a rule that the effective learning rate shouldn't change with width.
6. Muon: steepest descent in the spectral norm
So far the optimizer has been a black box, Let's open it. Strip away the moving averages and most optimizers solve the same small problem: find the step that lowers the linearized loss the most, within a ball of radius (Bernstein & Newhouse, 2024):
The choice of norm is the optimizer. The Frobenius norm gives normalized SGD. The max-entry norm gives , which is signSGD, basically Adam without the moving averages. The spectral norm gives , where : keep the gradient's singular vectors and throw away its singular values. That's Muon.
Why the spectral norm? Because a weight matrix isn't a bag of numbers; it's a linear map, and says the spectral norm is exactly the limit on how far any output can move. It's the natural trust region for a linear layer, and per-coordinate methods like Adam don't see it.
An SVD every step would be too slow, so Muon (Jordan et al., 2024) approximates with five iterations of an odd polynomial, starting from :
Each iteration applies to every singular value. The coefficients are tuned to pull small singular values up fast, at the price of landing them in roughly instead of exactly 1, which turns out to be fine.
Why would flattening the spectrum help? Jordan et al. noticed that the momentum updates of transformer weight matrices are dominated by a handful of directions. Orthogonalizing boosts the "rare directions" that are small in the update but still matter for learning.
In practice Muon handles only the 2D hidden matrices (embeddings, the output head and the gains stay on AdamW), and the update gets a shape factor: Moonlight (Liu et al., 2025) uses to match AdamW's update size, while μP suggests (the μP update size).
Muon's missing half
Muon controls the numerator of the ELR exactly: every step has a fixed spectral norm. It says nothing about the denominator, and without weight decay the weights drift. Moonlight found weight decay essential for scaling Muon, because without it the weights and layer outputs keep growing, eventually past what bf16 handles well. Large Muon runs have also hit exploding attention logits, the non-scale-invariant magnitude, usually patched with QK-norm, logit clipping or soft-capping. As Xie et al. (2026) put it, Muon is only "half-aligned" with μP: it controls the updates but lets the weights wander. That suggests an obvious next step.
7. MuonSphere and SSO: pin the weights too
If Muon controls the update, why not control the weights as well? Xie et al. (2026) put each hidden matrix on a spectral sphere: its top singular value is fixed at a radius (the μP target), and every update has spectral norm .
The simple version, MuonSphere, rescales the weights back onto the sphere before each step, , and then takes an ordinary Muon step of size : Normalize-and-Project, moved to the spectral norm. The catch is that part of each Muon step points straight outward or inward, and the next rescaling just undoes it.
The full Spectral Sphere Optimizer (SSO) removes the waste by only allowing steps that are tangent to the sphere. The direction that grows the spectral norm fastest is , built from the top singular vectors of , so the step has to satisfy . With a Lagrange multiplier , the best tangent step turns out to be another matrix sign:
Note that SSO pins only the top singular value, leaving the rest of the spectrum free, and that it drops weight decay on the hidden matrices entirely: pinning already bounds the whole matrix, so decay is no longer needed.
On the sphere, the ELR is the LR
Here's the payoff for our story. On the sphere,
The effective learning rate is no longer something that emerges from norm growth or settles after steps. It's literally the number in your learning-rate schedule. NaP and Hyperball do the same thing in the Frobenius norm; MuonSphere and SSO do it in the spectral norm.
Does it work?
On a 1.7B dense model trained on 100B tokens, with a learning rate tuned for AdamW, Muon reaches AdamW's final loss in 12% fewer steps and SSO in 19% fewer. Average downstream accuracy goes from 54.75 for AdamW to 55.26 for Muon, 56.19 for MuonSphere and 56.35 for SSO. SSO also gives the best expert load balance on an 8B mixture-of-experts model and the most stable training on a 200-layer model, and its activations stay put while AdamW's grow about 100× larger.
8. μP meets Adam, Muon and the sphere
Adam. Adam moves every entry of a weight matrix by about the same amount, roughly , however large or small its gradient is. A wider layer has more entries. Because the update is built from the layer's input, all those small changes push the output in the same direction and add up rather than cancel. So with a fixed learning rate, one Adam step changes the output more and more as the model gets wider. To keep the change constant, the learning rate has to shrink as the layer gets wider. That's where μP's rule comes from: Adam's learning rate for hidden layers scales like .
Muon. Muon normalizes the update as a whole matrix rather than entry by entry. Its update always has a spectral norm of exactly 1, whatever the matrix size. Multiply it by and you get the μP update size at every width, with the same learning rate. No width-dependent learning rate is needed.
But μP has two conditions: one on the size of the updates and one on the size of the weights. Muon only handles the first. During training the weights grow or shrink, and they do so at different rates in models of different widths. That changes the update-to-weight ratio μP is trying to hold fixed. Xie et al. (2026) see this directly. Across models from 70M to 1.8B parameters, Muon's best learning rate still shifts with width, while SSO's stays the same and reaches a lower loss.
MuonSphere and SSO. These keep the weights' spectral norm fixed at the μP target, and the update at the matching size, at every step. Both of μP's conditions hold for the whole of training, not just at the start. That's what "fully μP-aligned" means.
How weight decay fits μP
μP sets the right update-to-weight ratio at initialization. But as we showed, the weights don't stay near their initial size. After roughly steps, the weight norm is set by the balance between updates pushing it out and weight decay pulling it in. From then on, the relative update size is set by , not by the initialization.
Kosson et al. (2025) measured this in LLM training. They found that μP's assumptions hold only early on. For most of training, weight decay is what sets the relative update size.
This causes a problem with the usual setup. When a model gets times wider, μP divides Adam's hidden-layer learning rate by . If the weight decay stays the same, then also drops by . So later in training, wider models make relatively smaller updates than narrow ones, and the learning rate no longer transfers cleanly.
The fix is to multiply by at the same time, so stays the same at every width. With that change, μP's learning-rate scaling only affects the early part of training, where it acts like an extra warmup. That's the point of the paper's title: in practice, weight decay may matter more than μP for learning-rate transfer.
Putting this together, there are three ways to keep the update-to-weight ratio the same across widths:
- μP alone. The ratio is right at the start of training.
- μP plus weight decay scaled to keep fixed. The ratio is also right once the weight norm has settled.
- A sphere (MuonSphere, SSO, or Hyperball in the Frobenius norm). The ratio is right at every step, by construction.
9. Putting it all together
For the scale-invariant parts of a network, the weight norm's only job is to set the effective learning rate. Left alone, the norm grows and the ELR collapses. Weight decay is a tool that holds it near , and Cautious Weight Decay keeps the tool without the wall. Muon fixes the ELR's numerator, MuonSphere and SSO fix the denominator too, and μP in the theory that says both should scale as so the ratio doesn't change with width.
Cited as:
ElFadeel, Haytham. "Weight Decay Was Never About Small Weights: Effective Learning Rates, Cautious Decay, Muon, Spectral Spheres, and μP." (Sep 2026).
@article{elfadeel2026weightdecay,
title = {Weight Decay Was Never About Small Weights: Connecting Weight Decay, Learning Rate, Maximal Update Parametrization, Muon, and Muon Sphere together},
author = {ElFadeel, Haytham},
year = {2026},
month = {September}
}References
[1] A. Krogh and J. Hertz. "A Simple Weight Decay Can Improve Generalization." NeurIPS 1991.
[2] I. Loshchilov and F. Hutter. "Decoupled Weight Decay Regularization." ICLR 2019.
[3] T. van Laarhoven. "L2 Regularization versus Batch and Weight Normalization." arXiv:1706.05350, 2017.
[4] G. Zhang et al. "Three Mechanisms of Weight Decay Regularization." ICLR 2019.
[5] E. Hoffer et al. "Norm Matters: Efficient and Accurate Normalization Schemes in Deep Networks." NeurIPS 2018.
[6] Z. Li and S. Arora. "An Exponential Learning Rate Schedule for Deep Learning." ICLR 2020.
[7] R. Wan et al. "Spherical Motion Dynamics: Learning Dynamics of Normalized Neural Network using SGD and Weight Decay." NeurIPS 2021.
[8] A. Kosson, B. Messmer, and M. Jaggi. "Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks." ICML 2024.
[9] X. Wang and L. Aitchison. "How to Set AdamW's Weight Decay as You Scale Model and Dataset Size." arXiv:2405.13698, 2024.
[10] S. Xie and Z. Li. "Implicit Bias of AdamW: ℓ∞-Norm Constrained Optimization." ICML 2024.
[11] L. Chen, B. Liu, K. Liang, and Q. Liu. "Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts." ICLR 2024.
[12] L. Chen, J. Li, and Q. Liu. "Muon Optimizes Under Spectral Norm Constraints." arXiv:2506.15054, 2025.
[13] A. Power et al. "Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets." arXiv:2201.02177, 2022.
[14] Z. Liu, E. Michaud, and M. Tegmark. "Omnigrok: Grokking Beyond Algorithmic Data." ICLR 2023.
[15] V. Varma et al. "Explaining Grokking Through Circuit Efficiency." arXiv:2309.02390, 2023.
[16] C. Lyle, G. Sokar, R. Pascanu, and A. György. "What Can Grokking Teach Us About Learning Under Nonstationarity?" CoLLAs 2025.
[17] C. Lyle et al. "Normalization and Effective Learning Rates in Reinforcement Learning." NeurIPS 2024.
[18] C. Lyle et al. "Disentangling the Causes of Plasticity Loss in Neural Networks." arXiv:2402.18762, 2024.
[19] E. Nikishin et al. "The Primacy Bias in Deep Reinforcement Learning." ICML 2022.
[20] Z. Liu, R. Zheng, S. Zhang, C. Tian, K. Chen, Z. Zhang, and L. Wu. "Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining." arXiv:2608.24814, 2026.
[21] F. D'Angelo et al. "Why Do We Need Weight Decay in Modern Deep Learning?" NeurIPS 2024.
[22] K. Wen, X. Dang, K. Lyu, T. Ma, and P. Liang. "Fantastic Pretraining Optimizers and Where to Find Them II: Hyperball Optimization." arXiv:2606.16899, 2026.
[23] L. Chen et al. "Cautious Weight Decay." ICLR 2026.
[24] K. Liang, L. Chen, B. Liu, and Q. Liu. "Cautious Optimizers: Improving Training with One Line of Code." arXiv:2411.16085, 2024.
[25] K. Jordan et al. "Muon: An Optimizer for Hidden Layers in Neural Networks." Blog post, 2024.
[26] J. Bernstein and L. Newhouse. "Old Optimizer, New Norm: An Anthology." arXiv:2409.20325, 2024.
[27] N. Amsel, D. Persson, C. Musco, and R. Gower. "The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm." arXiv:2505.16932, 2025.
[28] J. Liu et al. "Muon is Scalable for LLM Training." arXiv:2502.16982, 2025.
[29] Kimi Team. "Kimi K2: Open Agentic Intelligence." arXiv:2507.20534, 2025.
[30] T. Pethick et al. "Training Deep Learning Models with Norm-Constrained LMOs." ICML 2025.
[31] T. Xie et al. "Controlled LLM Training on Spectral Sphere." arXiv:2601.08393, 2026.
[32] J. Bernstein. "Modular Manifolds." Thinking Machines blog, 2025.
[33] G. Yang and E. J. Hu. "Feature Learning in Infinite-Width Neural Networks" (Tensor Programs IV). ICML 2021.
[34] G. Yang, E. J. Hu, et al. "Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer." arXiv:2203.03466, 2022.
[35] G. Yang, J. B. Simon, and J. Bernstein. "A Spectral Condition for Feature Learning." arXiv:2310.17813, 2023.
[36] A. Kosson, J. Welborn, Y. Liu, M. Jaggi, and X. Chen. "Weight Decay may matter more than μP for Learning Rate Transfer in Practice." ICLR 2026.
[37] M. Wortsman et al. "Small-scale Proxies for Large-scale Transformer Training Instabilities." ICLR 2024.