Rounding in Preconditioner Space: Rethinking 4-bit AdamW Optimizer-State Quantization
A new paper recasts 4-bit AdamW state quantization as a rounding-space problem. It proposes ZIP-SR and ZE-EDEN, which cut the validation-loss gap to 32-bit AdamW by up to 70% from 130M to 2.7B parameters.
AdamW is the default optimizer for large-model training, and its cost is easy to underestimate. Beyond weights and gradients, every parameter carries two extra states, a first moment and a second moment. Stored in 32 bits, these states take about as much resident memory as the weights, and often more. Compressing them to 4 bits is one of the most direct ways to cut training memory. TorchAO already ships a 4-bit AdamW, yet a measurable validation-loss gap to 32-bit AdamW remains. The paper by Hanyang Li and four co-authors, posted to arXiv cs.LG on 8 October 2026, does not tune codebook shapes or block sizes. It asks a more basic question: when a quantizer chooses between two adjacent reconstruction levels, in which coordinate should it make that choice?
The authors call that coordinate the rounding space. The reason it matters is that quantization error does not stay put. It propagates through the moment recurrences and perturbs every later adaptive update. For the first moment, the update is roughly linear in the state, so rounding in state space is a mild problem. The second moment is different. Adam divides by the square root of the second moment, and that reciprocal is the preconditioner. The paper analyzes the quantization cell adjacent to zero and shows that a small mean error in the state does not imply a small mean error in the next step's preconditioner. A state that drifts slightly from zero is pushed through a steep reciprocal, so the error is magnified on the preconditioner side. A one-dimensional quadratic construction then shows that rounding in state space and rounding in preconditioner space produce qualitatively different optimization dynamics, not merely different numbers.
Those observations lead to the first method, Zero-Inclusive Preconditioner-space Stochastic Rounding, or ZIP-SR. It keeps zero in the second-moment codebook and computes the stochastic-rounding probabilities in preconditioner space rather than state space. Stochastic rounding is unbiased in expectation and turns a deterministic, systematic error into zero-mean noise. Choosing the right space places that unbiasedness on the quantity that actually sets the step size. The second method, Zero-Excluding EDEN calibration, or ZE-EDEN, takes a complementary route. Its second-moment codebook excludes zero, so quantized values have a positive floor and the reciprocal cannot blow up near zero. That floor itself distorts the preconditioner, so the method rescales each quantized second-moment block with an EDEN-style calibration to offset the distortion. Both configurations use 4-bit NormalFloat (NF4) for the first moment, and both apply targeted stochastic rounding to the LM-head first moment during the final 10% of training.
The experiments cover GPT-style and Llama-style pretraining at sizes from 130M to 2.7B parameters. At every evaluated size, both methods reduce the mean validation-loss gap between TorchAO 4-bit AdamW and 32-bit AdamW, and the largest reported reduction reaches 70%. In full-parameter supervised fine-tuning, both recipes reach lower validation loss than TorchAO while staying close to 32-bit AdamW on downstream tasks. The consistency matters more than any single large win. A gain that appears at every size suggests a mechanism at work, not luck in one run.
Our assessment is that the main contribution is a redrawing of the design space. Earlier work focused on how to design the codebook and how to cut blocks. This paper adds a third axis: the coordinate in which rounding happens. That view may well carry over to other optimizers that apply a nonlinear map to their stored state. Some restraint is still in order. The abstract gives no figures for memory saved or for throughput, the largest model is 2.7B parameters, and whether the gains hold at larger scale and longer training is a question for independent replication. For a team short on training memory, the sensible step is to treat ZIP-SR and ZE-EDEN as candidates for a small trial against TorchAO at their own scale, not as a drop-in change to a production recipe.
Sources
FAQ
What separates ZIP-SR from ZE-EDEN?
ZIP-SR keeps zero in the second-moment codebook and computes stochastic-rounding probabilities in preconditioner space. ZE-EDEN uses a codebook without zero and rescales each quantized second-moment block to offset the distortion from the positive floor. Both store the first moment in NF4.
Why does the rounding space matter more for the second moment?
Adam's update scales with the reciprocal square root of the second moment. Near zero that reciprocal is very steep, so a small state error becomes a large preconditioner error. The paper's local analysis shows small mean state error does not guarantee small mean preconditioner error.
Is the result ready for production use?
Not yet. The abstract reports up to a 70% cut in the validation-loss gap from 130M to 2.7B parameters and better fine-tuning loss than TorchAO. It gives no memory or throughput figures and no larger scales. Test against TorchAO at your own scale first.