TRACE: Rollout-Guided Quantization-Aware Training Brings FP4 Reinforcement Learning to MoE Language Models

Published · AI Daily — AI-assisted deep research, methodology & disclosure

TRACE is an FP4 reinforcement learning framework for Mixture-of-Experts language models. It uses quantization outcomes from the rollout side to guide FP4 rounding on the training side, which directly shrinks the gap between the two quantized paths. It caches only mantissa and scale information from deeper layers to limit overhead. According to the abstract, on four large MoE models, joint FP4 weight, activation and KV-cache rollout matches BF16 rollout in RL performance, with up to 5.4x faster rollout and better results than post-hoc FP4 quantization of BF16-trained policies.

The problem the paper targets

Reinforcement learning (RL) post-training is now a standard way to improve reasoning, coding and long-horizon behavior in large language models. It is also expensive. Every training step starts with rollouts: the current policy generates many samples, and only then does the model update. The authors state that rollout generation causes substantial compute and memory overhead. That is why teams look at low-precision rollout. FP4 is among the most aggressive floating-point formats that hardware supports today, and it can cover weights, activations and the KV cache.

The paper names a key weakness in existing FP4 RL methods. They mostly optimize quantization accuracy on the training path and on the rollout path independently. They do not directly reduce the gap between the two quantized execution paths. RL is very sensitive to that gap. The rollout engine samples trajectories with one set of FP4 numerics, and the trainer computes gradients on those trajectories with another. The mismatch acts like off-policy bias. Training becomes less stable, and final quality can drop.

The core idea of TRACE

TRACE stands for Train-Rollout Quantization Alignment via Compact GuidancE. It targets Mixture-of-Experts (MoE) language models and has two main parts.

The first part is rollout-guided quantization-aware training (QAT). In ordinary QAT, the training side decides by itself how each value rounds onto the FP4 grid. In TRACE, the quantization outcomes from the rollout side guide the FP4 rounding decisions on the training side. In plain terms, the trainer tries to reproduce the rounding that the rollout engine actually performed. This directly reduces the train-rollout discrepancy. It is the main difference from earlier work. The goal moves from "each path quantizes accurately on its own" to "both paths agree with each other".

The second part is an efficient quantization-information caching scheme. To guide the trainer, the system must store and move the rollout-side quantization information. That costs memory and communication. TRACE selectively keeps the mantissa and scale information from deeper layers, which reduces this extra cost. The abstract does not say which layers are kept or how much memory the cache adds. Readers should check the full paper for those details.

How it works, in engineering terms

FP4 formats usually use block scaling. A small group of values shares one scale factor, and each value inside the group uses only four bits. For each value, the rounding direction (to the nearest grid point above or below) is a discrete choice. If the two paths round the same weight or activation in different ways, the outputs differ by a small but systematic amount. Across many MoE layers, and with the discrete expert choice made by the router, such differences can grow.

A reasonable reading of TRACE is this. The rollout engine quantizes and records the outcome. The trainer then reads that record during its forward pass and uses it to pick the rounding, so that the two computation graphs match numerically as closely as possible. Errors in deeper layers affect the final output more directly, so caching deeper-layer information is a trade between accuracy and overhead. This is our interpretation of the abstract. The paper body is the authority on the exact algorithm.

Results and performance

The authors evaluate TRACE on four large-scale MoE language models, across reasoning, coding and long-horizon RL tasks. The abstract reports three main findings. First, TRACE enables joint FP4 weight and activation quantization together with an FP4 KV cache during rollout, while RL performance stays comparable to BF16 rollout. Low precision did not cost final quality in their experiments. Second, rollout speedup reaches up to 5.4x. Note the words "up to". This is a best-case figure. Real gains depend on the model, sequence length, batch size and hardware. Do not treat it as an average. Third, compared with post-hoc FP4 quantization of a policy that was trained in BF16, TRACE gives stronger final FP4 performance. This comparison matters. It shows that adapting the model to FP4 during training beats compressing it after training.

The abstract does not list benchmark scores, memory savings in percent, or model names. We do not guess them here.

Impact for developers and enterprises

For teams that run RL post-training, rollout often takes a large share of wall-clock time, especially with long chains of thought and long-horizon tasks. If FP4 rollout gives a multiple-fold speedup without hurting quality, the same GPU budget buys more iterations, or experiments finish sooner. MoE models have large weights and heavy KV-cache pressure, so they stand to gain more from FP4.

At the ecosystem level, the work needs FP4-capable hardware and inference kernels. Adoption has costs. The inference engine must export its quantization outcomes. The training framework must consume them. The data path between the two needs engineering work. The paper also supports a broader design rule: in an RL system, numerical consistency between training and inference should be a first-class goal, not a side effect.

Limits and what comes next

Several points call for caution. The 5.4x number is an upper bound. The experiments cover four MoE models, so it is unproven whether the method carries over to dense models or smaller scales. Rollout guidance couples the two sides more tightly: the trainer now depends on rollout output, which adds synchronization, storage and communication, even with the compressed cache. Finally, this is a fresh preprint without peer review, so wait for independent reproduction before you rely on it.

Open directions include extending the alignment idea to other number formats, combining it with asynchronous RL systems, testing how the cache scales to longer sequences, and checking whether safety and robustness behavior stay unchanged.

Summary

TRACE targets the real bottleneck of FP4 RL. The issue is not that one path quantizes too coarsely. The issue is that the training path and the rollout path do not line up.

With rollout-guided QAT and compact quantization-information caching, the abstract reports RL performance comparable to BF16 rollout and up to 5.4x rollout speedup. Whether it fits your stack depends on your hardware, framework and model scale. A small-scale reproduction is the sensible first step.

Sources