Learning to Steer, Steering to See: The Activation Geometry of RLVR in LLMs and the Alpha-Stabler Training Fix

Published · AI Daily — AI-assisted deep research, methodology & disclosure

This paper treats RLVR as an intervention problem in activation space. With the base model frozen and one trainable vector per controlled layer, it recovers over 85% of the full fine-tuning gain on most tasks. The authors find two geometric properties: the effective capacity is small but not infinitely compressible, and the control directions sit in the low-variance complement of the activation principal subspace. Building on this, Alpha-Stabler uses a PSI monitor as an early collapse warning and projects principal-subspace components out of activation gradients. It keeps training stable for 2,000 steps. The study covers text-only LLMs, not multimodal models. The code is public on GitHub.

Reinforcement learning is now a central tool for improving the reasoning of large language models, yet nobody has a clear picture of what it actually changes inside the network. Parameter updates are enormously high-dimensional, and optimization noise and configuration differences hide the changes that matter. A team from USTC, Tencent Hunyuan and BUPT, in a paper submitted on 28 September 2026 (arXiv 2609.34344, cs.LG), takes a different route. They stop looking at parameter space and treat reinforcement learning with verifiable rewards (RLVR) as an intervention problem in activation space.

One correction first. The background note supplied with this item described the work as covering vision-language reasoning. The arXiv title is "Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors", and the experiments use text-only LLMs on mathematical reasoning, scientific reasoning, code generation and instruction following. No multimodal benchmark appears in the abstract or the text we read. Everything below follows the paper itself.

The core method: one trainable vector as a probe

The authors freeze the base model and add a trainable vector to the output of selected layers: h'_l = h_l + delta_l. The vector is input-invariant. It is shared across all samples and all token positions, so each controlled layer gains just one d-dimensional vector, where d is the hidden size. The training signal is distillation, not RL directly. First a teacher is trained with GRPO or DAPO. Then the "frozen base plus vectors" student is trained to match the teacher. The study runs both on-policy and off-policy distillation, and compares three update forms: full-parameter training, LoRA and vector steering.

This turns a vague idea into a testable claim. If RL gains are truly low-dimensional, a handful of parameters should reproduce them. The models are DeepSeek-R1-Distill-Qwen-1.5B and 7B, plus Qwen3-4B, 8B and 14B. In total the paper covers 5 LLMs and 6 verifiable-reward tasks.

Property 1: effective manifold capacity is small, but not infinitely compressible

Under on-policy distillation, LoRA and full-parameter training perform about the same. A single steering vector per controlled layer still recovers over 85% of the gain from full fine-tuning on most evaluated tasks. Under off-policy distillation, with teacher-generated trajectories, recovery on the evaluation sets stays comparably high. So the compressibility does not depend on the student's own rollouts.

The limits are the interesting part. The paper names two. The first is the teacher-student distribution gap. When the two are close, one vector is enough. When the gap widens, a more flexible parameterization is needed. The second is insertion depth. Fixed vectors work well at low and middle layers and degrade near the output. A nonlinear analysis shows that this is not because the optimization signal is weak at depth. It is an expressiveness bottleneck: a low-dimensional, input-independent shift cannot express the input-dependent correction that high layers need. In short, the required capacity depends on how well the intervention form matches the correction.

Property 2: the control manifold is separated from the principal subspace

Next the authors extend one vector per layer to several mutually orthogonal vectors. Later vectors score lower on their own, yet each can still support the target behavior alone. The leading direction dominates but is not the only effective one. Within one task and base model, training runs with very different parameters converge, after compression, to a reproducible set of task-related directions. Across tasks, the alignment between direction vectors tracks capability transfer.

The key finding is where these directions sit. They lie mainly in the low-variance complement of the activation principal subspace. RL does not push on the high-variance directions that dominate the representation. It makes fine adjustments along quiet directions. The paper calls this control manifold separation.

Alpha-Stabler: turning geometry into a training tool

The authors then build Alpha-Stabler, a plug-and-play framework with two modules.

The Predictor monitors principal-subspace intrusion (PSI). A fixed principal subspace is estimated once from the frozen base model. During training, PSI measures the share of token-level activation-difference energy that falls inside that subspace: the squared Frobenius norm of the projected differences divided by the squared norm of all differences. In successful runs, activation shifts stay in the low-variance complement. Rising PSI goes together with degraded performance and collapse, so it works as an early warning.

The Controller acts only in the backward pass. It removes the principal-subspace component of the activation gradient and keeps the component in the orthogonal complement. It adds no trainable parameters and needs no change to the optimizer, the reward function or the architecture.

Results and cost

The paper reports that Alpha-Stabler keeps training stable for 2,000 steps and consistently improves RL gains. Reward growth is smoother than with gradient clipping and other gradient projection methods.

The authors stress that the gain comes from removing gradient components in a specific subspace, not from suppressing gradients in general. On cost, an appendix figure shows nearly identical score curves with and without Alpha-Stabler at equal wall-clock time, which points to very little overhead. One caveat: we did not find a single headline table of absolute benchmark scores in the material we read, so this article does not quote one.

What it means for developers and enterprises

Monitoring is the easiest win. PSI is a cheap metric that can sit in a training dashboard and trigger a rollback or a lower learning rate before the reward curve falls apart. Task adaptation is the second.

If an RL gain can be reproduced by one vector per layer, storing and switching task-specific behavior becomes far cheaper, which suits multi-task serving. Transfer prediction is the third. Because direction alignment correlates with transfer, vector geometry could give a rough guide to whether two tasks help each other before any large run. The code is public at github.com/caiyuchen-ustc/On_Policy_Vector_Training.

Limitations and next steps

The authors list their own limits. The theory rests on empirically motivated assumptions, not first-principles derivations. The study covers RLVR on tasks such as math and code. It does not test RLHF, multi-turn agentic decision making or open-ended generation. Future work includes neuron attribution and causal tracing to relax the assumptions, extension to RLHF and agent settings, and using off-principal localization as a design principle, for example with complement-space regularizers or PSI as an adaptive signal for learning rate and intervention strength.

Overall this is an analysis paper. Its value is a usable geometric view of RLVR and a low-cost stabilizer built on it. Readers should treat the conclusions as results for text-only LLMs. Whether they hold for multimodal models is not answered here.

Sources