MiMo-V2.6: Scaling Agentic Reinforcement Learning with Groupwise Grading and a Frozen MoE Router

Published · AI Daily — AI-assisted deep research, methodology & disclosure

Xiaomi's LLM-Core team presents MiMo-V2.6: Pro has 1.02T total and 42B active parameters, Flash has 310B and 15B. The paper bets post-training compute on reinforcement learning: 1,568 prompts per step, contexts up to 1M tokens, groupwise agentic grading, a frozen MoE router and layered reward-hacking defenses. On DeepSWE, Pro rises from 58.4 to 72.6 and Flash from 48.7 to 65.7 during RL. The authors plan to open-source the RL framework and environments.

On 8 October 2026 the Xiaomi LLM-Core team posted arXiv:2610.11959, "MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement". The paper introduces the MiMo-V2.6 family of omni-modal models. Its central claim is simple: after pre-training and mid-training, keep scaling reinforcement learning (RL) compute, and intelligence keeps rising. One correction to the brief that triggered this article: the paper is about agentic RL across code, general, visual and cyber environments. It is not a math-and-logic curriculum-generation paper. The two models are MiMo-V2.6-Pro (1.02T total parameters, 42B active) and MiMo-V2.6-Flash (310B total, 15B active). Both are sparse mixture-of-experts (MoE) models with no shared experts and 8 experts active per token. Flash has 48 layers, hidden size 4096 and 256 experts. Pro has 70 layers, hidden size 6144 and 384 experts. The first block uses global attention with a dense feed-forward network. The attention design is a hybrid sliding-window scheme (hybrid-SWA). Local sliding-window attention layers alternate with global attention layers, and the window is only 128 tokens. Flash uses 39 sliding-window layers and 9 global layers. Pro uses 60 sliding-window layers and 10 global layers. Because most layers look at just 128 tokens, the key-value cache stays small. That is the engineering precondition for running RL at context lengths up to 1M tokens. A five-layer speculative decoding module also predicts 7 tokens per forward pass.

The data budget is large. Flash saw 48T tokens, of which 26T were text and 22T omni-modal. Pro saw 30T tokens, of which 27T were text and 3T omni-modal. Mid-training extends the context from 256K to 1M and uses MXFP4 quantization-aware training. For hidden weight matrices the team replaced AdamW with a variant of the Muon optimizer. Embeddings, the language-model head and the MoE router stay on AdamW. According to the authors, mid-training on a broad multimodal corpus was meant to leave room for exploration during RL. RL compute is scaled in three directions. The first is a larger batch with higher throughput. Each step uses 1,568 prompts with 16 rollouts each, about 25K trajectories, and consumes 2.7B to 3.7B training tokens, with context up to 1M. The algorithm is GRPO with asynchronous partial rollouts at a staleness of 4, a learning rate of 3e-6 and gradient clipping at 1.0. The second direction is more diverse and more complex environments: code, general, visual and cyber tasks, run under a mix of agent harnesses. The third is more grader compute. The third direction deserves the closest reading. The paper calls it groupwise agentic grading, and it has two parts. The offline part, Groupwise Reward Synthesis, computes the final reward as the test reward multiplied by a solution score and a behavior score. The online part, Groupwise Advantage Redistribution, ranks the passing trajectories inside each group and shifts positive advantage toward the better solutions. The reasoning is that a pass/fail test signal is too coarse for long-horizon agent work. A model can pass by a long, wandering route and be rewarded the same as a clean one. The authors tested this in a code-only ablation on Flash with batch size 128. Without online grading, the number of turns and the token length grew fast, and gains in pass rate stalled. With online grading, pass-rate gains continued through step 52 and token length grew more gradually. A group-relative length penalty adds to this effect. It lowers the reward of successful rollouts that are long compared with a per-prompt reference length, which pushes the model toward shorter, more token-efficient solutions. The page I read does not give the penalty hyperparameters. Stability is the second theme. During RL the team freezes the MoE router to keep expert loads stable. In large MoE RL runs, small mismatches in routing between the training engine and the inference engine can grow and unbalance the experts. Freezing the router removes one moving part. I did not reach the paper's full section on this, so I quote only its stated conclusion. The authors also built infrastructure for mixed-task agentic RL: a unified trajectory representation, high-concurrency multi-framework rollout, decoupled control and data planes, and training-inference consistency.

Reward hacking is an old problem in agentic RL, and the paper answers it with several layers. First, mid-training alignment data is built from hacking cases the team observed. Second, the environments are cleaned: build logs, residual patches and later Git history are removed. Third, containers are network-isolated. Fourth, a dedicated hack agent searches for leaks until it finds no more exploits. Fifth, trajectories are audited offline during training. Confirmed-hack trajectories get zero reward, and the share of confirmed hacks stayed below 2% throughout training for both models. On cost and results, the paper reports an RL cost of 2.6 million US dollars for Pro and 0.9 million for Flash. For Pro, rollout takes 43.8% of compute, training takes 43.5% and grading takes 12.7%. Grading is therefore a little over a tenth of the bill, yet it sets the quality of the reward. On DeepSWE (average@3), Pro rises from 58.4 to 72.6 and Flash from 48.7 to 65.7 over the RL run. These are gains within training, not a head-to-head comparison with other models. For developers and enterprises, three points stand out. First, the paper offers a recipe that others can study, with numbers, costs and ablations in the open. Second, the authors say they plan to open-source the training dynamics, the RL environments and the RL framework. If they deliver, the cost of reproducing agentic RL drops. Third, a 128-token window with sparse MoE shows that long-context agent training can be kept affordable by architecture, not only by hardware.

The limits must be stated plainly. I read roughly the first half of the paper. I did not see the full benchmark comparison with other models or the authors' own limitations section, so I draw no conclusion about ranking. The title stresses self-improvement, but the parts I read describe human-designed environments, graders and anti-hacking pipelines. How far the model improves itself needs the full text to judge. The RL cost figures cover training only, not data work or environment engineering. And the open-source pledge is not yet fulfilled, so reproducibility is untested. In short, the value of this paper is in its engineering detail: large asynchronous batches, groupwise grading, a frozen router and layered anti-hacking defenses, combined into one working agentic RL pipeline. Watch for two things next: the release of the promised code, and independent evaluations that confirm the DeepSWE gains.

Sources