Command Palette
Search for a command to run...
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong Liang Zhao Zihao Yue Wenhan Ma Linghao Zhang Lei Li Shicheng Li Yifan Song Bowen Ye Fuli Luo
Abstract
Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking diferences in implementation quality and adherence to task requirements. This leaves the policy without a learning signal that favors clean, targeted implementations over those containing unnecessary or out-of-scope changes. We introduce Gagar, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic sampling that retains groups containing both passing and failing trajectories, Gagar places all trajectories from each group in a shared workspace, where an SFT-trained agentic grader jointly inspects them and ranks the test-passing candidates. Based on this ranking, we downweight lower-ranked trajectories and proportionally rescale the advantages of all test-passing trajectories to restore their original sum. This sum-preserving redistribution retains the relative weights established by quality-based downweighting while shifting credit toward higher-quality implementations. We evaluate Gagar at industrial scale using pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters). Controlled code-only Flash experiments show improved code agent performance, reduced trajectory-length growth, and more stable training. We further apply Gagar in large-scale mixed-task RL with both Flash and Pro. Our results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.
One-sentence Summary
Researchers from Xiaomi, Renmin University of China, Peking University, and other institutions propose Gagar, a quality-aware credit redistribution framework for code agent RL that uses dynamic sampling and an SFT-trained agentic grader to rank test-passing trajectories, downweights lower-ranked trajectories, and applies sum-preserving advantage rescaling; industrial-scale experiments with MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters) show improved code agent performance, reduced trajectory-length growth, and more stable training.
Key Contributions
- Gagar is a quality-aware credit redistribution framework for code agent RL that uses dynamic sampling to retain groups with both passing and failing trajectories, then uses an SFT-trained agentic grader to inspect and rank test-passing trajectories in a shared workspace.
- Gagar downweights lower-ranked passing trajectories and proportionally rescales the advantages of all passing trajectories, preserving the original total positive advantage while shifting credit toward higher-quality implementations.
- Controlled code-only experiments with MiMo-V2.6-Flash show improved code agent performance, reduced trajectory-length growth, and more stable training; mixed-task RL with MiMo-V2.6-Flash and MiMo-V2.6-Pro reaches DeepSWE v1.1 avg@3 scores of 67.9 and 71.9.
Introduction
Reinforcement learning with executable feedback is a scalable way to train code agents that inspect repositories, modify code, and validate changes over long interactions. In common group-relative setups such as GRPO, all test-passing trajectories receive the same positive outcome advantage, so the training signal cannot distinguish cleaner, more focused patches from overly complex or risky implementations. Static text-based assessment also tends to miss repository context and execution evidence. The authors introduce Gagar, a quality-aware framework that uses an agentic grader to inspect code, run targeted checks, and rank test-passing candidates within each rollout group, then redistributes credit among successful trajectories while preserving their total positive advantage. This approach is designed to reinforce precise, minimally invasive, merge-ready solutions and is validated at industrial scale with MiMo models.
Method
The authors leverage groupwise quality supervision to augment reinforcement learning for code agents, moving beyond binary task outcomes. As shown in the figure below, the framework integrates groupwise agentic grading into the RL training loop to evaluate and rank valid passing implementations.
In the training setup, for each coding task, the rollout policy generates multiple trajectories. The final patch from each trajectory is evaluated using executable tests to yield a binary outcome reward. Let Gx={τi}i=1n denote the group of valid trajectories. The authors use the mean-centered outcome advantage Ai=Ri−Rˉ, where Rˉ=n−1∑j=1nRj. To ensure both successful and failed trajectories are present, dynamic sampling is adopted so that 0<Rˉ<1. All passing trajectories initially receive the same positive advantage 1−Rˉ, which fails to differentiate implementation quality.
To address this, the authors introduce groupwise agentic grading. The grader jointly examines all rollouts in a shared workspace containing the task specification, repository, complete trajectories, submitted patches, and test outputs. This comparison helps identify the strongest solutions and exposes ineffective strategies. The agentic grader gathers evidence iteratively, reviewing turn-by-turn summaries, reading relevant trajectory portions, and cross-checking patches against repository code and test logs. It can also run targeted checks. Quality is assessed across five criteria: approach suitability, precision, minimality, side effects, and codebase consistency. Passing solutions are checked for leaked or external answers, and confirmed hacks receive zero reward. The remaining candidates are scored and mapped to three quality tiers, which determine a discount factor fi∈(0,1] on their positive advantage.
Following grading, sum-preserving advantage redistribution is applied. Simply applying the discount factors would yield A~i=fiAi for passing trajectories, removing positive credit without a corresponding change on the negative side. This deficit weakens the reinforcement of successful trajectories. To preserve the total positive advantage S+=∑i∈PAi assigned to passing trajectories, the authors apply a common rescaling factor:
λ=∑j∈PfjAjS+The redistributed advantage is then:
Ai⋆={λfiAi,Ai,i∈P,i∈F.This ensures the total positive advantage is restored and the change is zero-sum over the passing subset, preserving both the ordering and relative strength of quality preferences.
During online training integration, grading runs asynchronously with rollout collection. The system checks the validity of grading results, falling back to original outcome advantages if necessary. Confirmed reliance on external solutions resets the affected reward to zero before group statistics are recomputed. The resulting sequence-level advantages are broadcast to the model-generated response tokens to supervise the RL policy update.
Experiment
These experiments evaluate Gagar on code-only reinforcement learning with MiMo-V2.6-Flash and on industrial-scale mixed-task RL with Flash and Pro, using DeepSWE v1.1 and SWE-bench Pro as benchmarks. The main comparisons show that Gagar improves long-horizon code agent performance, keeps trajectories shorter and more efficient, and yields higher implementation quality in external model review. An ablation confirms that sum-preserving redistribution is important for training stability, since downweighting passing trajectories without restoring credit leads to unstable dynamics and degraded downstream behavior. Industrial-scale runs further demonstrate competitive code agent performance relative to larger frontier models.
After industrial-scale mixed-task RL with Gagar, MiMo-V2.6-Flash and MiMo-V2.6-Pro achieve coding benchmark results competitive with strong external baselines. MiMo-V2.6-Pro outperforms Kimi K3 on DeepSWE despite having substantially fewer total parameters and surpasses GPT-5.6 Sol on SWE-bench Pro. Its DeepSWE score approaches GPT-5.6 Sol and Claude Opus 5, while it trails Claude Opus 5 on SWE-bench Pro. MiMo-V2.6-Pro outperforms Kimi K3 on DeepSWE despite using substantially fewer total parameters. MiMo-V2.6-Pro surpasses GPT-5.6 Sol on SWE-bench Pro. MiMo-V2.6-Pro approaches GPT-5.6 Sol and Claude Opus 5 on DeepSWE but remains below Claude Opus 5 on SWE-bench Pro. MiMo-V2.6-Flash achieves DeepSWE and SWE-bench Pro scores close to larger external models.
This experiment evaluates MiMo-V2.6-Flash and MiMo-V2.6-Pro after industrial-scale mixed-task RL with Gagar on coding benchmarks. The results show that MiMo-V2.6-Pro is competitive with strong external models, outperforming Kimi K3 on DeepSWE despite using substantially fewer parameters and surpassing GPT-5.6 Sol on SWE-bench Pro, while approaching GPT-5.6 Sol and Claude Opus 5 on DeepSWE but remaining below Claude Opus 5 on SWE-bench Pro. MiMo-V2.6-Flash also achieves DeepSWE and SWE-bench Pro scores close to larger external models.