by Haytham ElFadeel - hfadeelm@gmail.com
July 1, 2025
Abstract
Group Relative Policy Optimization (GRPO) and REINFORCE Leave-One-Out (RLOO) has emerged as an effective and memory-efficient alternative to Proximal Policy Optimization (PPO) for reinforcement learning in large language models. By eliminating the need for a learned value function and instead computing advantages relative to a group of sampled responses, GRPO significantly reduces computational overhead. However, we identify four systematic issues in GRPO's formulation that can degrade training stability and generalization: (1) a response-level length bias introduced by per-response token normalization, (2) a question-level difficulty bias arising from per-prompt standard deviation normalization, (3) a difficulty-dependent sampling bias inherent to group-relative advantage estimation under finite rollout budgets, and (4) instability from unbounded advantage scaling in sparse-reward settings. To address this, we introduce Adaptive Policy Optimization (APO), an unbiased and adaptive optimization method that addresses all four issues. Empirical results on mathematical reasoning benchmarks indicate that APO exhibits improved training stability, token efficiency, and improves the average Pass@1 of GRPO by up to 8%
1. Introduction
Reinforcement learning with verified reward (RLVR) has become a central paradigm for training reasoning-oriented LLMs. Earlier approaches leveraged, Proximal Policy Optimization (PPO; Schulman et al., 2017), requires learning a separate value function to estimate per-token advantages—a significant computational and memory burden at the scale of modern LLMs. Group Relative Policy Optimization (GRPO; Shao et al., 2024) addresses this by sampling a group of G responses per prompt and computing advantages relative to the group's mean reward, entirely removing the critic network. This design has proven effective in practice, particularly for mathematical reasoning tasks such as those targeted by the DeepSeek series of models.
Despite its empirical success, we argue that GRPO's formulation introduces several systematic biases that become increasingly problematic as models and training regimes scale. These biases manifest as observable training artifacts including length distortion, uneven learning across difficulty levels, and training instability in sparse-reward settings.
We identify four distinct issues. First, GRPO normalizes each response's token-level losses by the response's own length , creating a coupling between sequence length and effective gradient magnitude. This favors brevity for correct responses and inadvertently reduces penalties for verbose incorrect responses. Second, GRPO normalizes advantages by the within-prompt reward standard deviation, which causes prompts with low reward variance (very easy or very hard questions) to receive disproportionately large update weights. Third, even without standard deviation normalization, group-relative advantage estimation under finite sampling is inherently biased as a function of prompt difficulty: the algorithm preferentially learns from mid-difficulty prompts while systematically under-learning from hard prompts. We formalize this in Proposition 1, showing that the expected positive advantage signal scales as for small success probability . Fourth, the linear dependence of policy gradients on advantage magnitude, combined with sparse or binary rewards and per-prompt reweighting, leads to heavy-tailed gradient distributions that can cause abrupt policy shifts.
Concurrent work by Liu et al. (2025) identifies the length and standard deviation normalization issues (our Issue 1, and 2) and proposes Dr. GRPO, which removes this normalization. However, Issues 3, and 4 remain unaddressed. As we argue, these issues are not independent: the finite-sampling difficulty bias (Issue 3) persists even after removing variance normalization, and the instability from unbounded advantages (Issue 4) can be exacerbated by difficulty corrections applied to address Issue 3.
To address all four issues in a unified framework, we propose Adaptive Policy Optimization (APO), a method built on three principles: (1) length-invariant gradient aggregation, which replaces per-response normalization with a fixed token budget constant; (2) a history-aware capability anchor that tracks the model's evolving success rate and applies signed, difficulty-aware corrections to advantage weights; and (3) a bounded, asymmetric soft gate that replaces hard ratio clipping with a smooth sigmoid-based trust region, providing robust gradient attenuation for off-policy samples without hard discontinuities.
APO retains the core computational advantages of GRPO—no learned value function, group-based advantage estimation—while correcting the identified biases and improving training stability. Table 1 summarizes the issues addressed by each method.
Table 1. Comparison of issues addressed by PPO, GRPO, Dr. GRPO, and APO.
Issue | PPO | GRPO | Dr. GRPO | APO |
No critic required | ✗ | ✓ | ✓ | ✓ |
Length bias (§3.1) | ✓ | ✗ | ✗ | ✓ |
Std normalization bias (§3.2) | N/A | ✗ | ✓ | ✓ |
Finite-sample difficulty bias (§3.3) | ✗ | ✗ | ✗ | ✓ |
Unbounded advantage instability (§3.4) | ✗ | ✗ | ✗ | ✓ |
Smooth trust region | ✗ | ✗ | ✗ | ✓ |
2. Background
2.1 Prior Work
Foundations. Policy gradient methods for reinforcement learning trace back to REINFORCE (Williams, 1992), which provides unbiased but high-variance gradient estimates. Trust Region Policy Optimization (TRPO; Schulman et al., 2015) introduced constrained optimization to stabilize updates, and Proximal Policy Optimization (PPO; Schulman et al., 2017) simplified this via a clipped surrogate objective that has become the dominant approach for RLHF in large language models. The effectiveness of RL-based post-training with verifiable rewards has been demonstrated at scale by DeepSeek-R1 (Guo et al., 2025a), which achieved significant reasoning improvements across mathematics, coding, and question-answering benchmarks.
Efficient policy optimization without value models. A key limitation of PPO is the requirement for a learned value function, which introduces substantial memory and compute overhead at LLM scale. GRPO (Shao et al., 2024; Guo et al., 2025a) addresses this by estimating advantages relative to a group of sampled responses, eliminating the critic entirely while maintaining strong performance. GPG (Chu et al., 2025) further simplifies the optimization pipeline by removing surrogate losses, critics, and KL constraints. RLOO (Ahmadian et al., 2024) takes a related approach, using leave-one-out baselines within a REINFORCE framework to reduce variance without a value function.
Bias and variance corrections. As group-based and critic-free methods have gained adoption, several works have identified specific failure modes. Dr. GRPO (Liu et al., 2025) identifies and mitigates length bias. DAPO (Yu et al., 2025) employs dynamic sampling strategies to improve rollout quality. OPO (Hao et al., 2025) derives an optimal baseline for group-based estimators to reduce gradient variance. Ahmadian et al. (2024) argue that variance is not a significant concern in LLM RL; we revisit this claim and show that variance remains problematic under limited rollout budgets and long-horizon generation (Section 3.4).
Despite this rapid progress, the systematic biases arising from the core design choices in group-based estimators—including length-dependent normalization, difficulty-dependent signal attenuation under finite sampling, and unbounded advantage scaling—remain largely uncharacterized. Existing corrections address individual symptoms in isolation. In this work, we provide a unified analysis of these interacting failure modes and propose a method that addresses them jointly.
2.2 Preliminaries
We consider the standard RL setting for large language model. Let denote a language model policy parameterized by , and let x denote a prompt sampled from a task distribution . The policy generates a response autoregressively, token by token:
A reward model assigns a scalar reward to each prompt–response pair.
Proximal Policy Optimization (PPO)
PPO (Schulman et al., 2017) optimizes the policy using a clipped surrogate objective. Given a reference (old) policy , the per-token likelihood ratio is:
and the clipped objective is:
where is an advantage estimate typically computed via Generalized Advantage Estimation (GAE) using a learned value function . The requirement to train and store introduces substantial memory and compute overhead, which becomes prohibitive for very large language models.
Group Relative Policy Optimization (GRPO)
GRPO (Shao et al., 2024) eliminates the value function by estimating advantages from a group of sampled responses. For each prompt x, GRPO samples G responses from and computes the group-relative advantage for each response as:
The GRPO objective is then:
where is the length of response i, and controls a KL penalty against a reference policy .
The critical design choices are: (i) advantage normalization by both the mean and standard deviation of within-group rewards, (ii) per-response length normalization by , and (iii) hard clipping of the likelihood ratio. As we show in the following section, each of these choices introduces systematic biases.
3. Issues and biases in GRPO
We now provide a detailed analysis of four systematic issues in GRPO's formulation. While the length and standard deviation normalization issue has been partially noted in concurrent work (Liu et al., 2025), we provide a unified treatment and identify additional biases that remain unaddressed even after removing the variance normalization.
3.1 Response-Level Length Bias from Per-Response Normalization
GRPO's surrogate objective averages each response's token-level losses by dividing by the response's own length . This normalization, absent in PPO's formulation, creates a coupling between the length of a generated sequence and the effective magnitude of its gradient contribution.
Mechanism. Consider the per-response contribution to the objective:
Since is a trajectory-level (response-level) quantity, the factor directly scales the effective weight of each trajectory in the gradient. For tokens near the on-policy regime (), the per-response contribution is approximately after the inner sum, but the normalization breaks the proportionality between the number of tokens contributing gradient signal and the trajectory weight. Concretely, this introduces two asymmetric effects:
- For positive advantages (): Shorter correct responses receive a larger per-trajectory update, since the same advantage is divided by a smaller denominator. This biases the model toward brevity for correct solutions.
- For negative advantages (): Longer incorrect responses receive a smaller per-trajectory penalty, since the negative signal is diluted across more tokens. This effectively reduces the penalization of verbose incorrect responses, potentially encouraging verbosity among failures.
Consequences. Over the course of training, this asymmetry can cause the model to conflate brevity with correctness and length with permissibility of errors. In reasoning tasks, this is particularly problematic: correct solutions that require extended chain-of-thought reasoning are under-reinforced relative to short correct answers, while long incorrect reasoning chains are insufficiently penalized. This creates a confound between genuine "reasoning quality" improvements and optimization artifacts driven by length bias.
3.2 Question-Level Difficulty Bias from Per-Prompt Standard Deviation Normalization
GRPO normalizes the centered reward by the within-prompt reward standard deviation:
This normalization is intended as a stabilization trick analogous to batch normalization in supervised learning. However, it introduces a difficulty-dependent bias.
Mechanism. For prompts where the model almost always succeeds or almost always fails, the within-group reward variance is near zero. Division by a near-zero standard deviation inflates the advantage magnitudes for these prompts, giving them disproportionately large influence on the gradient update. Conversely, prompts at intermediate difficulty—where the model produces a mix of correct and incorrect responses—have higher reward variance and thus receive attenuated advantages.
Consequences. The optimization landscape is implicitly reweighted by a function of prompt difficulty that bears no relation to the prompt's importance for learning. Very easy prompts, where the model already succeeds reliably, receive outsized gradient contributions despite offering minimal learning signal. Very hard prompts, where the model rarely succeeds, similarly dominate updates—but the signal from these prompts is noisy due to limited positive samples. Meanwhile, prompts at the frontier of the model's capability, which arguably carry the highest marginal value for learning, are systematically down-weighted. This distorts curriculum dynamics and can impair generalization.
Liu et al. (2025) identify this standard deviation normalization issue and propose removing it in Dr. GRPO. However, as we show next, simply removing the standard deviation is insufficient: the group-relative advantage estimator itself exhibits a difficulty-dependent bias under finite sampling that persists regardless of variance normalization.
3.3 Difficulty-Dependent Bias in Group-Relative Advantages Under Finite Sampling
Even without standard deviation normalization, the group-relative advantage estimator exhibits a systematic bias as a function of prompt difficulty when the group size G is finite and rewards are sparse or binary. This is a fundamental property of the estimator that cannot be resolved by simply removing the variance scaling.
Mechanism. For a given prompt, let denote the model's current probability of producing a correct response. With binary rewards () and sampled responses, the number of successes follows a binomial distribution . The total positive learning signal from a group—the sum of positive advantages—is proportional to . This quantity is maximized at and vanishes as or : the learning signal is a quadratic function of prompt difficulty, peaking at mid-difficulty and decaying symmetrically toward the extremes.
Implications:
- Hard prompts (): The group is almost always all-failures, producing and near-zero advantages for all responses. The rare event of a single success (probability for small ) is the only source of positive signal. The model systematically under-learns from these prompts.
- Easy prompts (): By symmetry, the group is almost always all-successes with near-zero advantages. However, the occasional failure generates a concentrated negative signal, potentially causing the model to over-attend to noise.
- Mid-difficulty prompts (): Mixed-outcome groups are frequent, providing consistent bidirectional learning signal.
3.4 Instability from Unbounded Advantage Scaling
In sequence-level policy optimization, the gradient is weighted linearly by the advantage magnitude. In settings with sparse or binary rewards, this can produce heavy-tailed or sharply bimodal gradient distributions.
Mechanism. After any per-prompt reweighting (including difficulty corrections that might be applied to address Issue 3), the effective advantage can take on extreme values, particularly when:
- Rewards are binary, producing bimodal advantage distributions within each group.
- Prompts vary widely in difficulty, and reweighting amplifies signals from rare-outcome groups.
- The policy is non-stationary, causing likelihood ratios to drift substantially from 1.
- Rollout budgets are limited, increasing variance in group-level statistics.
The hard clipping mechanism in PPO and GRPO addresses ratio drift but does not bound the gradient contribution from large advantages. Moreover, the hard clip introduces a discontinuity in the gradient that can cause optimization artifacts at the clip boundary.
Consequences. A small number of outlier samples can dominate the gradient, leading to large variance in updates, sensitivity to reward scaling, and susceptibility to abrupt policy shifts including mode collapse or reward hacking. Negative-dominated batches—common under sparse rewards where most responses are incorrect—can trigger overly aggressive avoidance updates that destabilize training. This issue is particularly acute for Mixture-of-Experts (MoE) architectures, where gradient spikes can disproportionately affect individual experts.
Contrary to prior claims (Ahmadian et al., 2024) that variance is not a significant concern in LLM RL, we argue that variance remains a critical issue under limited rollout budgets and long-horizon generation—precisely the regime where GRPO is most commonly deployed.
3.5 Interaction Between Issues
The four issues identified above are not independent. The standard deviation normalization (Issue 2) was partially intended to mitigate the finite-sampling difficulty bias (Issue 3) by equalizing advantage scales across prompts, but it does so in a way that introduces its own pathology (inflating signals for extreme-difficulty prompts). Removing the std normalization (as in Dr. GRPO) exposes the raw finite-sampling bias. Any correction for the finite-sampling bias (e.g., difficulty reweighting) can amplify advantage magnitudes, exacerbating Issue 4. And the length bias (Issue 1) operates orthogonally, distorting trajectory weights regardless of how advantages are computed. A principled solution must address all four issues jointly, which motivates the unified design of APO.
4. Adaptive Policy Optimization
We now present Adaptive Policy Optimization (APO), a method that addresses the four issues identified above through three complementary mechanisms while retaining GRPO's core advantage of eliminating the learned value function. Full pseudocode is provided in Appendix A.
4.1 Overview
APO is built around three design principles, each targeting one or more of the identified issues:
- Length-invariant gradient aggregation (addresses Issue 1). We replace the per-response normalization with a constant normalization factor tied to a fixed token generation budget , restoring invariance between sequence length and trajectory weight.
- History-aware difficulty correction (addresses Issues 2 and 3). We maintain a slowly-updated capability anchor that tracks the model's recent success rate across training. For each prompt, we compare its observed group success rate to and apply a signed difficulty correction that amplifies learning on prompts harder than the model's current capability while tempering updates on easier prompts. This replaces the per-prompt std normalization with a principled, sign-aware reweighting.
- Bounded, asymmetric advantage shaping (addresses Issue 4). We replace hard ratio clipping with a smooth sigmoid-based soft gate that provides continuous trust-region behavior: near on-policy updates are preserved at full strength, moderately off-policy updates are smoothly down-weighted, and extremely off-policy updates are heavily attenuated. This simultaneously bounds the effective gradient contribution per sample and eliminates the discontinuity introduced by hard clipping.
4.2 Data and Rollout
Let prompts . For each prompt x, we sample a group of G responses from the current behavior policy:
with maximum generation length . Each response receives a scalar reward . We define per-token likelihood ratios:
4.3 Centered Group Advantage Without Variance Normalization
We compute centered but not standard-deviation-normalized advantages:
.
By removing the normalization, we eliminate the difficulty-dependent inflation described in Section 3.2. The remaining difficulty-dependent bias from finite sampling (Section 3.3) is addressed by the capability anchor mechanism below.
4.4 History-Aware Capability Anchor
To correct for the finite-sampling difficulty bias, we maintain a scalar capability estimate that tracks the model's recent overall success rate.
Batch success rate. At each training iteration, we compute the observed success rate over all prompts and responses:
where is the number of prompts, the number of responses per prompt, and is a task-specific success indicator (identity for binary rewards; a monotone bounded transform otherwise).
Anchor update. The capability anchor is updated via exponential moving average with an adaptive rate:
The adaptive rate is designed to be more conservative when the anchor estimate is volatile. We maintain a history buffer H of recent anchor values (length m) and compute:
where is a base rate scale hyperparameter. This ensures the anchor tracks genuine capability changes while remaining stable under noisy estimates.
Motivation. The capability anchor serves as a global reference point that is absent from GRPO's per-prompt advantage computation. By comparing each prompt's observed success rate to the model's overall capability, we obtain a principled measure of relative difficulty that does not depend on the within-group reward distribution. This directly addresses the finite-sampling bias identified in Section 3.3: hard prompts (low relative to ) receive boosted positive signals.
4.5 Difficulty-Aware Reweighting
For each prompt, we compute the group success estimate and compare it to the capability anchor:
where is the success indicator for response .
We define a signed direction term that depends on both the advantage sign and the difficulty gap:
and the difficulty weight:
Interpretation. The sign structure of ensures appropriate behavior across all four quadrants of the difficulty–advantage space:
Prompt difficulty | Advantage sign | Effect on | Rationale | |
Hard () | Positive () | (boost) | Amplify rare correct signals on hard prompts | |
Hard () | Negative () |
(temper) | Avoid redundant penalization of expected failures | |
Easy () | Positive () | (dampen) | Reduce over-exploitation of easy prompts | |
Easy () | Negative () |
(amplify) | Maintain quality by learning from rare failures |
The effective advantage is then:
Since , the sign of is always preserved.
4.6 Soft-Gated Ratio Objective
Rather than the hard clip used in PPO and GRPO, we employ a smooth sigmoid-based gate that provides continuous trust-region behavior. This addresses Issue 4 (unbounded advantage scaling) while also eliminating the gradient discontinuity at the clip boundary.
Asymmetric temperature. We select a temperature based on the sign of the effective advantage:
Soft gate function. The gate is defined as:
where is the logistic sigmoid.
Properties. We analyze the behavior of the soft gate in detail in Appendix B. The key properties are:
- Unit sensitivity at on-policy: The derivative , ensuring that near on-policy updates are unmodified.
- Smooth decay: The effective weight decays smoothly as departs from , with the rate of decay controlled by .
- Bounded output: Unlike the linear or clipped-linear dependence on in PPO/GRPO, is bounded in , preventing any single token from contributing unbounded gradient.
The asymmetric temperatures () allow differential treatment of positive and negative updates. Setting provides a wider effective trust region for negative updates, which is beneficial under sparse rewards where negative advantages dominate and overly aggressive avoidance updates are a primary source of instability.
4.7 Length-Invariant Objective
The final APO objective combines all three mechanisms:
The key design choices embodied in this objective are:
- Constant normalization replaces the per-response , ensuring that the effective weight of a trajectory is independent of its realized length. This directly addresses Issue 1 (Section 3.1).
- No within-prompt standard deviation normalization in the advantage computation, addressing Issue 2 (Section 3.2).
- Difficulty-aware effective advantages that correct for finite-sampling bias via the capability anchor, addressing Issue 3 (Section 3.3).
- Smooth trust-region enforcement via the soft gate , replacing hard ratio clipping and bounding gradient contributions, addressing Issue 4 (Section 3.4).
4.8 Algorithm Summary
The complete APO training procedure operates as follows:
- Sample prompts from the training distribution .
- Group rollouts: For each prompt, sample responses from with maximum length .
- Reward evaluation: Compute for each response.
- Update capability anchor: Compute the batch success statistic and update using the adaptive rate .
- Compute centered group advantages: (no standard deviation normalization).
- Compute difficulty gap per prompt: .
- Compute difficulty weights: .
- Compute effective advantages: .
- Optimize for epochs over minibatches using the soft-gated objective with normalization.
- Synchronize: Set .
4.9 Hyperparameters
Hyperparameter | Meaning | Suggested Range | Notes |
G | Responses per prompt | 4–16 | 8 is a strong default for cost/variance tradeoff |
Constant token budget | 512–32,768+ | Set to match the rollout maximum generation length | |
Prompts per iteration | 64–2048 | Larger reduces variance; limited by throughput | |
Epochs per iteration | 1–4 | Keep small to limit drift from θold\theta_\text{old}
θold | |
Positive gate temperature | 0.6–1.2 | Lower = narrower trust region for positives | |
Negative gate temperature | 0.8–2.0 | Typically τ−>τ+\tau_- > \tau_+
τ−>τ+ to soften negative dominance | |
Difficulty weight scale | 1.1–2.0 | Larger values reallocate more learning to hard prompts | |
Anchor history window | 10–50 | Larger = smoother anchor; smaller = more adaptive | |
Anchor rate scale | 0.1–2.0 | Chosen so that ηt\eta_t
ηt rarely saturates at 1 |
5. Experiments
5.1 Implementation Details
Model. We evaluate the algorithm’s performance on reasoning tasks. Following Dr.GRPO (Liu et al., 2025), we use Qwen2.5-Math1.5B (Yang et al., 2024), and Qwen2.5-Math-7B as our base models to assess performance on mathematical tasks.
Training. Following the setup of Dr.GRPO (Liu et al., 2025), we use MATH (Hendrycks et al., 2021) Levels 3–5 as the training dataset for models under 7B, which contains 8,523 mathematical problems. For each question, we generate 8 rollouts and cap the model’s maximum response length at 3,000 tokens. During each RL training round, the old policy produces 1,024 rollouts, and the current policy is updated 8 times with a batch size of 128.
Evaluation. We evaluate our method on five mathematical reasoning benchmarks of varying difficulty following Dr.GRPO (Liu et al., 2025): AIME24, which consists of 30 high-school level olympiad problems from the American Invitational Mathematics Examination 2024; AMC, containing 83 intermediate difficulty multiple-choice problems; MATH500, a subset of 500 problems from the original MATH dataset covering algebra, geometry, and number theory; Minerva (Lewkowycz et al., 2022), featuring 272 graduate-level problems requiring multi-step reasoning; and Olympiad Bench (He et al.,2024), a collection of 675 high-difficulty olympiad problems. These benchmarks collectively cover a broad spectrum of problem types and difficulty levels.
5.2 Main Results
Model | AIME24 | AMC | MATH500 | Minerva | Olympiad Bench | Avg. |
1.5B - GRPO | 23.3 | 49.4 | 75.2 | 25.7 | 39.0 | 42.5 |
1.5B - Dr. GRPO | 23.0 | 52.1 | 77.5 | 30.1 | 38.5 | 44.2 |
1.5B - APO (Ours) | 25.4 | 54.3 | 78.9 | 31.3 | 40.2 | 46.0 |
7.0B - GRPO | 40.1 | 59.0 | 83.4 | 32.4 | 41.3 | 51.3 |
7.0B - Dr. GRPO | 43.3 | 62.7 | 80.7 | 33.5 | 42.9 | 52.6 |
7.0B - GMPO-7B (Ours) | 45.5 | 65.6 | 84.9 | 35.1 | 43.5 | 54.9 |
APO demonstrates consistent improvements across different base models. We also observed broad and consistently stable learning compared with GRPO and Dr. GRPO, both of which experience early training collapse. Additionally APO demonstrated similar token efficiency to Dr. GRPO and significantly higher token efficiency compare to GRPO.
6. Discussion
Computational overhead. APO introduces minimal computational overhead relative to GRPO. The capability anchor update requires maintaining a single scalar and a short history buffer . The difficulty weight computation adds scalar operations per iteration. The soft gate replaces the hard clip with a sigmoid evaluation, which has equivalent cost on modern hardware. The overall training cost remains dominated by the forward and backward passes through the language model, as in GRPO.
Choice of . The constant normalization factor should be set to the maximum generation length used during rollout. This ensures that the normalization is consistent across all responses: a response that uses the full token budget receives the same normalization as one that terminates early. The choice does not affect relative weighting between responses (since is constant), but it does affect the absolute scale of the objective, which interacts with the learning rate. In practice, setting equal to the maximum generation length and adjusting the learning rate accordingly is sufficient.
7. Limitations
We acknowledge several limitations of the current work. First, the capability anchor assumes a single global notion of "model capability," which may be overly simplistic when the training distribution contains heterogeneous task types with different difficulty profiles. Extending the anchor to maintain per-domain or per-task capability estimates is a natural direction for future work.
Second, the difficulty-aware reweighting relies on a success indicator that maps rewards to ; for tasks with continuous or multi-dimensional rewards, the choice of this mapping introduces a design decision that may require tuning.
Third, while we provide theoretical analysis of the biases in GRPO and the corrective properties of APO's components, a formal convergence analysis of the full APO objective under standard assumptions remains an open problem. Finally, the soft gate's asymmetric temperatures (, ) add two hyperparameters relative to the single clip parameter in PPO/GRPO; while we find these to be robust within the suggested ranges, further study of their sensitivity across diverse tasks would be valuable.
8. Conclusion
We identified four systematic issues in Group Relative Policy Optimization (GRPO) that introduce biases related to response length, prompt difficulty, finite-sample group estimation, and unbounded advantage scaling. These biases distort learning dynamics by conflating sequence length with gradient magnitude, creating implicit and unprincipled curriculum effects, under-weighting informative hard-prompt signals, and permitting catastrophic gradient spikes under sparse rewards. We showed that concurrent fixes addressing only the standard deviation normalization (Dr. GRPO) leave three of these four issues unresolved, and that the finite-sampling difficulty bias is a fundamental property of the group-relative estimator.
We proposed Adaptive Policy Optimization (APO), a method that corrects all four issues through three targeted mechanisms: constant-budget length normalization, a history-aware capability anchor with signed difficulty reweighting, and a smooth sigmoid-based soft gate replacing hard ratio clipping. Each mechanism addresses specific failure modes while preserving GRPO's core advantage of eliminating the learned value function. Together, they yield a bias-reduced, difficulty-aware, and stable training objective for group-based policy optimization of large language models.
References
- Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal Policy Optimization Algorithms. arXiv preprint arXiv:1707.06347.
- Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Zhang, M., Li, Y., Wu, Y., & Guo, D. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv preprint arXiv:2402.03300.
- Zichen Liu, Changyu Chen, Wenjun Li, Penghui Qi, Tianyu Pang, Chao Du, Wee Sun Lee, and Min Lin. (2025). Understanding R1-Zero-Like Training: A Critical Perspective. arXiv preprint arXiv:2503.20783.
- Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang,Xiao Bi, et al. (2025). Deepseek-R1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948 (2025).
- Ahmadian, A., Cremer, C., Gallé, M., Fadaee, M., Kreutzer, J., Pietquin, O., Üstün, A., & Hooker, S. (2024). Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL).