by Haytham ElFadeel - hfadeelm@gmail.com
Feb 2, 2025
Abstract
Reinforcement learning from verifiable rewards (RLVR) has emerged as a powerful paradigm for improving the reasoning capabilities of large language models (LLMs). However, outcome reward models (ORMs) provide only a single binary signal at the end of a complete reasoning trajectory, forcing all intermediate steps and tokens to share identical credit or blame regardless of their individual contribution. This sparse reward structure creates a fundamental credit assignment problem: the gradient signal is either discounted into oblivion for early tokens or diluted uniformly across all tokens, obscuring which actions actually mattered. Process reward models (PRMs) attempt to address this by providing per-step feedback, but existing approaches either rely on expensive human annotations that do not scale, or on Monte Carlo (MC) estimation methods that conflate future outcome potential with current-step correctness, leading to noisy and often misleading supervision signals.
We propose a hierarchical credit assignment framework that operates at three levels of granularity. At the step level, we introduce a progress reward that measures the change in probability of reaching a correct answer before and after each reasoning step, as evaluated by an independent evaluator policy. Unlike traditional PRMs that assess step correctness in isolation, this progress signal is grounded in actual outcome probabilities and captures whether a step made the problem more solvable. At the token level, we introduce an entropy-weighted advantage that uses the entropy of the policy's output distribution at each token position as a multiplicative modulator on the step-level advantage. The combined effective reward integrates outcome, progress, and token-level signals into a single coherent framework. We present the complete formulation, data collection procedures, model training methodology, and reinforcement learning integration. Experimental results are forthcoming.
1. Introduction
Large language models trained with reinforcement learning have achieved remarkable results across mathematical reasoning, code generation, and other structured problem-solving domains (Ouyang et al., 2022; Lightman et al., 2023; Shao et al., 2024). A central paradigm in this line of work is reinforcement learning from verifiable rewards (RLVR), where the model generates a complete chain-of-thought response and receives a binary outcome reward indicating whether the final answer is correct. This approach has proven effective when combined with algorithms such as GRPO (Shao et al., 2024), RLOO, or standard PPO. However, outcome reward models impose a fundamental limitation: they provide no information about which intermediate actions contributed to or detracted from the final result. Consider an LLM generating a 1,000-token chain-of-thought response to a mathematics problem. The model receives a single binary signal at the end. Standard temporal-difference methods face a dilemma when propagating this terminal reward backward through the trajectory.
1.1. The Credit Assignment Problem
The Discount Dilemma. With a typical discount factor (e.g., γ = 0.99), the effective credit arriving at the first token is γ⁹⁹⁹ ≈ 0.00043 of the terminal reward. Early decisions that establish the entire reasoning strategy receive negligible gradient signal. The reward effectively vanishes before reaching the tokens that matter most.
The Dilution Dilemma. The natural response is to set γ = 1, which is exactly what modern LLM RL methods do (GRPO, RLOO, and similar algorithms typically use undiscounted returns). This preserves signal magnitude but introduces reward dilution: by indiscriminately assigning identical credit to all 1,000 tokens for a single binary outcome, we obscure the causal link between early decisions and the final result. Every token—filler words, formatting characters, genuinely critical reasoning steps—receives the same advantage estimate. The result is high-variance gradient estimates that wash out the signal from the few tokens that actually mattered.
This is the core tension in LLM credit assignment: discount too aggressively and the signal vanishes; discount too little and the signal gets diluted across irrelevant actions. Neither extreme solves the credit assignment problem. What is needed are methods that can identify which actions mattered, not merely propagate a blanket signal backward.
1.2. Limitations of Existing Approaches
Process Reward Models (PRMs) attempt to provide finer-grained supervision by scoring each reasoning step independently. However, existing PRMs face two fundamental challenges. First, the labeling problem: human-annotated PRMs (Lightman et al., 2023) are expensive and do not scale, while automated methods based on Monte Carlo estimation (Wang et al., 2024) introduce systematic biases we discuss in Section 2. Second, and more fundamentally, step correctness is not the right signal for credit assignment. A step can be perfectly correct yet make zero progress toward the answer (e.g., restating the problem), and a step can appear unconventional yet represent the key insight that unlocks the solution. What matters for RL training is not whether a step is correct in isolation, but whether it brought the model closer to solving the problem.
1.3. Our Contributions
We propose a hierarchical credit assignment framework that addresses the limitations of both ORMs and traditional PRMs through three complementary mechanisms. First, we introduce a progress reward that measures the change in success probability before and after each step under an independent evaluator policy, providing a grounded, automated, and scalable per-step signal. Second, we propose an entropy-weighted token advantage that modulates the step-level signal at the individual token level based on decision entropy, directing more credit to tokens where the model made genuine choices. Third, we provide a unified effective reward formulation that integrates outcome, progress, and token-level signals, along with complete algorithms for data collection, model training, and RL integration.
2. Prior Work
2.1. Outcome Reward Models
Outcome reward models assign a single scalar score to an entire reasoning trajectory based on whether the final answer is correct (Cobbe et al., 2021). Given a mathematical problem p and a solution s, the ORM is trained with a binary cross-entropy loss to predict answer correctness. ORMs are straightforward to train since labels can be obtained automatically by comparing the model's final answer against a ground truth. This has made them the dominant approach in RLVR pipelines for mathematical reasoning.
However, ORMs suffer from a fundamental limitation: they provide only trajectory-level feedback. All steps in a correct trajectory receive identical positive signal, and all steps in an incorrect trajectory receive identical negative signal. When used as rewards in RL, this sparse signal leads to the credit assignment challenges described above. With enough rollouts per question, statistical averaging across trajectories can partially resolve step-level differences—if step 3 is consistently the point of failure across many incorrect rollouts, the averaged advantage for step 3 will eventually be lower than for other steps. But this requires a large number of rollouts per question to achieve reasonable variance reduction, making ORM-based RL sample-inefficient.
2.2. Process Reward Models
Process reward models provide per-step feedback by scoring each intermediate reasoning step individually (Uesato et al., 2022; Lightman et al., 2023). In principle, PRMs offer substantially richer signal than ORMs by identifying which specific steps are problematic rather than merely whether the overall trajectory succeeded.
Human-Annotated PRMs. Lightman et al. (2023) introduced PRM800K, a dataset of human-labeled step-level correctness annotations for mathematical reasoning. Human annotators assessed each step as correct, incorrect, or neutral, providing high-quality process supervision. Training PRMs on this data demonstrated significant improvements over ORMs in best-of-N selection for the MATH benchmark. However, collecting such annotations requires annotators with strong mathematical expertise and involves elaborate quality control procedures, making it prohibitively expensive to scale to large and diverse problem sets.
2.3. Monte Carlo Estimation for Automated Process Supervision
To address the scalability limitation of human annotation, researchers have explored automated methods for generating process supervision labels, most prominently through Monte Carlo (MC) estimation (Wang et al., 2024; Xiong et al., 2024; Luo et al., 2024). The core idea, popularized by Math-Shepherd (Wang et al., 2024), is as follows: given a reasoning trajectory with steps , the label for step is estimated by generating completions from the prefix using a completion model, scoring each completion against the ground truth, and computing the fraction that reach the correct answer. If any completion succeeds, the step is labeled positive (hard label), or the fraction of successes is used directly (soft label).
While this approach is fully automated and scalable, Zhang et al. (2025) identify several critical limitations through extensive experiments:
Conflation of correctness with outcome potential. MC estimation does not measure whether a step is correct; it measures the probability that a completion model can reach the correct answer from the current prefix. These are fundamentally different quantities. A completion model may generate a correct final answer despite an incorrect intermediate step (by implicitly correcting the error in subsequent steps), or fail to reach the correct answer despite a perfectly valid step (due to the completion model's own limitations). Zhang et al. (2025) draw a clear distinction between PRMs, which should function as deterministic evaluators of current-step correctness, and value models, which estimate future solution potential. MC estimation inherently trains the latter while claiming to produce the former, introducing systematic noise into the supervision signal.
Sensitivity to the completion model. The quality of MC-estimated labels depends critically on the completion model used for rollouts. If the completion model is too weak, most rollouts will fail regardless of step quality, producing near-zero labels even for correct steps. If it is too strong, most rollouts will succeed regardless, producing uniformly high labels. The completion model must be calibrated to the difficulty of the problems, creating a dependency that is difficult to manage in practice.
Inferior generalization. Zhang et al. (2025) demonstrate empirically that PRMs trained via MC estimation yield significantly inferior performance in step-wise error identification compared to PRMs trained on human-annotated or LLM-as-a-judge data, even when the MC-estimated training set is substantially larger. On the ProcessBench benchmark, their MC-estimated models achieved an average F1 of 40.2%, compared to 56.5% for a model trained on the much smaller PRM800K dataset with human labels. This performance gap persists despite careful tuning of both hard and soft label variants.
2.4. The Correctness–Progress Distinction
A deeper issue, which we argue is more fundamental than the technical limitations of MC estimation, is that step correctness itself is not the optimal signal for reinforcement learning. Consider three types of reasoning steps: (1) a step that is mathematically correct but merely restates the problem without advancing toward the solution; (2) a step that is correct and makes substantial progress by establishing a key relationship; and (3) a step that uses unconventional notation but contains the critical insight that leads to the answer. A correctness-based PRM assigns high scores to both (1) and (2) despite their vastly different contributions to solving the problem, and may penalize (3) despite its decisive role.
What RL training requires is not a judgment of whether each step is correct in isolation, but a measure of whether each step made the problem more solvable (a notion of progress). This observation motivates our progress-based approach, which we develop in the following section.
3. Approach
We present our approach to hierarchical credit assignment for LLM reinforcement learning. We first define the progress reward at the step level (Section 3.1), then introduce token-level entropy weighting (Section 3.2), and finally derive the unified effective reward formulation (Section 3.3).
3.1 Progress Reward
Consider an LLM generating a multi-step reasoning trajectory in response to a question . We define the state as the prefix after h reasoning steps, with denoting the initial state (question only).
Let denote a evaluator policy—a language model distinct from the base policy being trained. The evaluator serves as an independent evaluator of reasoning progress. We define:
The advantage measures the progress made by step : the change in the evaluator's probability of reaching a correct answer before and after taking the step. A positive value indicates the step made the problem more solvable; a negative value indicates it made the problem harder; a value near zero indicates the step was neutral (regardless of whether it was technically correct).
Why a separate evaluator policy? A critical design choice is that progress must be measured under a policy that is distinct from the base policy being trained. If we were to measure progress under the base policy itself, , the resulting signal would be redundant with what the outcome reward already provides. Since is exactly the expected outcome under the base policy's own continuations, and is already an unbiased Monte Carlo estimate of , adding to the reward reduces to a rescaled version of standard policy gradient with outcome reward. No new information is introduced. In contrast, a separate evaluator provides an independent assessment of step quality that is not already captured by the outcome reward, enabling genuinely improved credit assignment.
Evaluator policy. We define the evaluator policy to be best-of-K (BoK) policy derived from the base policy: sample K completions from π and return a correct one if any exists. With K in the range of 4–8, this creates a evaluator that is modestly stronger than the base policy and naturally complementary to it.
3.2. Token-Level Entropy Weighting
The progress reward operates at the step level, assigning a single scalar to each reasoning step. However, within a step, not all tokens contribute equally. Some tokens are near-deterministic (formatting, common phrases like "therefore" or "we have"), while others represent genuine decision points (choosing a specific numerical value, selecting a solution method, picking a variable name). We propose using the entropy of the policy's output distribution at each token position to modulate the step-level advantage.
For a token within step , let denote the Shannon entropy of the policy's next-token distribution at position i:
where is the context preceding token and the sum is over the vocabulary. To make entropies comparable across positions with different context lengths and vocabulary distributions, we normalize within each step:
where and are the mean and standard deviation of entropies within the step. We clamp the normalized values to [−1, 1] to prevent outliers from dominating:
The token-level weight is then defined as a multiplicative modulator:
where β ∈ [0.05, 0.2] is a small hyperparameter. With β = 0.1, the weight range is [0.9, 1.1], providing a ±10% modulation on the step-level advantage.
Why multiplicative, not additive. An additive formulation (token advantage = step advantage + β · entropy) would be unsafe because it creates reward signal independent of step quality: a high-entropy token in a terrible step would receive a positive bonus purely from its entropy, potentially inverting the sign of the advantage and creating reward hacking incentives. The multiplicative formulation is safe by construction: when the step advantage is zero, entropy contributes nothing; when positive, high-entropy tokens receive proportionally more credit; when negative, high-entropy tokens receive proportionally more blame. The sign of the advantage is never flipped by the entropy weight.
3.3. Effective Reward Formulation
Combining the three levels of credit assignment, the token-level advantage for token in step is:
where the step-level advantage is derived from the effective reward:
Here is the binary outcome reward (same for all steps in a trajectory). An alternative formulation for the step-level advantage is:
Here is the base policy's estimated probability of eventually reaching a correct answer from state after taking step . However this formulation require training MC-base PRM to estimate the and Progress Reward model so we don’t consider it.
The difference is: in the first formulation is either 0 or 1 (fixed for all actions in single trajectory), yet in the second formulation we use the Q value which vary across all actions in a single trajectory. For this work we use the first formulation.
is the step-specific progress under the evaluator, and controls the relative contribution of the progress signal. The step advantage is obtained by normalizing effective rewards across rollouts:
where the baseline and standard deviation are computed across all steps and all rollouts for a given question. The resulting policy gradient takes the form:
This gradient has a natural three-level interpretation: R_out provides trajectory-level signal separating correct from incorrect solutions; α · A_μ provides step-level signal differentiating helpful from harmful steps; and (1 + β · H̃ᵢ) provides token-level signal directing credit to decision points.
4. Methodology
4.1 Data Collection for Progress Reward Model
Training the progress reward model requires estimating for a diverse set of reasoning prefixes. We collect this data through Monte Carlo sampling under the evaluator policy, using the following procedure.
Given a set of training questions with verifiable ground-truth answers, a evaluator policy (e.g., best-of-K from the base policy), a rule-based outcome reward function that returns 1 for correct and 0 for incorrect answers, and hyperparameters (number of seed traces per question) and (number of MC rollouts per prefix), we proceed as follows:
Algorithm 1: Data Collection
Several design choices are important:
- Seed traces are sampled from the evaluator policy , not the base policy, since the model must generalize to prefixes encountered by the evaluator.
- The stored label is —a probability, not an advantage. The advantage is derived at RL time by subtracting two predictions.
- We include empty-prefix estimates (question-only) so the model can predict , which is needed to compute the advantage for the first step.
- We allocating more budget to MC rollouts compared to seed budget (, ), since the variance of the estimate decreases as .
4.2. Progress Reward Model Architecture and Training
The model is initialized from a pretrained language model of the same family as the base policy. We replace the language modeling head with a scalar value head consisting of a linear projection from the model's hidden dimension to a single scalar, followed by a sigmoid activation. The model takes a reasoning prefix as input and produces a prediction from the hidden state at the last token position.
Training uses binary cross-entropy (BCE) loss, which is the natural choice for a probability target:
where is the MC-estimated label. We note that BCE generalizes naturally to soft targets in [0, 1] via the cross-entropy between two Bernoulli distributions. At inference time during RL training, the advantage for step h is derived via two forward passes:
This requires two forward passes of the model per step: one for the prefix before the step and one for the prefix after. The advantage is a single subtraction of these two scalar outputs.
4.3 RL Training Procedure
We integrate the progress reward and token-level entropy weighting into a standard policy gradient framework. The complete RL training loop operates as follows:
Algorithm 2: RL Training with Progress and Token-Level Rewards
The entropy values are typically available at no additional computational cost, as the full next-token distribution is already computed during the forward pass used for sampling. The model forward passes constitute the primary additional cost compared to standard ORM-based RL, requiring two passes per step per rollout.
4.4. Comparison with Existing Approaches
Summary of the key differences between our approach and existing reward modeling paradigms:
Method | Signal Level | Grounded in Outcome? | Labels | Scalable? |
ORM | Trajectory only | Yes | Automatic | Yes |
PRM (human) | Step | No | Human | No |
PRM (MC) | Step | Indirect | Automatic | Yes |
Ours | Step + Token | Yes | Automatic | Yes |
Lists the key hyperparameters of our approach and their recommended ranges:
Hyperparameter | Role | Recommended Range |
K (BoK evaluator) | Evaluator strength | 4–8 |
MC rollouts per prefix (data collection) | 16–64 | |
Seed traces per question | 4–8 | |
Progress reward weight in effective reward | 0.5–2.0 | |
Entropy weight on token advantage | 0.05–0.2 |
5. Experimentation and Results
5.1. Experimental Setup
Base Models. We conduct all experiments using Qwen 2.5 3B and Qwen 2.5 7B (Qwen Team, 2024) as the base architecture for both the policy model and the reward model. Using the same base model for both components controls for capacity differences and isolates the effect of the reward signal itself.
Training Data. We construct reward model training data from the MATH training split (Hendrycks et al., 2021) following the procedure described in Section 4.1. We define the evaluator policy μ as a best-of-K policy derived from the base policy (K=4). For each of the 5K training problems, we sample seed trajectories from μ, split each into reasoning steps, and estimate for every resulting prefix via Monte Carlo rollouts under μ, scored by a rule-based outcome verifier. The reward model is trained to regress these soft probability targets using BCE loss, as detailed in Section 4.2. We emphasize that the stored labels are Q-values, not advantages — the advantage is derived at RL time via two forward passes of the trained reward model.
Baselines. We compare against two baselines, both trained on the same 200K trajectory samples and for the same number of training steps to ensure a fair comparison under equivalent data and compute budgets:
- MC-PRM (Math-Shepherd): A process reward model trained following the Math-Shepherd framework (Wang et al., 2024), which assigns a binary correctness label to each reasoning step based on the proportion of Monte Carlo rollouts from that step that reach the correct final answer.
- ORM: An outcome reward model trained to predict the correctness of the final answer only, without any intermediate step-level supervision.
All three reward models—our proposed method, MC-PRM, and ORM—share the same base architecture, training data, and training duration, differing only in the reward signal formulation.
Policy Optimization. We optimize the base policy using Group Relative Policy Optimization (GRPO; Shao et al., 2024) with each of the three reward models.
Evaluation. We evaluate on three benchmarks spanning different difficulty levels: MATH500 (Lightman et al., 2023), a held-out subset of MATH not used in training either the reward model or the policy; GSM8K (Cobbe et al., 2021), a grade-school mathematics benchmark; and GPQA (Rein et al., 2023), a graduate-level science reasoning benchmark. Including benchmarks of varying difficulty and domain allows us to assess both in-distribution generalization and out-of-distribution transfer of the learned reward signal.
Statistical Protocol. All experiments are run across 3 independent seeds. We report the mean and standard deviation of accuracy (%) averaged across the three benchmarks.
5.2 Results
Our progress-aware (Progress-Entropy) reward model consistently outperforms both baselines across both model scales. Relative to MC-PRM, our method achieves a 5 percentage point improvement in average accuracy; relative to ORM, the gain increases to 8 percentage points. These improvements are consistent across both the 3B and 7B settings, suggesting that the benefit of richer reward signals is not contingent on model capacity.
Beyond final performance, we evaluate sample efficiency by measuring accuracy as a function of policy training steps. Our reward model reaches the final performance of MC-PRM in approximately 3x fewer training steps, and the final performance of ORM in approximately 4x fewer steps. This efficiency gain is consistent across both model scales.
Ablation studies to disentangle the contributions of the two reward components and to compare different evaluator policy choices were done but outside the scope of this paper.
6. Discussion
Our framework introduces a hierarchical credit assignment structure that operates at three levels of granularity: trajectory (via the outcome reward), step (via the progress advantage), and token (via entropy weighting). Each level provides diminishing but meaningful improvement in signal specificity. The most substantial gain comes from the trajectory-to-step transition provided by the progress reward, which transforms the sparse binary outcome into a dense per-step signal grounded in actual success probability changes. The step-to-token transition provided by entropy weighting offers a more modest refinement, directing credit within each step toward the tokens that represent genuine decisions.
A key distinction from prior work on process reward models is that our progress signal measures a fundamentally different quantity than step correctness. Traditional PRMs ask whether each step is valid; our approach asks whether each step made the problem more solvable. This distinction has practical consequences: a correct but redundant step receives near-zero progress signal under our framework, while a traditional PRM would assign it a high score. Similarly, an unconventional step that happens to be the key insight receives high progress signal if it increases the evaluator’s success probability, regardless of whether a correctness judge would flag it.
The requirement for a separate evaluator policy introduces an additional component into the training pipeline. However, the evaluator need not be a specially trained model—the best-of-K strategy over the base policy itself provides an effective evaluator at the cost of K completions per seed trace during data collection. Moreover, the model is trained once and frozen during RL, so the additional cost is amortized over many RL iterations.
References
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. (2021). Training verifiers to solve math word problems. arXiv:2110.14168.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K. (2023). Let's verify step by step. arXiv:2305.20050.
Luo, H., Sun, Q., Xu, C., Zhao, P., Lou, J., Tao, C., Geng, X., Lin, Q., Chen, S., and Zhang, D. (2024). Improve mathematical reasoning in language models by automated process supervision. arXiv:2406.06592.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022). Training language models to follow instructions with human feedback. NeurIPS.
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Zhang, M., Li, Y., Wu, Y., and Guo, D. (2024). DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv:2402.03300.
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Ciosek, K., Mocanu, D., Oliveira, R., and others (2022). Solving math word problems with process- and outcome-based feedback. arXiv:2211.14275.
Wang, P., Li, L., Shao, Z., Xu, R.X., Dai, D., Li, Y., Chen, D., Wu, Y., and Sui, Z. (2024). Math-Shepherd: Verify and reinforce LLMs step-by-step without human annotations. arXiv:2312.08935.
Xiong, W., Zhang, H., Ye, C., Chen, L., Jiang, N., and Zhang, T. (2024). An implementation of generative PRM. RLHF-Reward-Modeling, GitHub.
Zhang, Z., Zheng, C., Wu, Y., Zhang, B., Lin, R., Yu, B., Liu, D., Zhou, J., and Lin, J. (2025). The lessons of developing process reward models in mathematical reasoning. arXiv:2501.07301.