Meet POPO, a simple drop-in fix that fills every RLVR training batch with useful data, without generating a single extra sample.
RL with verifiable rewards (RLVR) works like a very patient tutor. Give the model a problem, let it try k times, check each answer, and nudge it toward the attempts that beat the group average. GRPO turns this into an advantage: each response's reward minus the group mean, divided by the group's standard deviation.
Now picture a problem the model already solves every time, or one it never solves. Every attempt gets the same score. There is no "better than average" in that group, so every advantage is zero. The model spent GPU-minutes writing eight long chains of thought and learned exactly nothing.
Try it yourself below. Click the attempts to flip them between right and wrong.
This isn't an edge case. In the paper's experiments, vanilla GRPO typically trains on batches where fewer than half the groups are effective. That holds across math (DeepScaleR), arithmetic planning (Countdown) and visual geometry. As the model improves, more prompts turn into easy "all correct" groups, so the problem gets worse over time.
Roll out extra prompts, drop the zero-variance ones, and repeat until the batch is full.
Catch: the extra rollouts can cost more than the RL update itself.
Guess which prompts are "medium difficulty" before rolling out, e.g. MoPPS or GRESO.
Catch: guessing is hard when the model keeps changing and each prompt has little history.
Stick an old correct answer into an all-wrong group, e.g. ARPO.
Catch: all-correct groups stay dead, and mixing policies inside a group muddies the advantage.
Think of a football coach with a strict rule: every slot on the pitch must be filled by someone who can actually play. When a starter turns up injured, the coach doesn't recruit a stranger from the street, which is what DAPO's extra rollouts amount to. They also don't patch one half-fit player into the lineup, which is what trajectory replay does. They bring in a substitute who played well in the last match.
That's prioritized group replay. It follows two rules:
Why is recency a good stand-in for "close to the current policy"? PPO-style clipping keeps each update small. So the gap between two policies, measured by total variation distance, grows at most linearly with the number of steps between them: it is at most εn/2 after n steps. Freshness is therefore a free proxy for similarity, with no KL computations needed.
Two details make this cheap and clean. First, the buffer never holds more than one batch of groups, so memory overhead is negligible. Second, because an entire group is replayed together, every answer in it came from the same policy. The group's advantage is still well defined, and the behavior policy is known exactly. That second point matters a lot for the next step.
A replayed group was generated by an older policy, call it πβ. If you just pretend it's fresh, your gradient is biased. Prior work handled this in two ways, and POPO takes a third.
✗ Biased whenever πβ ≠ πold
In the ablation it collapses.
✓ Unbiased
✗ Leashes the model to a worse policy, and a different leash for every replayed sample. Clipping triggers too often.
✓ Unbiased: the first factor fixes the distribution
✓ Same leash as on-policy data: the clip always anchors to πold
The trick is that the full importance ratio πθ/πβ factors into two pieces, each with its own job:
The orange factor corrects where the data came from. The blue ratio sets how far this update may move, and it uses exactly the same trust region as fresh data.
The paper proves that wherever neither ratio is clipped, this objective gives the same value and the same gradient as the fully corrected πβ version. The only thing that changes is which policy the leash is tied to. It's like a GPS that accounts for where your map was drawn, but still measures your next turn from where you're standing now.
Correct for the past, but stay anchored to the present.
In code, POPO adds a filter, a tiny queue and one multiplication:
buffer = FIFO(capacity=B) # stores groups + their rollout-time log-probs for step in range(T): groups = rollout(policy, sample_prompts(B), k) # same cost as GRPO on_eff = [g for g in groups if std(g.rewards) > 0] # drop dead groups off_eff = buffer.newest(B - len(on_eff)) # refill with most recent loss = grpo_loss(on_eff) w = exp(logp_old(off_eff) - off_eff.logp_beta).clamp(max=2.0).detach() loss += (w * grpo_token_loss(off_eff)).aggregate() # the decoupled correction optimize(policy, loss) buffer.push(on_eff) # today's effective groups
In practice, the paper also caps the importance weight w at 2.0 to suppress rare outliers. It reports no variance-related instability.
POPO was compared against GRPO, DAPO (the costly gold standard), MoPPS (predictive sampling) and ARPO (trajectory replay). The comparisons cover four settings: Qwen2.5-3B on Countdown, Qwen2.5-VL-3B on Geometry3k, and DeepSeek-R1-Distill-Qwen 1.5B and 7B on DeepScaleR math.
The pattern holds everywhere. POPO lands on top of or right next to DAPO and clearly ahead of GRPO, ARPO and MoPPS. Its cost stays close to plain GRPO. The small runtime bump over GRPO is actually a good sign: POPO's models write longer reasoning chains, in a way that tracks DAPO's, and longer reasoning is a known correlate of stronger reasoning.
A few highlights worth pausing on:
Every piece of POPO earns its place. Here is what happens on Countdown when you remove or swap each one:
About the same accuracy, but slower: 4.8h versus 3.2h. Recency is a free and effective stand-in for distance.
Only slightly better than GRPO. Smaller batches mean noisier gradients, so refilling the batch is what matters.
Substantial degradation. The quality filter matters.
Collapses. A 1× random buffer is only slightly worse, so freshness matters, but POPO isn't brittle to small changes in recency.
Learning slows down. This is the "leash tied to the past" problem in action.
Performance collapses because of the uncorrected bias.
POPO is a plug-in rather than a whole new algorithm, and the paper tests it that way:
A buffer that only remembers the latest batch is deliberately forgetful. It trades reuse of older, possibly valuable groups for a tight off-policy gap, and it doesn't try to pick the most informative group. The authors suggest a hybrid as future work: a short-term buffer plus a small long-term "reservoir" of standout groups. They also suggest generating fewer fresh rollouts per step and leaning harder on replay. Finally, the zero-variance test assumes binary-ish rewards. With continuous rewards, it would become a low-variance threshold instead.
RLVR's most expensive ingredient is generation, and a surprisingly large share of it gets thrown away as zero-gradient groups. POPO's answer is refreshingly small. It keeps the last batch's useful groups, swaps them in for today's dead ones, and applies one importance weight so the math stays honest. The payoff is DAPO-level results on a GRPO-sized rollout budget.
Every slot in the batch should teach the model something, and it shouldn't cost you another rollout to make that happen.