- Sample $K$ responses to a prompt and they form a tiny probability simplex. On it, every group-based policy gradient implicitly targets a reward-weighted softmax, $w^* = \mathrm{softmax}(R/\tau)$, where the "temperature" $\tau$ is whatever the advantage normalizer divides by.
- The usual policy-gradient step is only a first-order approximation of a reverse-KL projection toward that target, and it drifts once updates go off-policy.
- LPO splits each update into two clean steps: compute the target in closed form, then project onto it exactly with a divergence of your choice. Gradients become bounded, zero-sum and self-correcting, at no extra compute.
- Across logic, math, code and multimodal geometry (1.5B–14B models, 4 LLM families), LPO beats its matched PG baseline in the large majority of settings, while keeping entropy up and gradient norms down.
1 · A group of answers is a little universe
Here's the standard RLVR loop: give the model a math problem, sample $K=8$ attempts, let a verifier mark each one right or wrong, and nudge the model toward the good ones. GRPO does the nudging with group-relative advantages: subtract the group's mean reward, divide by its standard deviation.
Now change the viewpoint. Instead of looking at each response in isolation, ask: among these eight, how much does the current policy prefer each one, relative to the policy that sampled them? That gives a distribution over the group:
$$P_{\theta,k} \;=\; \mathrm{softmax}(s_\theta)_k,\qquad s_{\theta,k} = \log\frac{\pi_\theta(y_k\mid x)}{\pi_b(y_k\mid x)}.$$Before any update, $\pi_\theta=\pi_b$ and this is just uniform, $1/K$ each. As training moves the policy, $P_\theta$ slides around a $(K-1)$-dimensional simplex, the response simplex. That tiny, finite space is where the whole story plays out.
2 · The target nobody wrote down
Here's the first result. Take any zero-mean advantage vector $A$ and define $w^* = \mathrm{softmax}(A)$. Then, at the on-policy point,
$$\underbrace{\tfrac{1}{K}\textstyle\sum_k A_k\,\nabla_\theta \log\pi_\theta(y_k\mid x)}_{\text{the policy gradient you already use}} \;=\; -\nabla_\theta\, D_{\mathrm{KL}}\!\left(P_\theta \,\|\, w^*\right).$$…but only exactly at the on-policy point. The error grows with off-policy drift.
Since advantages have the shape $(R_k-\mu)/\tau$ and softmax ignores the shift $\mu$, the hidden target is always $\mathrm{softmax}(R/\tau)$. Methods that look different on paper turn out to aim at the same family of targets, and differ only in how sharp that target is:
| Method | Advantage | Implicit target | Temperature τ |
|---|---|---|---|
| GRPO / DAPO | $(R_k-\mu_G)/\sigma_G$ | $\mathrm{softmax}(R/\sigma_G)$ | group std $\sigma_G$ |
| Dr.GRPO / RLOO | $R_k-\mu_G$ | $\mathrm{softmax}(R)$ | ≈ 1 |
| MaxRL | $(R_k-\mu_G)/\mu_G$ | $\mathrm{softmax}(R/\mu_G)$ | success rate $\mu_G$ |
Click responses to mark them correct ✓ or wrong ✗, then switch the normalizer. The bars show how much probability each method's hidden target puts on each response.
Play with it and you'll see something neat. MaxRL's target gets extremely sharp on hard prompts: with one lucky success out of eight, τ = 1/8 and that response is weighted about 3000× more than each failure. On easy prompts it relaxes to about 3×. GRPO is symmetric instead: its target is softest for a 50/50 group (σ is largest there) and sharpens toward both extremes. The normalizer isn't just a variance trick; it's a choice of where you're aiming.
3 · From "aim roughly" to "aim exactly": LPO
Classic RL-as-inference methods (REPS, MPO, AWR) also build a reward-weighted target and move the policy toward it, but with continuous actions they must approximate everything. LLM RLVR has a lucky structural gift: the sampled responses form a finite simplex. Both the target and the projection can be computed in closed form. So LPO does exactly that, in two decoupled steps:
What to aim for
Maximize expected reward on the simplex, inside a KL trust region around the current policy:
$$\max_{w\in\Delta^{K-1}} \textstyle\sum_k w_k R_k - \tau\, D_{\mathrm{KL}}(w\|P_t)$$Unique closed-form answer, a listwise Gibbs target:
$$w^*_k = \mathrm{softmax}\!\left(\tfrac{R_k}{\tau} + s_{t,k}\right)$$How to get there
Minimize a divergence between target and listwise policy. Every choice gives $\nabla = \sum_k c_k \nabla_\theta\log\pi_\theta(y_k|x)$:
LPOfwd · $D_{\mathrm{KL}}(w^*\|P_\theta)$
$c_k = P_{\theta,k} - w^*_k$
LPOrev · $D_{\mathrm{KL}}(P_\theta\|w^*)$
$c_k = P_{\theta,k}(d_k-\bar d),\ \ d_k = s_{\theta,k}-\phi_k$
Two nice things fall out immediately. On-policy, Step 1 reproduces exactly the hidden targets in the table above, except now $\tau$ has a real meaning (a trust-region strength) instead of being a side effect of normalization. And because target and projection are separate, the divergence becomes a design knob that policy gradient never exposed.
Three sampled responses. Each corner is "all probability on that response"; the centre is the untouched policy. We run several inner updates on the same samples (like PPO-style epochs), which is exactly where off-policy drift kicks in. Toy model: each response's log-probability is a free parameter; PG uses unclipped importance ratios.
Why the exact projection behaves so well
Zero-sum
Coefficients sum to 0: pushing one response up automatically pushes others down. A built-in control variate, for any divergence you pick.
Bounded
For LPOfwd, $|c_k|\le 1$ and $\sum_k|c_k|\le 2$ regardless of reward scale. No exploding updates from lopsided groups.
Self-correcting
Coefficients vanish as $P_\theta\to w^*$. The update knows when it has arrived instead of relying on a clip to stop it.
On top of that, iterating target-then-project gives a monotonic improvement guarantee: the listwise reward goes up each iteration by at least $\tau$ times the Jeffreys divergence between old policy and target, minus a term for imperfect projection. And the two KL flavours bring their own personalities. Forward KL is mode-covering: any response the target cares about gets a guaranteed probability floor, a log-barrier against collapse. Reverse KL secretly contains an entropy bonus: it decomposes into expected target logit plus $H(P_\theta)$.
4 · Does it work?
The comparison is deliberately strict. For each baseline (GRPO, Dr.GRPO, MaxRL), LPO uses the exact same temperature as that baseline's implicit target. Only the projection changes, so any difference is down to exact listwise projection, not tuning. Tasks: Countdown (logic), MATH (math), PRIME (code) and Geometry3k (vision-language), with Qwen3 1.7B/4B/8B/14B, Qwen2.5-VL-3B, and cross-family checks on DeepSeek-R1-Distill, Llama-3.1 and Mistral.
Pass@1 averages all six benchmarks (AIME24/25, AMC23, MATH500, Minerva, OlympiadBench); Pass@k averages the four with k>1. Each group compares a PG baseline with its temperature-matched LPO variants. Axis is truncated to make differences visible. Not every cell is a win: on Qwen3-8B, Dr.GRPO's own Pass@k stays on top.
- Diversity is preserved. Both LPO variants keep response entropy higher than PG throughout training, the direct antidote to the entropy collapse that plagues RLVR, and the reason Pass@k gains are so robust.
- Training is calmer. Gradient norms are lower and more stable than PG's, exactly what bounded, self-correcting coefficients predict.
- Forward vs. reverse KL has a personality. Head to head, LPOfwd beats LPOrev on Pass@k in 13/15 settings. That's its mode-covering nature keeping many valid solution paths alive. LPO responses also run longer than PG's, and LPOfwd's are the longest.
- Small groups benefit most. Across $K\in\{2,4,8,16,32\}$ the gains are largest at small $K$, where noisy advantages hurt PG the most.
- Theory checks out. In a strictly on-policy run (one update per batch), LPOrev and GRPO curves are practically identical, just as the equivalence predicts. The gap opens once updates go off-policy; even in that on-policy setup, LPOfwd is more sample-efficient early and ends higher on Pass@k.
5 · The bigger picture
Seen through this lens, a lot of RLVR design debates reorganize themselves. Arguments about how to normalize advantages become arguments about which target to aim at. Clipping and trust regions become questions about how to project. Keeping the two separate opens a design space that was hidden before:
- Any divergence. Zero-sum gradients hold for any differentiable divergence on the simplex, so Jensen–Shannon and other f-divergences, or schedules (forward KL early to explore, reverse KL late to exploit, or annealing τ) are all fair game.
- DPO as the K=2 special case. With two responses, LPOfwd becomes a binary cross-entropy with soft, τ-controlled labels: an online, trust-region cousin of DPO. As $K\to\infty$ you recover classical KL-regularized RL.
- Plug-and-play. Engineering tricks like DAPO's dynamic sampling and asymmetric clipping are orthogonal to LPO and can be layered on top. The paper deliberately uses a minimal shared pipeline so the gains are clearly due to the projection.
- What's next. Step-level listwise projection for multi-turn agents and process rewards, and off-policy replay where listwise normalization acts as self-normalized importance sampling.
LPO just stops approximating.
Cite
@article{qu2026lpo,
title = {Listwise Policy Optimization: Group-based RLVR as
Target-Projection on the LLM Response Simplex},
author = {Qu, Yun and Wang, Qi and Mao, Yixiu and Zou, Heming and
Jiang, Yuhang and Li, Yingyue and Xu, Wutong and Cai, Lizhou and
Liu, Weijie and Bai, Clive and Yang, Kai and Chen, Yangkun and
Yang, Saiyong and Ji, Xiangyang},
journal = {arXiv preprint arXiv:2605.06139},
year = {2026}
}