Research blog · RL for LLM reasoning

NIPS26 LPO: Your GRPO update is secretly chasing a target.
LPO just aims at it.

Group-based RLVR methods (GRPO, Dr.GRPO, MaxRL, and friends) look like tweaks to advantage normalization. Underneath, they all share one geometric move: build a target distribution over the sampled responses, then take a first-order step toward it. Listwise Policy Optimization (LPO) makes that move explicit and exact.

Yun Qu, Qi Wang, Yixiu Mao, Heming Zou, Yuhang Jiang, Yingyue Li, Wutong Xu, Lizhou Cai, Weijie Liu, Clive Bai, Kai Yang, Yangkun Chen, Saiyong Yang, Xiangyang Ji
Department of Automation, Tsinghua University · LLM Department, Tencent · arXiv, May 2026
TL;DR

1 · A group of answers is a little universe

Here's the standard RLVR loop: give the model a math problem, sample $K=8$ attempts, let a verifier mark each one right or wrong, and nudge the model toward the good ones. GRPO does the nudging with group-relative advantages: subtract the group's mean reward, divide by its standard deviation.

Now change the viewpoint. Instead of looking at each response in isolation, ask: among these eight, how much does the current policy prefer each one, relative to the policy that sampled them? That gives a distribution over the group:

$$P_{\theta,k} \;=\; \mathrm{softmax}(s_\theta)_k,\qquad s_{\theta,k} = \log\frac{\pi_\theta(y_k\mid x)}{\pi_b(y_k\mid x)}.$$

Before any update, $\pi_\theta=\pi_b$ and this is just uniform, $1/K$ each. As training moves the policy, $P_\theta$ slides around a $(K-1)$-dimensional simplex, the response simplex. That tiny, finite space is where the whole story plays out.

2 · The target nobody wrote down

Here's the first result. Take any zero-mean advantage vector $A$ and define $w^* = \mathrm{softmax}(A)$. Then, at the on-policy point,

$$\underbrace{\tfrac{1}{K}\textstyle\sum_k A_k\,\nabla_\theta \log\pi_\theta(y_k\mid x)}_{\text{the policy gradient you already use}} \;=\; -\nabla_\theta\, D_{\mathrm{KL}}\!\left(P_\theta \,\|\, w^*\right).$$
A policy-gradient step is a gradient step on reverse KL toward a hidden target.
…but only exactly at the on-policy point. The error grows with off-policy drift.

Since advantages have the shape $(R_k-\mu)/\tau$ and softmax ignores the shift $\mu$, the hidden target is always $\mathrm{softmax}(R/\tau)$. Methods that look different on paper turn out to aim at the same family of targets, and differ only in how sharp that target is:

MethodAdvantageImplicit targetTemperature τ
GRPO / DAPO$(R_k-\mu_G)/\sigma_G$$\mathrm{softmax}(R/\sigma_G)$group std $\sigma_G$
Dr.GRPO / RLOO$R_k-\mu_G$$\mathrm{softmax}(R)$≈ 1
MaxRL$(R_k-\mu_G)/\mu_G$$\mathrm{softmax}(R/\mu_G)$success rate $\mu_G$
▶ Play: meet the hidden target

Click responses to mark them correct ✓ or wrong ✗, then switch the normalizer. The bars show how much probability each method's hidden target puts on each response.

uniform 1/K

Play with it and you'll see something neat. MaxRL's target gets extremely sharp on hard prompts: with one lucky success out of eight, τ = 1/8 and that response is weighted about 3000× more than each failure. On easy prompts it relaxes to about 3×. GRPO is symmetric instead: its target is softest for a 50/50 group (σ is largest there) and sharpens toward both extremes. The normalizer isn't just a variance trick; it's a choice of where you're aiming.

3 · From "aim roughly" to "aim exactly": LPO

Classic RL-as-inference methods (REPS, MPO, AWR) also build a reward-weighted target and move the policy toward it, but with continuous actions they must approximate everything. LLM RLVR has a lucky structural gift: the sampled responses form a finite simplex. Both the target and the projection can be computed in closed form. So LPO does exactly that, in two decoupled steps:

STEP 1 · TARGET

What to aim for

Maximize expected reward on the simplex, inside a KL trust region around the current policy:

$$\max_{w\in\Delta^{K-1}} \textstyle\sum_k w_k R_k - \tau\, D_{\mathrm{KL}}(w\|P_t)$$

Unique closed-form answer, a listwise Gibbs target:

$$w^*_k = \mathrm{softmax}\!\left(\tfrac{R_k}{\tau} + s_{t,k}\right)$$
STEP 2 · PROJECTION

How to get there

Minimize a divergence between target and listwise policy. Every choice gives $\nabla = \sum_k c_k \nabla_\theta\log\pi_\theta(y_k|x)$:

LPOfwd · $D_{\mathrm{KL}}(w^*\|P_\theta)$
$c_k = P_{\theta,k} - w^*_k$

LPOrev · $D_{\mathrm{KL}}(P_\theta\|w^*)$
$c_k = P_{\theta,k}(d_k-\bar d),\ \ d_k = s_{\theta,k}-\phi_k$

Two nice things fall out immediately. On-policy, Step 1 reproduces exactly the hidden targets in the table above, except now $\tau$ has a real meaning (a trust-region strength) instead of being a side effect of normalization. And because target and projection are separate, the divergence becomes a design knob that policy gradient never exposed.

▶ Play: race to the target on a 3-response simplex

Three sampled responses. Each corner is "all probability on that response"; the centre is the untouched policy. We run several inner updates on the same samples (like PPO-style epochs), which is exactly where off-policy drift kicks in. Toy model: each response's log-probability is a free parameter; PG uses unclipped importance ratios.

Policy gradient LPOfwd LPOrev ★ target $w^*$
Rewards (click to cycle 0 → 0.5 → 1)
iteration 0
Hit Run iteration. Notice that the very first grey step lands exactly on the first yellow one: on-policy, PG is reverse-KL projection. After that PG's step size never shrinks (the importance ratio even amplifies it), so it blows straight past the star into the corner. PG's E[R] of 1.00 looks great, but it got there by putting everything on one response from a group of three, ignoring the trust region. In a real LLM that's the overconfident, entropy-collapsing update clipping has to fight. The LPO paths slow down and stop on the target. Run a few more iterations to see LPO climb toward the best response one controlled step at a time.

Why the exact projection behaves so well

⚖️

Zero-sum

Coefficients sum to 0: pushing one response up automatically pushes others down. A built-in control variate, for any divergence you pick.

📏

Bounded

For LPOfwd, $|c_k|\le 1$ and $\sum_k|c_k|\le 2$ regardless of reward scale. No exploding updates from lopsided groups.

🎯

Self-correcting

Coefficients vanish as $P_\theta\to w^*$. The update knows when it has arrived instead of relying on a clip to stop it.

On top of that, iterating target-then-project gives a monotonic improvement guarantee: the listwise reward goes up each iteration by at least $\tau$ times the Jeffreys divergence between old policy and target, minus a term for imperfect projection. And the two KL flavours bring their own personalities. Forward KL is mode-covering: any response the target cares about gets a guaranteed probability floor, a log-barrier against collapse. Reverse KL secretly contains an entropy bonus: it decomposes into expected target logit plus $H(P_\theta)$.

Is "listwise" really the key? The authors ablate it: keep the same target, but fit it pointwise with a weighted log-likelihood (the MPO/AWR style). Performance drops severely and optimization turns unstable. In the pointwise loss every response gets pushed up, just by different amounts, and the push never switches off: there's no competition between responses and no built-in control variate. Exact target fitting only pays off when paired with listwise normalization.

4 · Does it work?

The comparison is deliberately strict. For each baseline (GRPO, Dr.GRPO, MaxRL), LPO uses the exact same temperature as that baseline's implicit target. Only the projection changes, so any difference is down to exact listwise projection, not tuning. Tasks: Countdown (logic), MATH (math), PRIME (code) and Geometry3k (vision-language), with Qwen3 1.7B/4B/8B/14B, Qwen2.5-VL-3B, and cross-family checks on DeepSeek-R1-Distill, Llama-3.1 and Mistral.

13/15
settings where each LPO variant beats its matched PG baseline on Pass@1 training curves
15/15
settings where LPOfwd beats PG on Pass@k (LPOrev: 11/15)
~70
steps for LPOfwd to reach GRPO's 200-step peak (Qwen3-14B, Polaris)
0
extra compute: same rollouts, same pipeline, only the loss coefficients change
Final math benchmark scores (average over 6 benchmarks)

Pass@1 averages all six benchmarks (AIME24/25, AMC23, MATH500, Minerva, OlympiadBench); Pass@k averages the four with k>1. Each group compares a PG baseline with its temperature-matched LPO variants. Axis is truncated to make differences visible. Not every cell is a win: on Qwen3-8B, Dr.GRPO's own Pass@k stays on top.

5 · The bigger picture

Seen through this lens, a lot of RLVR design debates reorganize themselves. Arguments about how to normalize advantages become arguments about which target to aim at. Clipping and trust regions become questions about how to project. Keeping the two separate opens a design space that was hidden before:

Group-based RL was always target-then-project.
LPO just stops approximating.

Cite

@article{qu2026lpo,
  title   = {Listwise Policy Optimization: Group-based RLVR as
             Target-Projection on the LLM Response Simplex},
  author  = {Qu, Yun and Wang, Qi and Mao, Yixiu and Zou, Heming and
             Jiang, Yuhang and Li, Yingyue and Xu, Wutong and Cai, Lizhou and
             Liu, Weijie and Bai, Clive and Yang, Kai and Chen, Yangkun and
             Yang, Saiyong and Ji, Xiangyang},
  journal = {arXiv preprint arXiv:2605.06139},
  year    = {2026}
}