The short version
- Outcome-only RL learns by comparison. A group of rollouts that all succeed, or all fail, produces no gradient at all.
- In multi-turn agents, every thought-action-observation turn is a natural save point. Replaying from the right save point turns one terminal reward into a controlled side-by-side comparison.
- TRACE uses a small learned predictor to spend a fixed budget on the prompts and prefixes most likely to split into success and failure. Same budget, more signal: up to +2.8 points on Qwen3-14B multi-hop QA over GRPO, and more than double the share of useful training groups on math.
Eight rollouts, zero signal
Most scalable RL pipelines for language models reward only the final outcome: the answer is right or it isn't. GRPO-style training squeezes a learning signal out of that by comparison. Sample a group of rollouts for the same prompt, score each one 1 or 0, and push the policy toward the rollouts that beat the group average.
That works well until the group agrees. If all eight rollouts succeed, every advantage is exactly zero. The same happens when all eight fail. You paid for eight long, tool-calling trajectories and got a gradient of nothing. Try it below: push the success rate toward either end and watch the signal disappear.
When does a group actually teach anything?
Each circle is one rollout. The number under it is its group-relative advantage, (reward − mean) ÷ std.
There is a second, quieter problem. In a multi-turn agent, one terminal reward is stamped onto every turn of the trajectory. A search agent that issued three sharp queries and then fumbled the final answer receives the same 0 as one that went wrong at the very first step. The reward says something failed, not where.
Prior work on prompt selection and rollout allocation attacks the first problem at the prompt level: skip prompts that are too easy or too hard, and give more rollouts to the promising ones. But once a prompt is chosen, each rollout is still sampled as one indivisible trajectory. TRACE goes one level deeper.
Your rollout is secretly a tree
A ReAct-style agent works in turns: think, act (search, run code, call an API), read the observation. Each finished turn is a clean, meaningful save point. Load that save point, keep the history fixed, and sample a fresh future. If the replay ends differently from the original, you have a controlled experiment: same past, different continuation, opposite outcome. That is local credit assignment without a hand-built process reward.
The catch is choosing which save points to replay. Some are already doomed: the search returned junk and the agent is going nowhere. Some are already won: the answer is in hand. Replaying those just buys more copies of a known ending. The valuable ones are knife-edge states where the next few decisions still determine the result.
Seen this way, the whole rollout budget is a tree-building budget. Prompts are depth-zero anchors, visited prefixes are deeper anchors, and every sampling decision is a choice of which anchor gets more descendants.
One rule, at two scales
TRACE's allocation principle fits in one sentence: give budget to the anchors whose descendants are most likely to contain both a success and a failure. It applies twice.
At the prompt
If a predictor estimates that prompt x succeeds with probability v, then m fresh rollouts form a mixed group with probability
TRACE picks a count m ∈ {0, 2, 3, …} for every candidate prompt so the total equals the root budget. Zero means skip the prompt; two or more means keep it and sets how many rollouts it gets. One knob replaces both prompt filtering and rollout-count allocation.
At a prefix
For a prefix, we already know how the original rollout ended: reward r. Let q be the predicted chance that a fresh continuation repeats that ending. Then k replays produce at least one flip with probability
Each active prompt gets a local continuation budget proportional to its rollouts, spread across its visited prefixes. Both problems are solved exactly with a small dynamic program whose cost is negligible next to generation.
Play the allocator yourself. Six prompts, one budget. Uniform sampling spreads rollouts evenly; TRACE solves the root-allocation problem above.
You have a rollout budget. Where does it go?
Illustrative prompts with made-up predicted success rates. Each square is one rollout.
| Prompt | Predicted success | Uniform | TRACE |
|---|
Notice what the optimum does. At a tight budget, everything goes to the prompts in the middle and the lopsided ones are skipped. As the budget grows, the coin-flip prompts saturate (five rollouts already give them a better than 90% chance of a mixed group), so extra rollouts flow to the 90% and 15% prompts, where each one still buys real contrast. The 99% prompt is never worth it, and the 2% prompt only earns rollouts near the top of the slider, once a lucky success becomes plausible.
Why the middle is where the action is
Picture a win-probability meter in a sports broadcast. Early in the game it hovers near 50% and swings with every play. Late in a blowout it sits at 98% and barely twitches. A policy's chance of solving a task behaves the same way as a rollout unfolds turn by turn, and the paper makes this precise with three results.
More history, better forecasts. Predicting how a group of continuations will score can only get easier as you observe more turns: the best achievable squared error never increases with depth. Scoring prefixes is therefore at least as informed as scoring prompts. On Qwen3-8B HotpotQA, the predictor's group-success error falls to roughly a third of its prompt-level value by the fourth turn.
Uncertainty is remaining contrast. Treat the success probability V as that live meter. From any prefix, the expected total squared movement of the meter until the rollout ends equals exactly
So V(1 − V) is more than a static uncertainty score: it measures how much the outcome can still move below that prefix. The empirical picture agrees. Many prompts and prefixes sit near 0% or 100%, and ranking anchors by this contrast lets a small fraction of the budget capture most of the available pairwise contrast.
Contrast is what turns on the gradient. With binary rewards, both pairwise and group-relative updates vanish below an anchor unless its descendants include a success and a failure. The expected squared gradient therefore factors into an activation probability times a gradient scale. Under the paper's normalization assumption, maximizing activation probability, which is exactly what TRACE's two objectives do, yields at least as much expected gradient energy as uniform allocation at every stage.
A 0.6B model tells a 14B model where to look
Every formula above needs a success estimate for a prompt or a partial trajectory before sampling. TRACE uses one shared predictor for both: a Qwen3-0.6B critic that reads the serialized prompt plus interaction history and outputs a score between 0 and 1.
It trains online from the trees TRACE collects. Every node gets a target equal to the success rate of the leaves beneath it, computed bottom-up, and the predictor regresses onto those targets after each step. Training is dominated by prompt-level examples, with prefixes making up only 6% of each predictor batch. Even so, its rank correlation with real outcomes stays positive at the prefix level, which suggests it learns a history-conditioned sense of difficulty rather than memorizing prompt identities.
And it is cheap. On HotpotQA, predictor scoring plus updates take about 3% of wall-clock time:
How one training step runs
- Score the candidate pool. The predictor estimates a success rate for every candidate prompt.
- Allocate root rollouts. Solve the budgeted root problem: most candidates get zero, active prompts get two or more bare rollouts.
- Expand locally, right away. As soon as one prompt's rollouts return, score its visited prefixes and spend its continuation budget there. No waiting for the slowest prompt in the batch.
- Refresh the predictor. Compute bottom-up success targets on the finished trees and take a regression step.
- Update the policy. Hand the trees to any tree-aware optimizer. The paper uses TreeRPO for math and QA and Tree-GRPO for function calling.
The per-prompt design in step 3 is deliberately systems-aware. A fully global prefix allocator would have to wait for every prompt in the batch to finish; TRACE only needs one prompt's rollouts, which already live on the same worker.
Results at equal budget
The evaluation covers three agentic settings: mathematical reasoning with a Python interpreter (trained on DeepScaleR), multi-hop QA with a local Wikipedia retriever (trained on HotpotQA), and multi-turn function calling (BFCL v4). Baselines are GRPO, PCL (predictive prompt selection), and TreePO, which uses the same tree rollouts and tree-aware updates as TRACE but branches at random. Every method uses the same rollout-budget accounting, so differences come from where samples go, not how many there are.
Accuracy gain over GRPO
Bars show points above GRPO at the end of training; labels on the left give the absolute scores.
TRACE posts the best average in every setting and at both Qwen scales, and the gap is widest where trajectories are longest and most branching-friendly: multi-hop QA and function calling. The comparison with TreePO is the telling one. Both methods build trees and use the same optimizer; the only difference is whether branches go to random prefixes or predicted ones. The same pattern holds on Llama-3.2-3B-Instruct, where TRACE beats GRPO by 3.8 points and random branching by 3.1.
Many more groups that actually teach
The clearest effect is on the effective ratio: the share of prompts in a batch whose rollout trees end in both successes and failures. On DeepScaleR math:
Both stages pull their weight
Swapping each learned allocator for a uniform one on Qwen3-8B HotpotQA shows the gains stack: root allocation picks prompts likely to split, prefix allocation spends replays where they can still reveal contrast.
| Prompt stage | Prefix stage | Avg. accuracy | Effective ratio |
|---|---|---|---|
| Uniform | Uniform | 49.5 | 42.8% |
| Adaptive | Uniform | 49.8 | 49.1% |
| Uniform | Adaptive | 50.0 | 47.3% |
| Adaptive | Adaptive | 50.6 | 52.3% |
The shape of the budget matters, not just its size
A mid-trajectory replay costs about half a full rollout, so a configuration with M root rollouts and N continuations per root costs roughly M(1 + N/2) trajectories. TRACE beats random branching at every budget, lifting the effective ratio by about ten points each time.
| Budget | Roots M, replays N | TreePO acc. / eff. | TRACE acc. / eff. |
|---|---|---|---|
| 1024 | 512, 2 | 48.8 / 32.2% | 49.7 / 42.4% |
| 2048 | 512, 6 | 49.4 / 37.7% | 50.3 / 47.8% |
| 2048 | 1024, 2 | 49.5 / 42.8% | 50.6 / 52.3% |
At the same 2048 budget, broader root coverage beats deeper replay from fewer roots. The bottleneck isn't the raw number of rollouts; it is whether they reach states where rewards can disagree.
Where the budget ends up
Watching TRACE allocate on BFCL function calling, where episodes can run to dozens of turns, makes the behavior concrete. At the prompt level it is selective: about 71% of candidate prompts are skipped with Qwen3-8B and about 65% with Qwen3-14B, and active prompts typically receive five to seven rollouts. Importantly, the number of distinct prompts trained on never exceeds the fixed-size baselines, so the gains don't come from seeing more data.
At the prefix level, replays are spread along the whole trajectory. The earliest turns get the largest single share, but meaningful budget reaches prefixes as deep as turn 15. Stage 2 isn't quietly acting as extra root sampling; it is finding decision points throughout long interactions.
Limits and next steps
TRACE is built for outcome-verifiable tasks; settings without a clear terminal check would need the mixed-outcome objective rethought. The predictor here is intentionally plain, and a sharper one should place budget better, especially since it learns against targets that shift as the policy improves, a natural fit for continual-learning techniques. Experiments cover math, multi-hop QA and function calling with Qwen3-8B, Qwen3-14B and Llama-3.2-3B, leaving longer-horizon and non-stationary agent environments for future work.
The broader takeaway is a reframing. For agentic RL, the sampling question isn't only how many independent rollouts to draw. It is where in the rollout tree to branch.
Cite this work
@article{zou2026trace,
title = {TRACE: A Unified Rollout Budget Allocation Framework for
Efficient Agentic Reinforcement Learning},
author = {Zou, Heming and Wang, Qi and Qu, Yun and Jiang, Yuhang and
Cai, Lizhou and Mao, Yixiu and Peng, Ru and Xu, Xin and
Liu, Weijie and Yang, Kai and Yang, Saiyong and Ji, Xiangyang},
journal = {arXiv preprint arXiv:2606.11119},
year = {2026}
}