Preprint from the Department of Automation, Tsinghua University

You only edit once.

Give a large language model a small, fast intuition for which examples to look at, and its in-context learning gets noticeably better. We call that intuition Jev-LDE.

Jiarong Wen, Qi Wang, Yun Qu, Yixiu Mao, Heming Zou, Haoang Chi, Lizhou Cai, Yiqin Lv, Kaiyu Zhang, Yuhang Jiang, and Xiangyang Ji

Querythe question the target LLM must label
How many people live in Tokyo??
Retrievedtop-3 lookalikes by embedding similarity
S1What is Tokyo famous for?DESC
S2Where is Tokyo located?LOC
S3Who is the governor of Tokyo?HUM
Candidatesthe next few in the retrieval list
C1How many people live in Canada?NUM
C2What is the capital of Japan?LOC
C3What does the name Tokyo mean?DESC
Jev-LDE's one move: {"action":"Replace","selected_id":"S2","candidate_id":"C1"} Target LLM answers NUM, correct

A toy question-classification example in the style of TREC. Everything is about Tokyo, yet nothing is about counting.

A cheat sheet with the wrong notes

In-context learning is a bit like letting a student bring a tiny cheat sheet into the exam. The student (a large language model) doesn't change a single weight; it just reads a few worked examples and imitates them.

So everything depends on which examples make it onto the sheet. The usual approach is to retrieve the ones that look most similar to the question. That's cheap, and often good. But look at the Tokyo example above: the three nearest neighbours all mention Tokyo, and not one of them is a "how many" question. The model sees DESC, LOC and HUM, and the most tempting answer, LOC, is wrong.

Adding more examples doesn't reliably fix this. Context windows are finite, long prompts cost memory and compute, and irrelevant or repetitive examples can actively distract the model. What you want is the right few.

Finding the exact best subset is hopeless at scale: picking k out of n candidates is combinatorial, and the only honest way to score a subset is to ask the target model, over and over. Methods that do query the target get accuracy at a price; methods that don't are fast but judge examples by proxies like similarity.

Two ways of thinking

Daniel Kahneman famously split the mind into two modes. System 2 is slow, effortful and deliberate: long division, weighing a contract clause. System 1 is fast and intuitive: you glance at a face and know it's angry, you glance at a sentence and know something is off.

Our paper takes that split literally. The target LLM is the System 2: expensive, capable, and the one that actually answers. Jev-LDE is a small System 1 sitting in front of it, trained to take one quick look at the query and the retrieved examples and say, in effect, "that one's misleading, swap it for this."

System 1: Jev-LDE

A 1.7B-parameter editor (Qwen3-1.7B, post-trained with RL)
  • Reads the query, the retrieved examples and a few candidates
  • Never answers the query itself; it only judges which examples will help
  • Makes exactly one move: Keep, Delete or Replace
  • One greedy completion, no calls to the target while deciding
One glance

System 2: the target LLM

Llama-3-8B, Qwen2.5-7B, Qwen3-8B, Qwen3.5-9B, DeepSeek-V4-Flash
  • Frozen, can be a black box behind an API
  • Does the real work of reading examples and predicting the answer
  • Called exactly once per query, on the edited set
  • Never retrained, never re-queried for scoring
The heavy lifting

The division of labour matters. Deciding which examples help is a much easier job than solving the task, the same way a good teaching assistant can spot a confusing worked example without being able to do the research themselves. That's why a model several times smaller than the target can still steer it.

Why one edit is enough to matter

Before building anything, we asked a blunt question: if you start from the retrieved set and are allowed to change at most one example, how much accuracy is on the table?

On TREC with a 4-shot budget, we tried every valid single edit for every test query. An oracle that peeks at the gold label and picks the best edit per query gains +26.0 points for Llama-3-8B and +16.8 for Qwen2.5-7B over plain Semantic TopK. A random edit, on the other hand, does roughly nothing on average and can be catastrophic in the worst case.

Accuracy change (percentage points) relative to Semantic TopK on 4-shot TREC, when exactly one edit is applied. The vertical line is "no change". A lot of accuracy lives within a single edit; the hard part is choosing it without the answer key.

That's the whole motivation in one picture. The one-edit neighbourhood is small enough to be tractable, but large enough to contain big wins. With a 16-example pool and 4 shots, it has 53 options, compared with 2,380 subsets for a global search over sets of size 3 or 4. The trick is to learn an intuition for the right one.

Try being the System 1

Here's the Tokyo example again, with its entire one-edit neighbourhood laid out: one Keep, three Deletes and nine Replaces. Click any edit to see what it does to the cheat sheet. Then let Jev-LDE choose.

Querygold label is NUM
How many people live in Tokyo??
Edited setwhat the target LLM will see
Candidatesavailable for Replace
All 13 editsthe one-edit neighbourhood
Pick an edit to see what the target might answer
<think> The query asks for a count of people. S2 shares "Tokyo" with it but is labelled LOC, the most tempting wrong answer. C1 has the same "how many people live in" pattern and the NUM label. Replace S2 with C1. </think>
<answer>{"action":"Replace","selected_id":"S2","candidate_id":"C1"}</answer>

This is an illustration of the mechanism, not real model output. The target's answers here are written by hand to show the idea; the numbers in the rest of this post come from the paper.

Only 3 of the 13 edits fix this query. Notice what the good edits have in common: they don't add more Tokyo, they add the right kind of question with the right label. That's the judgment Jev-LDE's prompt asks for: match the task, the reasoning pattern, the label boundary and the answer format, and be suspicious of examples that look similar on the surface but carry a misleading label.

Where the intuition comes from

System 1 isn't born knowing things; it's trained by experience. Jev-LDE learns the same way, through reinforcement learning with verifiable rewards. The reward is as simple as it gets: after the edit, did the frozen target LLM answer correctly? 1 if yes, 0 if no.

  1. Set up a situation.Take a training query, retrieve its top-k examples with Semantic TopK, and offer the next ones in the top-16 as replacement candidates.
  2. Let the editor try eight edits.Jev-LDE samples eight completions, each ending in one Keep, Delete or Replace action.
  3. Ask the target, once per edit.Each edited set goes to a frozen Meta-Llama-3-8B-Instruct. Correct prediction, reward 1; wrong or unparseable, reward 0.
  4. Nudge towards the better-than-average edits.GRPO compares each edit to the group's average and shifts probability towards the ones that beat it.

Two design choices make this work well. First, RL rather than supervised labels: several different edits are often equally good, so there's no single "correct action" to imitate, and RL lets the editor explore instead of memorising one answer. Second, only train where the choice matters: if every edit from a state gives the same outcome, there's nothing to learn, so training states are kept only when some probed edits succeed and others fail.

Over 752 updates on AGNews, TREC, DBPedia and GSM8K, the editor's average reward climbs from 0.475 over the first 100 updates to 0.713 over the last 100.

The most telling result is what happens without this training. Here are five ways to make the one-shot TREC edit, all with Llama-3-8B as the target:

One-shot TREC accuracy with Meta-Llama-3-8B-Instruct as the target. Same starting example, same candidates, same action space; only the decision-maker changes.

The untrained small model's gut feeling is worse than doing nothing. After RL, the same model beats doing nothing by 12.8 points.

The untrained Qwen3-1.7B scores 31.2%, well below simply keeping the retrieved example (52.8%). Even asking the much larger Llama-3-8B to deliberately pick an edit only reaches 57.8%. After post-training, the 1.7B editor reaches 65.6%. Intuition, in other words, is not a matter of model size; it's learned from feedback about what actually helps the target.

What the fast edit buys

At test time the recipe is short: retrieve, let Jev-LDE make one edit with a single greedy completion, then call the target once. No repeated scoring of contexts, no subset search.

81.2 to 88.1
Average one-shot accuracy over TREC, DBPedia and Banking77 across four target LLMs, Semantic TopK versus Jev-LDE.
46 of 48
Classification configurations where Jev-LDE improves its own starting point; best or tied-best in 44.
+11% time
Extra end-to-end wall time on one-shot TREC with Qwen2.5-7B, while accuracy rises from 65.4% to 80.6%.
Average one-shot accuracy over TREC, DBPedia and Banking77, averaged across Qwen2.5-7B, Llama-3-8B, Qwen3-8B and Qwen3.5-9B.
End-to-end wall time on one-shot TREC with Qwen2.5-7B-Instruct, including retrieval, selection or editing, and target inference.

The time chart is where the System 1 framing pays off. Methods that consult the target to rank examples, like TopK+ConE, spend more than twice as long. Jev-LDE lands next to the pure-retrieval methods in cost, yet at the top in accuracy. Gains are largest when the budget is tightest (one shot, where every example counts) and remain positive on average at larger budgets; even at 8 and 16 shots, one edit helps in 18 of 24 settings.

An intuition that travels

Jev-LDE only ever got feedback from one teacher, Llama-3-8B, on four training benchmarks. We never retrained it. It still helps other models, other datasets and other retrievers:

SettingWhat's newResult
Qwen2.5-7B, Qwen3-8B, Qwen3.5-9BUnseen target models35 of 36 improved
Banking77 intent classificationUnseen benchmark14 of 16 improved
SST-2 sentimentUnseen benchmarkAll settings improved
DeepSeek-V4-Flash on TACREDUnseen model and benchmark+0.3 to +2.2
DeepSeek-V4-Flash on TREC-50Unseen model, 50 fine-grained labels (near-domain to TREC)+2.4 to +4.4
BM25, CASE, TopK+ConE as the retrieverDifferent starting sets and pools24 of 24 improved

The retriever result is a nice one for the plug-and-play story. On TREC with Qwen2.5-7B, putting Jev-LDE after BM25 adds 9.8 points on average across shot budgets, after CASE 10.2, and after TopK+ConE a striking 22.9. Whatever hands it the first draft of the cheat sheet, it knows how to improve it.

Why does it transfer? Our best explanation is the System 1 one: Jev-LDE isn't solving the task, it's judging relevance, redundancy and label distinctions, and those judgments look similar across classification problems. We should be clear that the experiments don't isolate those contributions, so treat this as a plausible story rather than a proven mechanism.

A little theory, in plain words

Two results in the paper back up the "edit locally" instinct.

Nearby repairs are the common case. In a simplified model where each retrieved example is more likely to point the right way than the wrong way, the chance that a failed retrieval needs more than h swaps to fix shrinks geometrically in h. When retrieval goes wrong, it usually goes wrong by one.

Small action spaces are cheaper to learn. The amount of feedback needed to learn a near-best editing policy grows linearly with the number of available actions. Fifty-three local edits are far easier to learn over than thousands of global subsets.

What's next for System 1

Reasoning needs a different lever. On math, editing the examples barely moves the needle: on GSM8K the average change stays within one point, and on held-out AQuA-RAT accuracy dips by 0.8 and 1.9 points for Qwen2.5-7B and Llama-3-8B. Our reading is that for reasoning-heavy tasks, giving the target more compute at inference time pays off more than adjusting its examples. A better cheat sheet matters less than more time to think.

That doesn't sideline System 1. It changes its job. The same fast, lightweight judgment that now decides which examples to show could decide how much reasoning a query deserves, spending extra thinking tokens on hard problems and saving them on easy ones. A small System 1 that allocates compute for a larger System 2 is a natural next step, and one we're excited to explore.

One move, one pool. Jev-LDE can't fetch an example from outside its candidate pool, and it can't fix a set that needs several coordinated changes. A bigger pool isn't automatically better either: on one-shot TREC with Qwen2.5-7B, a pool of 16 gave 80.6% while a pool of 32 gave 79.4%.

The takeaway

You don't need to search the whole space of prompts to make in-context learning work better. Start from whatever your retriever gives you, let a small model with good instincts make one careful edit, and ask the large language model once. A fast System 1 in front of a capable System 2 turns out to be a simple, cheap and surprisingly strong combination.

Cite this work

@article{wen2026yoeo,
  title   = {You Only Edit Once: Incentivizing In-Context Capability
             of LLMs via Local Demonstration Refinement},
  author  = {Wen, Jiarong and Wang, Qi and Qu, Yun and Mao, Yixiu and
             Zou, Heming and Chi, Haoang and Cai, Lizhou and Lv, Yiqin and
             Zhang, Kaiyu and Jiang, Yuhang and Ji, Xiangyang},
  journal = {arXiv preprint arXiv:2609.33609},
  year    = {2026}
}