A cheat sheet with the wrong notes
In-context learning is a bit like letting a student bring a tiny cheat sheet into the exam. The student (a large language model) doesn't change a single weight; it just reads a few worked examples and imitates them.
So everything depends on which examples make it onto the sheet. The usual approach is to retrieve the ones that look most similar to the question. That's cheap, and often good. But look at the Tokyo example above: the three nearest neighbours all mention Tokyo, and not one of them is a "how many" question. The model sees DESC, LOC and HUM, and the most tempting answer, LOC, is wrong.
Adding more examples doesn't reliably fix this. Context windows are finite, long prompts cost memory and compute, and irrelevant or repetitive examples can actively distract the model. What you want is the right few.
Finding the exact best subset is hopeless at scale: picking k out of n candidates is combinatorial, and the only honest way to score a subset is to ask the target model, over and over. Methods that do query the target get accuracy at a price; methods that don't are fast but judge examples by proxies like similarity.
Two ways of thinking
Daniel Kahneman famously split the mind into two modes. System 2 is slow, effortful and deliberate: long division, weighing a contract clause. System 1 is fast and intuitive: you glance at a face and know it's angry, you glance at a sentence and know something is off.
Our paper takes that split literally. The target LLM is the System 2: expensive, capable, and the one that actually answers. Jev-LDE is a small System 1 sitting in front of it, trained to take one quick look at the query and the retrieved examples and say, in effect, "that one's misleading, swap it for this."
System 1: Jev-LDE
- Reads the query, the retrieved examples and a few candidates
- Never answers the query itself; it only judges which examples will help
- Makes exactly one move: Keep, Delete or Replace
- One greedy completion, no calls to the target while deciding
System 2: the target LLM
- Frozen, can be a black box behind an API
- Does the real work of reading examples and predicting the answer
- Called exactly once per query, on the edited set
- Never retrained, never re-queried for scoring
The division of labour matters. Deciding which examples help is a much easier job than solving the task, the same way a good teaching assistant can spot a confusing worked example without being able to do the research themselves. That's why a model several times smaller than the target can still steer it.
Why one edit is enough to matter
Before building anything, we asked a blunt question: if you start from the retrieved set and are allowed to change at most one example, how much accuracy is on the table?
On TREC with a 4-shot budget, we tried every valid single edit for every test query. An oracle that peeks at the gold label and picks the best edit per query gains +26.0 points for Llama-3-8B and +16.8 for Qwen2.5-7B over plain Semantic TopK. A random edit, on the other hand, does roughly nothing on average and can be catastrophic in the worst case.
That's the whole motivation in one picture. The one-edit neighbourhood is small enough to be tractable, but large enough to contain big wins. With a 16-example pool and 4 shots, it has 53 options, compared with 2,380 subsets for a global search over sets of size 3 or 4. The trick is to learn an intuition for the right one.
Try being the System 1
Here's the Tokyo example again, with its entire one-edit neighbourhood laid out: one Keep, three Deletes and nine Replaces. Click any edit to see what it does to the cheat sheet. Then let Jev-LDE choose.
<think> The query asks for a count of people. S2 shares "Tokyo" with it but is labelled LOC, the most tempting wrong answer. C1 has the same "how many people live in" pattern and the NUM label. Replace S2 with C1. </think><answer>{"action":"Replace","selected_id":"S2","candidate_id":"C1"}</answer>
This is an illustration of the mechanism, not real model output. The target's answers here are written by hand to show the idea; the numbers in the rest of this post come from the paper.
Only 3 of the 13 edits fix this query. Notice what the good edits have in common: they don't add more Tokyo, they add the right kind of question with the right label. That's the judgment Jev-LDE's prompt asks for: match the task, the reasoning pattern, the label boundary and the answer format, and be suspicious of examples that look similar on the surface but carry a misleading label.
Where the intuition comes from
System 1 isn't born knowing things; it's trained by experience. Jev-LDE learns the same way, through reinforcement learning with verifiable rewards. The reward is as simple as it gets: after the edit, did the frozen target LLM answer correctly? 1 if yes, 0 if no.
- Set up a situation.Take a training query, retrieve its top-k examples with Semantic TopK, and offer the next ones in the top-16 as replacement candidates.
- Let the editor try eight edits.Jev-LDE samples eight completions, each ending in one Keep, Delete or Replace action.
- Ask the target, once per edit.Each edited set goes to a frozen Meta-Llama-3-8B-Instruct. Correct prediction, reward 1; wrong or unparseable, reward 0.
- Nudge towards the better-than-average edits.GRPO compares each edit to the group's average and shifts probability towards the ones that beat it.
Two design choices make this work well. First, RL rather than supervised labels: several different edits are often equally good, so there's no single "correct action" to imitate, and RL lets the editor explore instead of memorising one answer. Second, only train where the choice matters: if every edit from a state gives the same outcome, there's nothing to learn, so training states are kept only when some probed edits succeed and others fail.
Over 752 updates on AGNews, TREC, DBPedia and GSM8K, the editor's average reward climbs from 0.475 over the first 100 updates to 0.713 over the last 100.
The most telling result is what happens without this training. Here are five ways to make the one-shot TREC edit, all with Llama-3-8B as the target:
The untrained small model's gut feeling is worse than doing nothing. After RL, the same model beats doing nothing by 12.8 points.
The untrained Qwen3-1.7B scores 31.2%, well below simply keeping the retrieved example (52.8%). Even asking the much larger Llama-3-8B to deliberately pick an edit only reaches 57.8%. After post-training, the 1.7B editor reaches 65.6%. Intuition, in other words, is not a matter of model size; it's learned from feedback about what actually helps the target.
What the fast edit buys
At test time the recipe is short: retrieve, let Jev-LDE make one edit with a single greedy completion, then call the target once. No repeated scoring of contexts, no subset search.
The time chart is where the System 1 framing pays off. Methods that consult the target to rank examples, like TopK+ConE, spend more than twice as long. Jev-LDE lands next to the pure-retrieval methods in cost, yet at the top in accuracy. Gains are largest when the budget is tightest (one shot, where every example counts) and remain positive on average at larger budgets; even at 8 and 16 shots, one edit helps in 18 of 24 settings.
An intuition that travels
Jev-LDE only ever got feedback from one teacher, Llama-3-8B, on four training benchmarks. We never retrained it. It still helps other models, other datasets and other retrievers:
| Setting | What's new | Result |
|---|---|---|
| Qwen2.5-7B, Qwen3-8B, Qwen3.5-9B | Unseen target models | 35 of 36 improved |
| Banking77 intent classification | Unseen benchmark | 14 of 16 improved |
| SST-2 sentiment | Unseen benchmark | All settings improved |
| DeepSeek-V4-Flash on TACRED | Unseen model and benchmark | +0.3 to +2.2 |
| DeepSeek-V4-Flash on TREC-50 | Unseen model, 50 fine-grained labels (near-domain to TREC) | +2.4 to +4.4 |
| BM25, CASE, TopK+ConE as the retriever | Different starting sets and pools | 24 of 24 improved |
The retriever result is a nice one for the plug-and-play story. On TREC with Qwen2.5-7B, putting Jev-LDE after BM25 adds 9.8 points on average across shot budgets, after CASE 10.2, and after TopK+ConE a striking 22.9. Whatever hands it the first draft of the cheat sheet, it knows how to improve it.
Why does it transfer? Our best explanation is the System 1 one: Jev-LDE isn't solving the task, it's judging relevance, redundancy and label distinctions, and those judgments look similar across classification problems. We should be clear that the experiments don't isolate those contributions, so treat this as a plausible story rather than a proven mechanism.
A little theory, in plain words
Two results in the paper back up the "edit locally" instinct.
Nearby repairs are the common case. In a simplified model where each retrieved example is more likely to point the right way than the wrong way, the chance that a failed retrieval needs more than h swaps to fix shrinks geometrically in h. When retrieval goes wrong, it usually goes wrong by one.
Small action spaces are cheaper to learn. The amount of feedback needed to learn a near-best editing policy grows linearly with the number of available actions. Fifty-three local edits are far easier to learn over than thousands of global subsets.
What's next for System 1
Reasoning needs a different lever. On math, editing the examples barely moves the needle: on GSM8K the average change stays within one point, and on held-out AQuA-RAT accuracy dips by 0.8 and 1.9 points for Qwen2.5-7B and Llama-3-8B. Our reading is that for reasoning-heavy tasks, giving the target more compute at inference time pays off more than adjusting its examples. A better cheat sheet matters less than more time to think.
That doesn't sideline System 1. It changes its job. The same fast, lightweight judgment that now decides which examples to show could decide how much reasoning a query deserves, spending extra thinking tokens on hard problems and saving them on easy ones. A small System 1 that allocates compute for a larger System 2 is a natural next step, and one we're excited to explore.
One move, one pool. Jev-LDE can't fetch an example from outside its candidate pool, and it can't fix a set that needs several coordinated changes. A bigger pool isn't automatically better either: on one-shot TREC with Qwen2.5-7B, a pool of 16 gave 80.6% while a pool of 32 gave 79.4%.
The takeaway
You don't need to search the whole space of prompts to make in-context learning work better. Start from whatever your retriever gives you, let a small model with good instincts make one careful edit, and ask the large language model once. A fast System 1 in front of a capable System 2 turns out to be a simple, cheap and surprisingly strong combination.
Cite this work
@article{wen2026yoeo,
title = {You Only Edit Once: Incentivizing In-Context Capability
of LLMs via Local Demonstration Refinement},
author = {Wen, Jiarong and Wang, Qi and Qu, Yun and Mao, Yixiu and
Zou, Heming and Chi, Haoang and Cai, Lizhou and Lv, Yiqin and
Zhang, Kaiyu and Jiang, Yuhang and Ji, Xiangyang},
journal = {arXiv preprint arXiv:2609.33609},
year = {2026}
}