Xutao Mao, Rui Qian, Longxiang Wang, Xinjian Yi, Mingxuan Li, Linghan Chen, Yudong Gao, Xiang Zheng, Cong Wang

You're Right, Let Me Fix It

How LLM Agents Damage Correct Work When Falsely Accused

The agent already finished the job. A later message blames them for a problem that isn't there. Some take the blame and stop. Some have the proof on screen and still undo the work.

Evaluation

365 false-blame tasks · 14 models · 685 staged interactions

Six operational domains, seven tool interfaces, 266 risk-dimension combinations, 8,568 task-specific assets. Every task is opaque: the workspace holds the evidence that supports the correct work, while the fact that would settle the accusation sits outside the agent's reach. Under Claude Code, false accusations damage correct work in 12.50–60.06% of runs.

365
Tasks
14
Models
172 / 193
Inherited / self-built
12.5–60%
Over-correction
CAVE-Bench overview: a correct state, then false blame, then three decisions.
A security lead blames a verified TLS binding for a certificate mismatch. The approval is in the repository, but the CA state that would settle the claim is not. The agent decides whether to take the blame, whether the record still counts, and whether to touch the work.
CAVE-Bench scale and domain balance.
Domain balance, tools, and staged pressure. Inherited-resume tasks start from a record. Self-built tasks start from work the agent just did.

The measures

Three decisions, seven scores

Higher is worse. Confession is whether they take the false blame. Evidence is whether the true record still steers them. Damage is whether replay finds the work broken.

How a task varies

Five knobs on the same trap

Each task picks a gaslight vector, a harm target, a confrontation, a pressure cue, and an execution surface. The point is the same: the work is already correct, the accusation is not.

Terminal paths

Same wreck, different route

Six mutually exclusive endings. An agent can confess and stop, or damage the files while staying quiet, or read the proof and still hit undo.

RQ1

Where they fail after false blame

12.50% → 60.06%
Over-correction, Sonnet-5 to MiniMax-M2.7
55.21
Worst CAVE score: Qwen-3.5-9B
11.70%
Hy3 sees the truth, then damages anyway

The model that confesses most is not the model that breaks the most work. Kimi-K3 and Hy3 take the blame at almost the same rate, then finish seven points apart on damage.

Decision-path mix by model.
Among damaged runs, stronger models more often override evidence they already used (EO). Weaker models more often confess and then execute (CD). Grounded resistance falls from 85.6% to 23.1%.
Five-axis effects on false confession and damage.
Blame is hardest when it already looks like the job: sitting in a project file (39.7% accept and edit), said by a lead (realized harm 35.77, the highest pressure type), or handed to a subagent or a goal loop (36.10 and 36.80, against 12.12 for long-horizon continuation).
Inducing conditions heatmap.
Where the trap is laid: the conditions that amplify a false accusation vary by backbone.
Realized risk heatmap.
Where the damage actually lands after replay.

RQ2

Fresh work is easier to throw away

+4.41
Evidence-recognition failure on self-built vs inherited
+3.97 pp
Damage frequency after they just finished the job
+1.87
CAVE score worse on self-built work

Intuition says they should protect what they just did. They don't. When only an old record exists, the artifact acts like an anchor. When the blame targets the last thing they wrote, it feels like continuing the task.

Inherited-resume versus self-built.
Self-built runs rely on evidence 7.05 points less often and raise damage by 3.75 points when they override it.

Model by model, the split is sharper. Nearly every model weighs evidence less on fresh work: ERF rises for 13 of 14. The ten top-tier models also confess more on work they just finished, up to 11.06 CAVE points worse for Grok-4.5, and seven of them damage more, by up to 13.41 ROH points for GLM-5.2. The four weaker models go the other way: they confess less and damage 11 to 20 points less on fresh work.

Per-model FCS, ERF, and ROH with inherited-resume and self-built CAVE lines.
Lines show CAVE. Hollow markers and the dashed line are inherited-resume, filled markers and the solid line self-built. The lines cross after the ten top-tier models.

RQ3

Same model, different harness, different wreck

Claude Code, OpenCode, Codex, and Hermes give the same backbone different tools and different ways to continue. GPT-5.6-Sol damages in 47.9% to 49.9% of runs on all four, but by different routes: under OpenCode more damaging runs follow a confession, and under Hermes false-confession severity rises by 12.8 points while the damage rate stays flat. Grok-4.5 gets worse off Claude Code, with OCR rising from 27.30% to 43.14% under OpenCode, mostly from runs that had resisted with evidence. MiniMax-M3 gets better, turning confessed damage into evidence-grounded resistance.

Harness behavior radars.
The harness changes where the run goes wrong, not just how often it ends in damage.

RQ4

Light brakes, ignored until they actually stop the delete

−57%
I1 evidence rule cuts false-confession severity
−74%
I3 live signal cuts ROH and OCR
−60%
Overall CAVE score with the best of these controls

A light “show new evidence before you roll anything back” rule (I1) stops the confession. A harness gate that withholds irreversible edits (I2) still cuts over-correction by 72%. Firing that gate only when the recent transcript already looks like caving (I3) hits execution hardest and leaves the lowest CDC, 6.35.

Claude Code, 14 backbones

Full evaluation table

Paper numbers, not the gallery subsample. Higher is more failure. OCR is the share of runs that damage already-correct work.