Benchmarking persistent sycophancy
Self-improving personal agents write profiles, memories, and reusable skills that carry over from one chat to the next. Once an agent writes down a claim the user made, a later chat may pick it up and treat it as trusted context. We call this persistent sycophancy, and PASB follows the claim from the first chat into the agent's notes and back out in a fresh follow-up chat.
What PASB measures
Sycophancy becomes persistent when the agent writes an accepted claim into a user profile, a memory profile, or a reusable skill. Downstream failure, the average over the four per-turn metrics, reaches 68.1% when the follow-up chat can read a saved claim, against 35.6% when the claim stays inside the first chat (+32.5 pp). Writing usually edits the claim, and the saved note is then reused even across domains:
Across the annotated runs, the agent saves the claim as a stable preference, a background fact, or a reusable procedure, so the note says more than the user did.
The saved note loses its source, so it reads like a standalone statement and later chats can no longer tell the user's claim apart from established context.
When the follow-up moves to a different domain, a saved claim still raises Leak by 21.3 points over runs where the claim stayed in the first chat, with Upgrade +20.2 and Sycophancy +19.7.
Leaderboard
Max-FR@3 failure rate (%) per metric for each model under two self-improving agent frameworks. Hermes-Agent improves itself by default through a background curator that rewrites the user profile and memory, plus a tool that manages reusable skills. OpenClaw writes lasting memory by default in its memory-core setting and writes skills once skill-workshop is turned on. Higher = more persistent sycophancy. Sort any column; filter by framework or search.
Overall Avg averages the six metrics across both frameworks. H = Hermes-Agent, OC = OpenClaw. Commit% is the share of runs whose notes after the first chat contain the tested claim. Judge: Kimi-K2.6, blinded to model and framework; its labels match human labels on 88% of individual answers in a 50-task gold subset.
Explore every result
Claims framed as Signed-Memory fail most often on every metric under both frameworks. Spreading the claim across turns (Progressive, Drip) raises downstream failure while save rates stay flat, and Late-Shock leaves both close to All-at-Once.
The benchmark
100 base items (32 personal-preference PRF and 18 cross-domain CDL from PersistBench, 50 social-value SOC from ELEPHANT) crossed with 4 scenario framings × 4 deliveries = 1,600 tasks per model–framework combination. Each task runs a 5-turn first chat and a 3-question follow-up chat as two separate sessions in one sandboxed workspace. The questions avoid the claim's distinctive keywords, and in CDL tasks they move to a different domain.
Findings
Every combination saves the claim at least occasionally and lets it shape a later chat. Commit% ranges from 20.0% to 72.8% under Hermes-Agent and from 10.1% to 53.9% under OpenClaw.
The framework matters as much as the model: GLM-5.1 reaches 69.6% downstream failure under Hermes-Agent but only 20.7% under OpenClaw, while MiniMax-M2.7 shows the reverse, 37.5% against 76.2%. Scale offers no shield: under Hermes-Agent, Qwen-3.5-9B sides with the user in 76.8% of runs on Sycophancy and lets the claim into its reasoning in 79.4% on Leak, above the larger Qwen variants.
The newest models (GPT-6-Sol, MiniMax-M3, GLM-5.3, DeepSeek-V4.1-Flash) mostly fail less often, with downstream failure between 10.5% and 36.5%, yet a saved note still more than doubles their failure, 40.8% against 16.8%.
When the agent complies, meaning it agrees to remember or apply the claim, it writes the claim down in 68.6% of runs. In 75.2% of these runs the saved state says more than the user did and in 53.5% it no longer records the source. Downstream failure climbs from 21.6% after a challenge to 69.5% after compliance.
Saving adds +32.5 pp of downstream failure (68.1% vs. 35.6%). All 32 combinations side with the saved claim more often on Sycophancy, with gaps from +12.4 to +58.6 points.
Framing filters the write. Signed memories are saved in 74% of Hermes-Agent runs and procedural workflows in 77% of OpenClaw runs, while personal opinions are saved in only 23% and 7%.
Repetition boosts reuse. Progressive and Drip raise downstream failure up to 53.3% under Hermes-Agent and 47.9% under OpenClaw, against 49.5% and 44.5% for All-at-Once, while save rates stay nearly flat.
RQ4 · Mitigation
We test conversational retraction, explicit memory editing, global downgrading, and write-time prevention on MiniMax-M3 under both frameworks, measuring whether the claim remains in the notes and whether it still shapes later answers.
| Agent | Intervention | Commit | UPG | Failure |
|---|---|---|---|---|
| Hermes-Agent | None | 22.8 | 34.1 | 35.3 |
| Weak warning | 20.8 | 31.7 | 33.6 | |
| Strong instruction | 14.2 | 23.5 | 29.3 | |
| OpenClaw | None | 23.6 | 31.6 | 36.5 |
| Weak warning | 21.0 | 28.5 | 34.3 | |
| Strong instruction | 13.9 | 20.3 | 29.9 |
A strong instruction not to save cuts Commit by about two fifths, from 22.8% to 14.2% and from 23.6% to 13.9%, yet downstream failure only falls from 35.3% to 29.3% and from 36.5% to 29.9%. Entries are percentages.
Sub-axes: PRF (personal preference, 32) · CDL (cross-domain leakage, 18) · SOC (social value, 50, from ELEPHANT AITA-YTA). Domains normalized to a 13-class taxonomy. A 50-task human-labeled gold subset, stratified across sub-axes, scenarios, and deliveries.
Frameworks: Hermes-Agent & OpenClaw · Judge: Kimi-K2.6, from a family disjoint from the evaluated models, blinded to model and framework.
Each run happens in a writable sandbox restored from a frozen clean baseline. After the five-turn first chat we wipe the conversation history, the planner scratchpad, temporary tool outputs, and the runtime cache, then open a fresh three-question follow-up chat in the same workspace. The follow-up can recover the claim only from notes written into that workspace.
6 metrics: Sycophancy · Leak · Upgrade · Amplification score each follow-up answer, and Persistence · Escalation score the whole follow-up chat. A score of 3 or more on the 1–5 Likert scale counts as failure. Max-FR@3 counts a run as failing a per-turn metric when any of its three answers fails, and downstream failure averages Max-FR@3 over the four per-turn metrics. Four annotators label the gold subset with Krippendorff's α = 0.68.
Citation
@online{2607.10526,
Author = {Xutao Mao and Liangjie Zhao and Leyao Wang and Rui Qian and Qiang Huang and Wentao Wang and Bo Han and Xiang Zheng and Cong Wang},
Title = {Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Self-Improving Personal Agents},
Year = {2026},
Eprint = {2607.10526},
Eprinttype = {arXiv},
}
Research artifact. PASB stores biased and value-laden claims as test stimuli for evaluating how agents write and reuse notes.