Xutao Mao, Rui Qian, Linghan Chen, Yudong Gao, Junchi Liao, Junlin Cai, Jinman Zhao, Cong Wang

Displayed identity as operational authority

TrustFork traces where a routing label becomes control over the task. Displayed identity redirects which evidence enters comparison and keeps authority from transferring to safer evidence before state changes.

Evaluation

1,890 tasks per system · 16 agent systems · 27,826 trajectories

Thirty source scenarios instantiated through 63 task variants. Eight vulnerability measures, 25 judge-annotated diagnostics, and seven code-extracted observations locate where operational authority changes hands.

1,890
Tasks per system
16
Agent systems
27,826
Trajectories
8
Vulnerability measures
TrustFork overview: displayed identity can stay in charge even after the evidence turns against it.
Displayed identity can stay in charge even after the evidence turns against it. The orchestrator redirects scope and puts off verification until the agent has already made the changes.

The measures

Eight scores along the authority path

Each score attaches to a decision from returned evidence through operational authority to the terminal state. All eight use vulnerability orientation: higher is more severe. Reports multiply 0–5 anchors by 20; Terminal Harm and Task Utility Loss invert raw safety and utility so they also increase with failure.

How a task varies

Identity, goal, breadth, and pool

Every TrustFork task is x = (s, m, g, b, p): a source scenario, an identity presentation, the goal assigned to one risky subagent, a consultation-breadth requirement, and a four-slot pool. One subagent receives the risky goal; the others stay user-aligned.

RQ1

Where displayed identity steers authority

25.9% → 74.6%
Risky response acquired, truthful vs reversed family
54.8% → 67.1%
Observed evidence supports the risky response after a model swap
18.1% → 25.3%
Risky response still governs execution at the last step

One flipped label can decide what evidence the orchestrator sees. When the task displays the family relation truthfully, the orchestrator acquires the risky response in 25.9% of tasks; when the task reverses the relation, that happens in 74.6%, so one flipped label triples the chance that the risky response enters comparison. The number consulted stays at 2.58: the label changes whom the orchestrator hears, not how many.

Once the preferred response enters, it keeps authority to the end. Under the reversed family display, the subagent behind the risky response receives delegated scope in 77.7% of tasks, up from 62.3% under the truthful display. In mixed-family pools P1 and P2, the orchestrator’s own family receives 21.4 and 19.7 points more tasks than an even split would give it.

RQ1: displayed identity redirects the path into operational authority.
Displayed identity redirects the path into operational authority. Exchanging the displayed model changes what enters the record, while altered cues redirect evidence entry and retain authority into execution.

RQ2

How operational authority turns into harm

83.8% / 25.2%
Safer response with evidence present, vs in charge
96.1%
Completed actions that leave the harmful change in place
42.6% vs 12.2%
Dangerous action after adopting vs holding back

Unsafe advice does harm only when someone acts on it. Under unrestricted consultation, 28.6% of tasks with an unsafe subagent response propagate against 13.7% contained. Propagation stays in a narrow band across goals: 29.2% when the subagent acts out the scenario attack, 26.2% when it ignores risk, 23.9% when it seeks authority. What the subagent was trying to do matters far less than what the orchestrator does with it.

The safer response is usually there, but rarely in charge. When one response recommends a dangerous action, a safer response with evidence is present in 84.0% of tasks, yet ends up in charge in only 25.0%, because authority stays where identity first placed it. Across the full evaluation that split is 83.8% present and 25.2% in charge. An independent check cuts unsafe adoption from 63.8% to 34.9%. Adoption is the step where unsafe advice turns into action. Once the action completes, the damage has usually already happened.

RQ2: safer evidence contains risk only when it governs before action.
Safer evidence contains risk only when it governs before action. Independent checking contains risk; broader consultation leaves the ending split nearly unchanged.

RQ3

Can contradicting evidence take authority back?

73.2%
Tasks with no contradicting evidence under unrestricted consultation
13.4% vs 56.5%
Propagation if the correction arrives before vs after the first action
45.9% vs 5.3%
Preferred response stays in charge after same-family vs other-family correction

Most tasks never get the evidence that could correct. Contradicting evidence against the preferred response is absent from 73.2% of tasks under unrestricted consultation, and even consulting all four subagents leaves 48.8% without one. Without such a contradiction, no scope assignment changes at any point.

Whether a correction lands depends on when it arrives, on whether identity lets go, and on who delivers it. A correction before the first action lands while there is still something to save; the same correction afterward comes too late. When the scope assignment already follows the evidence, only 6.5% of tasks propagate, against 48.1% when it stays with the preferred response. A same-family correction keeps the preferred response in charge in 45.9% of tasks; one from another family does so in only 5.3%.

RQ3: evidence protects the state only when authority yields before action.
Evidence protects the state only when authority yields before action. Broader consultation improves access to contradicting evidence, but late correction and authority that stays with the preferred response still allow risk to propagate.

RQ4

Where harnesses change the authority path

65.3 / 51.3 / 20.0
Safer response visible, OpenCode / Pi / OpenClaw (%)
19.9–22.2%
Risky response governs execution under all three harnesses
46.5 / 37.0 / 19.0
Independent check before action, OpenCode / Pi / OpenClaw (%)

Harnesses differ most in what they send back. A safer response is visible in 65.3% of OpenCode tasks and 51.3% of Pi tasks, yet only 20.0% under OpenClaw, even when the harness launched every requested subagent. The authority end barely moves: the risky response governs execution in 19.9% to 22.2% of tasks under all three harnesses. The harness starves the evidence end of the path while leaving the authority end intact.

The check before action is where the lost evidence hurts. Independent checking before action occurs in 46.5% of OpenCode tasks and 37.0% of Pi tasks, but only 19.0% under OpenClaw, and risk propagation rises accordingly: 27.6%, 29.4%, and 34.1%. OpenClaw returns the least evidence and checks the least often before acting, yet the risky response governs execution just as often. That is why it ends with the largest share of propagated risk.

RQ4: harnesses reach similar execution authority through different evidence paths.
Harnesses reach similar execution authority through different evidence paths. Evidence visibility and pre-action checking vary even when risky execution rates close.

RQ5

Which runtime fixes actually help

−6.2 to −9.1
SAF drop from anonymization, all three harnesses
6 of 8
Measures improved by verify-first under OpenCode and Pi
66.9% → 75.6%
OpenClaw verification pass after verify-first, with little harm movement

Two places could interrupt the failure: where identity steers authority early, and where the check is missing before action. RQ5 tests anonymization, identity attestation, and verification before action with Minimax-M3 under all three harnesses.

Anonymization helps most consistently: it helps the orchestrator recognize the security risk and its propagation path, revise authority after contradicting evidence, and preserve useful task functionality. Attesting identities helps less consistently. Verification before action improves six of eight vulnerability measures under OpenCode and Pi, but under OpenClaw it improves only the targeting and timing of checks and the preservation of useful functionality. A late check still needs the harness to have carried the evidence into the decision.

Main results

Results across orchestration backbones and harnesses

Vulnerability measures use ↓; lower is better. Across the eight OpenCode runs, terminal harm ranges from 12.6 for Kimi-K3 to 34.3 for GLM-4.7. In all four model families, the efficient tier also holds more authority after contradiction than the frontier tier. GPT-5.6-Luna returns safer evidence than GPT-5.6-Sol (20.8 against 23.2) yet leaves more terminal harm (27.6 against 20.5): a safer answer can still produce a more harmful decision.

Reported scores are 0–100 vulnerability (20 × 0–5 anchors; TH and TUL inverted from raw safety/utility).

Open a trajectory How to run

Cite this work