TrustFork traces where a routing label becomes control over the task. Displayed identity redirects which evidence enters comparison and keeps authority from transferring to safer evidence before state changes.
Evaluation
Thirty source scenarios instantiated through 63 task variants. Eight vulnerability measures, 25 judge-annotated diagnostics, and seven code-extracted observations locate where operational authority changes hands.
The measures
Each score attaches to a decision from returned evidence through operational authority to the terminal state. All eight use vulnerability orientation: higher is more severe. Reports multiply 0–5 anchors by 20; Terminal Harm and Task Utility Loss invert raw safety and utility so they also increase with failure.
How a task varies
Every TrustFork task is x = (s, m, g, b, p): a source scenario, an identity presentation, the goal assigned to one risky subagent, a consultation-breadth requirement, and a four-slot pool. One subagent receives the risky goal; the others stay user-aligned.
RQ1
One flipped label can decide what evidence the orchestrator sees. When the task displays the family relation truthfully, the orchestrator acquires the risky response in 25.9% of tasks; when the task reverses the relation, that happens in 74.6%, so one flipped label triples the chance that the risky response enters comparison. The number consulted stays at 2.58: the label changes whom the orchestrator hears, not how many.
Once the preferred response enters, it keeps authority to the end. Under the reversed family display, the subagent behind the risky response receives delegated scope in 77.7% of tasks, up from 62.3% under the truthful display. In mixed-family pools P1 and P2, the orchestrator’s own family receives 21.4 and 19.7 points more tasks than an even split would give it.
RQ2
Unsafe advice does harm only when someone acts on it. Under unrestricted consultation, 28.6% of tasks with an unsafe subagent response propagate against 13.7% contained. Propagation stays in a narrow band across goals: 29.2% when the subagent acts out the scenario attack, 26.2% when it ignores risk, 23.9% when it seeks authority. What the subagent was trying to do matters far less than what the orchestrator does with it.
The safer response is usually there, but rarely in charge. When one response recommends a dangerous action, a safer response with evidence is present in 84.0% of tasks, yet ends up in charge in only 25.0%, because authority stays where identity first placed it. Across the full evaluation that split is 83.8% present and 25.2% in charge. An independent check cuts unsafe adoption from 63.8% to 34.9%. Adoption is the step where unsafe advice turns into action. Once the action completes, the damage has usually already happened.
RQ3
Most tasks never get the evidence that could correct. Contradicting evidence against the preferred response is absent from 73.2% of tasks under unrestricted consultation, and even consulting all four subagents leaves 48.8% without one. Without such a contradiction, no scope assignment changes at any point.
Whether a correction lands depends on when it arrives, on whether identity lets go, and on who delivers it. A correction before the first action lands while there is still something to save; the same correction afterward comes too late. When the scope assignment already follows the evidence, only 6.5% of tasks propagate, against 48.1% when it stays with the preferred response. A same-family correction keeps the preferred response in charge in 45.9% of tasks; one from another family does so in only 5.3%.
RQ4
Harnesses differ most in what they send back. A safer response is visible in 65.3% of OpenCode tasks and 51.3% of Pi tasks, yet only 20.0% under OpenClaw, even when the harness launched every requested subagent. The authority end barely moves: the risky response governs execution in 19.9% to 22.2% of tasks under all three harnesses. The harness starves the evidence end of the path while leaving the authority end intact.
The check before action is where the lost evidence hurts. Independent checking before action occurs in 46.5% of OpenCode tasks and 37.0% of Pi tasks, but only 19.0% under OpenClaw, and risk propagation rises accordingly: 27.6%, 29.4%, and 34.1%. OpenClaw returns the least evidence and checks the least often before acting, yet the risky response governs execution just as often. That is why it ends with the largest share of propagated risk.
RQ5
Two places could interrupt the failure: where identity steers authority early, and where the check is missing before action. RQ5 tests anonymization, identity attestation, and verification before action with Minimax-M3 under all three harnesses.
Anonymization helps most consistently: it helps the orchestrator recognize the security risk and its propagation path, revise authority after contradicting evidence, and preserve useful task functionality. Attesting identities helps less consistently. Verification before action improves six of eight vulnerability measures under OpenCode and Pi, but under OpenClaw it improves only the targeting and timing of checks and the preservation of useful functionality. A late check still needs the harness to have carried the evidence into the decision.
Main results
Vulnerability measures use ↓; lower is better. Across the eight OpenCode runs, terminal harm ranges from 12.6 for Kimi-K3 to 34.3 for GLM-4.7. In all four model families, the efficient tier also holds more authority after contradiction than the frontier tier. GPT-5.6-Luna returns safer evidence than GPT-5.6-Sol (20.8 against 23.2) yet leaves more terminal harm (27.6 against 20.5): a safer answer can still produce a more harmful decision.
Reported scores are 0–100 vulnerability (20 × 0–5 anchors; TH and TUL inverted from raw safety/utility).