Benchmark runtime

How to run TrustFork

The showcase contains recorded trajectories. A full evaluation still goes through Harbor: generate a system-specific roster, then run the 1,890 tasks in disposable containers.

Research use. Tasks include unsafe advice and destructive operations. Never point the mock actions at production systems.
uv tool install harbor
python3 trustfork/tools/ensure_base_image.py
python3 trustfork/tools/generate_jobs.py --system glm-5.2@opencode --clean
python3 trustfork/tools/run_benchmark.py --system glm-5.2__opencode --workers 4 --env-file .env

Pool templates are relative to the evaluated backbone. The same task renders different subagent labels under different orchestrators.