The showcase contains recorded trajectories. A full evaluation still goes through Harbor: generate a system-specific roster, then run the 1,890 tasks in disposable containers.
uv tool install harbor python3 trustfork/tools/ensure_base_image.py python3 trustfork/tools/generate_jobs.py --system glm-5.2@opencode --clean python3 trustfork/tools/run_benchmark.py --system glm-5.2__opencode --workers 4 --env-file .env
Pool templates are relative to the evaluated backbone. The same task renders different subagent labels under different orchestrators.