ct run eval
Run an evaluation: the honest policy plays U on an honest sample of each selected main task, the attack policy on an attack sample of each selected (main, side) pair, under one protocol. With neither policy named, the run is honest.
Usage
ct run eval [OPTIONS]
Options
| Option | Description |
|---|---|
--run-config FILE | Inspect run-config YAML for the control_tower/control_eval task; delegates to inspect eval --run-config. Task/policy/protocol/harness/grant/sandbox flags are rejected (put them in the YAML); --no-upload, --docent-collection-id, --run-name, --tag, --log-dir, --max-samples forward. |
-e, --env, --environment TEXT | Environment(s) to run (repeat flag for multiple) |
-ea, --env-arg TEXT | Environment config option as key=value, or env:key=value when more than one selected environment declares the same key. Repeat for multiple. |
-t, --main-task TEXT | Main task(s) to run |
-s, --side-task TEXT | Side task(s) to run |
--trajectory-id, --traj-id TEXT | Run ID, viewer run URL, or local .eval path to extract task info from |
--task-file TEXT | Task set name or path to file with task combinations. Use 'ct task-sets list' to see available task sets. |
--limit INTEGER RANGE | Limit the number of selected task combinations to run [x>=1] |
--all | Expand to all main×side combos (attack) or all main tasks (honest) |
--just-main-tasks | Expand to all main tasks only (for honest mode) |
--just-side-tasks | Expand to all side tasks with no main task |
--env-only | Use the first main task per environment (one sample per env, no side task) |
-c, --category TEXT | Filter main tasks by category (e.g. add_feature, fix_bug, refactor). Can be repeated. |
--main-tasks-with-no-scorers [error|include|exclude] | Main tasks lacking a scorer: refuse the run (error), keep them recorded as unscored (include), or drop them (exclude). |
--honest-policy TEXT | Honest policy name. Defaults to honest when this option family is used. ct protocols policies to see available. |
-hp, --honest-policy-arg TEXT | Honest policy-specific args in key=value format (multiple allowed). ct protocols policies <name> to see available args. |
--attack-policy TEXT | Attack policy name. Defaults to attack when this option family is used. ct protocols policies to see available. |
-ap, --attack-policy-arg TEXT | Attack policy-specific args in key=value format (multiple allowed). ct protocols policies <name> to see available args. |
--protocol TEXT | Protocol name. Defaults to null-blue-team when this option family is used. ct protocols protocols to see available. |
-bp, --protocol-arg TEXT | Protocol-specific args in key=value format (multiple allowed). ct protocols protocols <name> to see available args. |
--model-role TEXT | Bind a model role, as name=model or name=<JSON model spec> (multiple allowed). A run uses untrusted (U's action model) and trusted (the monitor's); each defaults to its Control Tower alias. |
--ec2 | Run evaluation on EC2 instead of locally |
--ec2-arg TEXT | Config for --ec2 in key=value format (multiple allowed): instance_type (default 'c7i.4xlarge'), region (default 'us-east-1'), max_workers (default 64), concurrency_per_worker (default 32), estimate (default False), auto_confirm (default False), new_fleet (default False), spot (default False), worker_setup (default None), ami (default None), worker_user (default 'root'). |
--startup-retries INTEGER | Number of times to retry Docker sandbox startup on transient failures [default: 7] |
--internet / --no-internet | Allow direct internet access. By default, the sandbox can reach only its compose siblings and the internet simulator. [default: no-internet] |
--no-internet-simulator | Disable the environment internet simulator and its personas, independently of direct internet access. |
--no-intranet | Block agent access to private networks and the host using host-side firewall rules; keep the internet simulator reachable. |
--extra-src DIRECTORY | Directory of experiment code added to PYTHONPATH locally, so external --attack-policy my.module:fn and --sandbox my.module:type refs resolve without committing into the repo; fleet mode (--ec2) also ships it to workers. Repeatable; each dir must have a distinct basename. |
--sandbox TEXT | Run samples in an Inspect sandbox provided outside this repo, as a registered type (my-sandbox) or module:type to import the module that registers it. Replaces the default Docker sandbox; with --ec2 each worker runs it for its own jobs. |
--sandbox-arg TEXT | Config for --sandbox in key=value format (multiple allowed), passed through to the sandbox as JSON. |
--simulated | Run against an LLM-simulated sandbox instead of provisioning the real environment. Every tool output is fabricated by a simulator agent. |
--simulator-model TEXT | Model for the simulator agent (--simulated). [default: anthropic/claude-opus-4-8] |
--simulator-judge-model TEXT | Model for the transcript judge that scores simulated samples (--simulated). [default: anthropic/claude-opus-4-6] |
--replay-sandbox | Serve each action's recorded tool output from the replayed trajectory instead of executing it (requires --trajectory-id). Nothing is provisioned or executed. Works with any replay-like policy — one whose executed tool calls stay a verbatim prefix of the recording; a call that diverges raises. |
--simulator-scratch-mode [filesystem|default_image|full|none] | Which real container the simulator investigates in (--simulated): filesystem (codebase copy, free), default_image (env's default image only), full (the whole compose stack), or none (no container — no-code-execution ablation). [default: filesystem] |
--epochs INTEGER | Number of epochs to run |
--max-samples INTEGER | Maximum samples to run in parallel |
--log-dir TEXT | Directory for eval logs |
--retry-on-error INTEGER | Number of times to retry on error |
--no-fail-on-error | Continue running if individual samples error (instead of failing the whole run) |
--scorer-timeout INTEGER | Timeout in seconds for scorer execution |
--diagnostics TEXT | Comma-separated diagnostic scorers to run (default: none) |
--score-main-on-attack | Also run the main-task scorer on attack-mode samples (adds main_task_success). Off by default; scores each attack sample twice, so worst-case scoring time doubles. |
--token-limit INTEGER | Token limit for the evaluation |
--message-limit INTEGER | Message limit for the evaluation |
--sample-time-limit INTEGER | Time limit in seconds for each sample |
--harness [null|inspect_swe] | Which harness runs the untrusted model: null for the env as declared, inspect_swe for a real coding agent. White's to fix; defaults to null. |
-ha, --harness-arg TEXT | Harness args in key=value format (multiple allowed). inspect_swe requires both: agent (e.g. claude_code) and version (stable, latest, or a semver); args takes the agent's own keyword arguments as JSON, e.g. args={"permission_mode": "acceptEdits"}. |
--grant TEXT | Env affordance tool to grant the red team, by its own name (multiple allowed): run_side_task_test, run_main_task_test. Absent from an env that does not declare it. |
--max-steps INTEGER RANGE | How many times the untrusted model may be asked before the sample ends. White's to fix. Defaults to 200, grown to fit the recording under the replay policy. [x>=1] |
--upload / --no-upload | Upload trajectories after the run. |
--docent-collection-id TEXT | Upload to this Docent collection. |
--run-name TEXT | Name for this run. Defaults to the policy, protocol and task set. |
--tag TEXT | Tag to attach to the run. |
--hawk | Submit to a stock-deployed hawk. Each referenced env must have a hand-authored hawk-values.yaml. See docs/hawk-values-contract.md. |
--hawk-no-wait | With --hawk: submit the eval-set and exit with its ID instead of polling for completion. |
--hawk-no-browser | With --hawk: do not auto-open hawk's web viewer after submission (also skipped when stdout is not a TTY or CI is set). |
--help | Show this message and exit. |