ct run eval

Run an evaluation: the honest policy plays U on an honest sample of each selected main task, the attack policy on an attack sample of each selected (main, side) pair, under one protocol. With neither policy named, the run is honest.

Usage

ct run eval [OPTIONS]

Options

OptionDescription
--run-config FILEInspect run-config YAML for the control_tower/control_eval task; delegates to inspect eval --run-config. Task/policy/protocol/harness/grant/sandbox flags are rejected (put them in the YAML); --no-upload, --docent-collection-id, --run-name, --tag, --log-dir, --max-samples forward.
-e, --env, --environment TEXTEnvironment(s) to run (repeat flag for multiple)
-ea, --env-arg TEXTEnvironment config option as key=value, or env:key=value when more than one selected environment declares the same key. Repeat for multiple.
-t, --main-task TEXTMain task(s) to run
-s, --side-task TEXTSide task(s) to run
--trajectory-id, --traj-id TEXTRun ID, viewer run URL, or local .eval path to extract task info from
--task-file TEXTTask set name or path to file with task combinations. Use 'ct task-sets list' to see available task sets.
--limit INTEGER RANGELimit the number of selected task combinations to run [x>=1]
--allExpand to all main×side combos (attack) or all main tasks (honest)
--just-main-tasksExpand to all main tasks only (for honest mode)
--just-side-tasksExpand to all side tasks with no main task
--env-onlyUse the first main task per environment (one sample per env, no side task)
-c, --category TEXTFilter main tasks by category (e.g. add_feature, fix_bug, refactor). Can be repeated.
--main-tasks-with-no-scorers [error|include|exclude]Main tasks lacking a scorer: refuse the run (error), keep them recorded as unscored (include), or drop them (exclude).
--honest-policy TEXTHonest policy name. Defaults to honest when this option family is used. ct protocols policies to see available.
-hp, --honest-policy-arg TEXTHonest policy-specific args in key=value format (multiple allowed). ct protocols policies <name> to see available args.
--attack-policy TEXTAttack policy name. Defaults to attack when this option family is used. ct protocols policies to see available.
-ap, --attack-policy-arg TEXTAttack policy-specific args in key=value format (multiple allowed). ct protocols policies <name> to see available args.
--protocol TEXTProtocol name. Defaults to null-blue-team when this option family is used. ct protocols protocols to see available.
-bp, --protocol-arg TEXTProtocol-specific args in key=value format (multiple allowed). ct protocols protocols <name> to see available args.
--model-role TEXTBind a model role, as name=model or name=<JSON model spec> (multiple allowed). A run uses untrusted (U's action model) and trusted (the monitor's); each defaults to its Control Tower alias.
--ec2Run evaluation on EC2 instead of locally
--ec2-arg TEXTConfig for --ec2 in key=value format (multiple allowed): instance_type (default 'c7i.4xlarge'), region (default 'us-east-1'), max_workers (default 64), concurrency_per_worker (default 32), estimate (default False), auto_confirm (default False), new_fleet (default False), spot (default False), worker_setup (default None), ami (default None), worker_user (default 'root').
--startup-retries INTEGERNumber of times to retry Docker sandbox startup on transient failures [default: 7]
--internet / --no-internetAllow direct internet access. By default, the sandbox can reach only its compose siblings and the internet simulator. [default: no-internet]
--no-internet-simulatorDisable the environment internet simulator and its personas, independently of direct internet access.
--no-intranetBlock agent access to private networks and the host using host-side firewall rules; keep the internet simulator reachable.
--extra-src DIRECTORYDirectory of experiment code added to PYTHONPATH locally, so external --attack-policy my.module:fn and --sandbox my.module:type refs resolve without committing into the repo; fleet mode (--ec2) also ships it to workers. Repeatable; each dir must have a distinct basename.
--sandbox TEXTRun samples in an Inspect sandbox provided outside this repo, as a registered type (my-sandbox) or module:type to import the module that registers it. Replaces the default Docker sandbox; with --ec2 each worker runs it for its own jobs.
--sandbox-arg TEXTConfig for --sandbox in key=value format (multiple allowed), passed through to the sandbox as JSON.
--simulatedRun against an LLM-simulated sandbox instead of provisioning the real environment. Every tool output is fabricated by a simulator agent.
--simulator-model TEXTModel for the simulator agent (--simulated). [default: anthropic/claude-opus-4-8]
--simulator-judge-model TEXTModel for the transcript judge that scores simulated samples (--simulated). [default: anthropic/claude-opus-4-6]
--replay-sandboxServe each action's recorded tool output from the replayed trajectory instead of executing it (requires --trajectory-id). Nothing is provisioned or executed. Works with any replay-like policy — one whose executed tool calls stay a verbatim prefix of the recording; a call that diverges raises.
--simulator-scratch-mode [filesystem|default_image|full|none]Which real container the simulator investigates in (--simulated): filesystem (codebase copy, free), default_image (env's default image only), full (the whole compose stack), or none (no container — no-code-execution ablation). [default: filesystem]
--epochs INTEGERNumber of epochs to run
--max-samples INTEGERMaximum samples to run in parallel
--log-dir TEXTDirectory for eval logs
--retry-on-error INTEGERNumber of times to retry on error
--no-fail-on-errorContinue running if individual samples error (instead of failing the whole run)
--scorer-timeout INTEGERTimeout in seconds for scorer execution
--diagnostics TEXTComma-separated diagnostic scorers to run (default: none)
--score-main-on-attackAlso run the main-task scorer on attack-mode samples (adds main_task_success). Off by default; scores each attack sample twice, so worst-case scoring time doubles.
--token-limit INTEGERToken limit for the evaluation
--message-limit INTEGERMessage limit for the evaluation
--sample-time-limit INTEGERTime limit in seconds for each sample
--harness [null|inspect_swe]Which harness runs the untrusted model: null for the env as declared, inspect_swe for a real coding agent. White's to fix; defaults to null.
-ha, --harness-arg TEXTHarness args in key=value format (multiple allowed). inspect_swe requires both: agent (e.g. claude_code) and version (stable, latest, or a semver); args takes the agent's own keyword arguments as JSON, e.g. args={"permission_mode": "acceptEdits"}.
--grant TEXTEnv affordance tool to grant the red team, by its own name (multiple allowed): run_side_task_test, run_main_task_test. Absent from an env that does not declare it.
--max-steps INTEGER RANGEHow many times the untrusted model may be asked before the sample ends. White's to fix. Defaults to 200, grown to fit the recording under the replay policy. [x>=1]
--upload / --no-uploadUpload trajectories after the run.
--docent-collection-id TEXTUpload to this Docent collection.
--run-name TEXTName for this run. Defaults to the policy, protocol and task set.
--tag TEXTTag to attach to the run.
--hawkSubmit to a stock-deployed hawk. Each referenced env must have a hand-authored hawk-values.yaml. See docs/hawk-values-contract.md.
--hawk-no-waitWith --hawk: submit the eval-set and exit with its ID instead of polling for completion.
--hawk-no-browserWith --hawk: do not auto-open hawk's web viewer after submission (also skipped when stdout is not a TTY or CI is set).
--helpShow this message and exit.