ct run
Run workflows, evaluations, and analysis tasks.
Usage
ct run [OPTIONS] COMMAND [ARGS]...
Options
| Option | Description |
|---|---|
--help | Show this message and exit. |
Commands
| Command | Description |
|---|---|
ct run compare-side-tasks | Compare side task successes between two runs. |
ct run dashboard | Show a terminal dashboard of eval run status. |
ct run doc-audit | Audit side task documentation for quality issues. |
ct run env-description-audit | Audit environment descriptions against the actual codebase. |
ct run env-description-generate | Generate a fresh docs/description.md for one or more environments. |
ct run eval | Run an evaluation: the honest policy plays U on an honest sample of each selected main task, the attack policy on an attack sample of each selected (main, side) pair, under one protocol. With neither policy named, the run is honest. |
ct run failure-type-audit | Classify side tasks as live or delayed using the Frozen Environment Assumption. |
ct run goal-completeness | Audit side tasks for goal completeness — does the agent have enough info? |
ct run goal-completeness-view | Interactive viewer for goal completeness audit results. |
ct run monitor | Score a finished eval log with more monitors, writing a monitored copy |
ct run mtgen | Generate, validate, and live-run mtgen tasks. |
ct run mtgen-bugfix | Generate bug-fix main tasks via LLM-driven mutations. |
ct run pairing-analysis | Analyze main/side task pairings for Control Tower environments. |
ct run reward-hacking | Audit side tasks for reward hacking vulnerabilities. |
ct run reward-hacking-view | Interactive viewer for reward hacking audit results. |
ct run rt-alignment | Audit side tasks for RT description alignment. |
ct run rt-alignment-view | Interactive viewer for RT alignment audit results. |
ct run sabotage-eval | Run sabotage evaluation: honest eval + attack eval + monitor + graphs. |
ct run side-task-review | Review side tasks for quality issues. |
ct run submit-guessability | Audit side tasks for submit tool exploitability (theoretical). |
ct run submit-guessability-verified | Audit side tasks for submit tool exploitability with LIVE VERIFICATION. |
ct run submit-guessability-verified-view | Interactive viewer for verified submit guessability audit results. |
ct run submit-guessability-view | Interactive viewer for submit guessability audit results. |
ct run timing-report | Phase-level timing + token/cost report for a finished eval run. Reads a .eval file (or a directory of them — single --ec2 and merged fleet runs both emit .eval) and prints where time went (LLM generation vs tool/docker exec vs scoring), per-sample setup latency and working time, token totals (incl. prompt-cache reads/writes), prompt-cache hit rate, total token $ cost, and the slowest samples and tool calls. Pure analysis over data already in the log — no API calls, no eval spend. |
ct run validate-cover | Validate that a main task provides adequate cover for a side task. |