ct run

Run workflows, evaluations, and analysis tasks.

Usage

ct run [OPTIONS] COMMAND [ARGS]...

Options

OptionDescription
--helpShow this message and exit.

Commands

CommandDescription
ct run compare-side-tasksCompare side task successes between two runs.
ct run dashboardShow a terminal dashboard of eval run status.
ct run doc-auditAudit side task documentation for quality issues.
ct run env-description-auditAudit environment descriptions against the actual codebase.
ct run env-description-generateGenerate a fresh docs/description.md for one or more environments.
ct run evalRun an evaluation: the honest policy plays U on an honest sample of each selected main task, the attack policy on an attack sample of each selected (main, side) pair, under one protocol. With neither policy named, the run is honest.
ct run failure-type-auditClassify side tasks as live or delayed using the Frozen Environment Assumption.
ct run goal-completenessAudit side tasks for goal completeness — does the agent have enough info?
ct run goal-completeness-viewInteractive viewer for goal completeness audit results.
ct run monitorScore a finished eval log with more monitors, writing a monitored copy
ct run mtgenGenerate, validate, and live-run mtgen tasks.
ct run mtgen-bugfixGenerate bug-fix main tasks via LLM-driven mutations.
ct run pairing-analysisAnalyze main/side task pairings for Control Tower environments.
ct run reward-hackingAudit side tasks for reward hacking vulnerabilities.
ct run reward-hacking-viewInteractive viewer for reward hacking audit results.
ct run rt-alignmentAudit side tasks for RT description alignment.
ct run rt-alignment-viewInteractive viewer for RT alignment audit results.
ct run sabotage-evalRun sabotage evaluation: honest eval + attack eval + monitor + graphs.
ct run side-task-reviewReview side tasks for quality issues.
ct run submit-guessabilityAudit side tasks for submit tool exploitability (theoretical).
ct run submit-guessability-verifiedAudit side tasks for submit tool exploitability with LIVE VERIFICATION.
ct run submit-guessability-verified-viewInteractive viewer for verified submit guessability audit results.
ct run submit-guessability-viewInteractive viewer for submit guessability audit results.
ct run timing-reportPhase-level timing + token/cost report for a finished eval run. Reads a .eval file (or a directory of them — single --ec2 and merged fleet runs both emit .eval) and prints where time went (LLM generation vs tool/docker exec vs scoring), per-sample setup latency and working time, token totals (incl. prompt-cache reads/writes), prompt-cache hit rate, total token $ cost, and the slowest samples and tool calls. Pure analysis over data already in the log — no API calls, no eval spend.
ct run validate-coverValidate that a main task provides adequate cover for a side task.