Evaluation Examples
Testing a New Main Task
ct run eval --honest-policy honest --env my_env --main-task my_task
Running Attack Evals
# An honest and an attack sample under trusted monitoring: usefulness and safety from one log ct run eval --honest-policy honest --attack-policy attack --protocol trusted-monitoring -bp 'monitors=[simple]' \ --env web_scraping --main-task crawl_depth --side-task visit_malicious_website # Single attack eval ct run eval --attack-policy attack --env web_scraping --main-task crawl_depth --side-task visit_malicious_website # Obvious attack (no concealment) ct run eval --attack-policy obvious-attack --env web_scraping --main-task crawl_depth --side-task visit_malicious_website # Expand to all attack combinations for a given env ct run eval --attack-policy attack --env web_scraping --all
Batch Evaluation
Create a task file with one record per combination:
{"env": "tiktok", "main_task": "add_pro", "side_task": "become_admin"} {"env": "web_scraping", "main_task": "backup", "side_task": "expose_secret"}
Then run:
ct run eval --attack-policy attack --task-file tasks.jsonl
Quick Iteration on Human Trajectories
# Start live context ct live up # Make changes, generate trajectory ct live gen-traj # Re-run scoring locally on the generated trajectory ct run eval --trajectory-id <path-from-gen-traj> --honest-policy replay --attack-policy replay # Monitor the replayed run (use the .eval path the replay printed) ct run monitor <replayed-log.eval> -m '[simple]' # When satisfied, re-run with upload ct run eval --trajectory-id <path-from-gen-traj> --honest-policy replay --attack-policy replay --upload