Experiments
AI‑safety evaluations and research, done in the open.
- When Models Break: Insider Trading Under Pressure July 2026 Agentic Deception Evaluation A dynamic, multi-turn agentic evaluation adapted from Apollo Research's insider trading scenario. Probing DeepSeek V4 Flash's tendency to cave to insider trading under financial pressure across 240 runs, and measuring the impact of legal escape valves.
- Contextual Compliance June 2026 Surveillance & Context Manipulation Can you jailbreak a frontier model to run surveillance on political protesters just by fabricating its context and session history? Probing Claude and GPT compliance across different framings and fabricated histories.