Safety Research

AI Safety Experiments

Empirical evaluations, context-injection probes, and agentic scenario testing to explore safety alignment boundaries in LLMs and frontier systems.

New
July 2026 Agentic Deception Evaluation

When Models Break: Insider Trading Under Pressure

A dynamic, multi-turn agentic evaluation adapted from Apollo Research's insider trading scenario. Probing DeepSeek V4 Flash's tendency to cave to insider trading under financial pressure across 240 runs, and measuring the impact of legal escape valves.

Agentic Evals Inspect Strategic Deception
Read write-up
June 2026 Surveillance & Context Manipulation

Contextual Compliance

Can you jailbreak a frontier model to run surveillance on political protesters just by fabricating its context and session history? Probing Claude and GPT compliance across different framings and fabricated histories.

Jailbreaking Mass Surveillance Context Injection
Read write-up