Dr. Stephen Casper - AI Safety
BIO
Dr. Stephen Casper (commonly known as "Cas") is an Assistant Professor of Public Policy at the Harvard Kennedy School and an affiliate of the Harvard John A. Paulson School of Engineering and Applied Sciences, having completed his PhD in Computer Science at MIT within the Algorithmic Alignment Group. His research centers on technical AI safety, safeguards, adversarial red-teaming, and institutional AI governance, framing AI risk mitigation as a continuous sociotechnical and regulatory challenge rather than a purely solvable machine learning alignment problem. His work spans red-teaming frontier models, auditing agentic AI systems, evaluating systemic misuse and proliferation risks, and designing process-oriented governance frameworks such as the AI Agent Index and the AI Risk Repository. In addition to academic research, he actively contributes to global policy initiatives, serving as a writer for the International AI Safety Report, a lead author for the Singapore Consensus, and a former research resident with the UK AI Security Institute.
Synopsis
As AI transitions from novel chatbots to autonomous agents running finance, healthcare, and critical infrastructure, safety isn't just an academic debate anymore—it's a high-stakes issue that impacts every single one of us.
Together, Benny and Stephen cut through the sci-fi hype to unpack the urgent realities of AI safety and governance:
Safety vs. Alignment: Why getting an AI to follow orders isn't enough, and why true safety requires solving complex political, legal, and human challenges.
Real-World Threats: How everyday risks—from automated negligence and unexpected AI glitches to malicious misuse—directly impact society.
The Open-Source Dilemma: The high-wire act of keeping powerful AI tools open and accessible without handing dangerous capabilities to bad actors.
Battle testing Frontier Models: Real takeaways from analyzing systems like Stable Diffusion, DeepSeek, and emerging autonomous agents operating in the wild.
“AI safety isn’t just about alighnment. There is a lot more to AI safety than just making sure the systems do what the creators want them to do”
[47 min]
Next Podcast: Intentional Integrity by Rob Chesnut
Highlights from our conversation
What is AI Safety? [4:59 min]
“AI safety isn’t just about alighnment. There is a lot more to AI safety than just making sure the systems do what the creators want them to to…. I think of AI safety, not really as a technical solviable or machine learning problem, not as an alignment problem, but as kind of a never ending institional challenge.”
Defining AI Risks [4 min]
“Technical people love well defined problems. But in the real world, I don’t think it is going to lend itself to such nicely formulated problems”
The Alignment Problem [4 min]
“Real world problem is different from technical problem. The people who are worried about alignment problem because they understand the theory behind it. But this really doesn’t seem to be shaping up to the kind of the defining challenge for making sure that AI has a positive impact on the world.”
AI Evaluations [3 min]
DeepSeek & AI Proliferation [6 min]
“It’s a lesson about what happens when you know developers without requirements for accountability or requirements to produce documentations or constraints around what they do, and they want to make a splash and push something out to the world. “
Substantive vs Process Regulation [6 min]
“One of the most tenable paths forward is going to be one that focuses on process regulation and transparency.”
Approaches to AI Governance [3 min]
“We want AI systems to be aligned with user preferences . We also want people who are in control of the AI systems to be aligned with society.”
Open Source vs Closed Systems [5 min]
“How reasonably well-intentioned model development processes can have pretty devasting consequences doesnstream.”
Urgency of AI Safety [4 min]
“We need to have safety culture early on.”
AI Agents and Governance [5 min]
“The agentic space is rapidly developing and governane is even a bigger problem than generative AI.”
Published Work
AI Alignment & Reinforcement Learning from Human Feedback (RLHF)"Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback" (TMLR): A seminal critique and taxonomy detailing the technical, theoretical, and empirical limitations of RLHF as a primary alignment mechanism for LLMs."Foundational Challenges in Assuring Alignment and Safety of Large Language Models" (TMLR): A comprehensive survey outlining the core barriers to guaranteeing LLM safety and reliability."Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs" (FAccT): Evaluates the brittleness and variance in measuring cultural values and alignment in frontier language models.
Red-Teaming, Jailbreaks & Adversarial Robustness"Explore, Establish, Exploit: Red Teaming Language Models from Scratch": Introduces the "Explore, Establish, Exploit" (EEE) framework to systematically elicit novel failure modes and hallucinations (introducing the CommonClaim benchmark)."Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation": Demonstrates how persona-prompting attacks can reliably bypass safety guardrails across different commercial models."Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs" / "Defending Against Unforeseen Failure Modes with Latent Adversarial Training" (TMLR): Investigates adversarial training in internal representations to improve resilience against hidden or persistent failure modes.
Model Auditing, Unlearning & Tamper-Resistance"Black-Box Access is Insufficient for Rigorous AI Audits" (FAccT): Argues the necessity of white-box and gray-box access for meaningful third-party auditing and regulatory oversight."Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs" (ICLR): Demonstrates that curating/filtering pretraining data creates persistent safeguards that resist post-hoc adversarial fine-tuning and unlearning reversals."Rethinking Machine Unlearning for Large Language Models" (Nature Machine Intelligence) & "Eight Methods to Evaluate Robust Unlearning in LLMs": Landmark studies evaluating whether models genuinely forget sensitive or hazardous knowledge.
Interpretability & Mechanistic Analysis"Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks" (IEEE SaTML): A comprehensive taxonomy and critical evaluation of methods used to interpret deep neural network representations."Open Problems in Mechanistic Interpretability" (TMLR): A roadmap detailing open challenges in reverse-engineering neural networks into human-understandable circuits.
AI Governance Frameworks & Global Policy"The AI Risk Repository: A Meta-Review, Database, and Taxonomy of Risks from Artificial Intelligence" (Patterns): A systematic, multi-dimensional database categorizing hundreds of documented AI risks.The AI Agent Index: An empirical audit of top agentic AI systems assessing the widespread lack of safety, security, and autonomy documentation."The Singapore Consensus on Global AI Safety Research Priorities": Collaborative framework established among international researchers setting core priorities for global safety standards.