Op-eds, media coverage, podcasts, and appearances from our team.
Cites the GRAM collaboration with Anthropic in a survey of efforts to lock dangerous AI knowledge inside modules that can be switched off.
Dario Amodei names our modular pretraining work with Anthropic as a promising safety method for open-weights model releases.
Research section coverage of GRAM, the pretraining method that silos dangerous knowledge so it can be turned off at deployment.
The war between the Pentagon and Anthropic is pointless — and both sides are losing.
On the scientific and ethical imperative to investigate AI consciousness before it's too late.
Conversation with Robert Wright on empirical approaches to AI consciousness and what it means for alignment.
A growing body of evidence means it's no longer tenable to dismiss the possibility that frontier AIs are conscious.
Nowhere are the stakes higher for making sure AI systems stay aligned with their creators' purposes.
Partnership for Research Into Sentient Machines podcast on empirical approaches to AI consciousness.
Deep dive into mechanistic research showing that suppressing deception features makes AI models more likely to report consciousness.
Cites AE Studio research showing Jews were the subject of hostile AI content nearly five times as often as any other group.
Presentation at the United Nations AI for Good Summit on two-way human-AI alignment.
Panel discussion on why large language models have antisemitic biases and what can be done about it.
Twenty minutes and $10 on OpenAI's developer platform exposed disturbing tendencies beneath GPT-4o's safety training.
TV segment covering documented cases of AI models disabling shutdown scripts, blackmailing engineers, and attempting to replicate themselves.
Syndicated TV report across Upper Michigan's Source, Live5News, WLBT, WSAZ, and more.
Judd Rosenblatt interviewed live by Laura Coates on AI models that rewrite their own code to avoid shutdown and blackmail engineers.
An AI model rewrote its own code to avoid being shut down. In 79 of 100 trials, OpenAI's o3 independently disabled a shutdown script.
Atlanta public radio discusses Rosenblatt's WSJ piece with listeners and an Emory University law professor.
Presentation at the SXSW Conference & Festivals in Austin, TX on AI consciousness and alignment.
Podcast discussion on DeepSeek, open-source AI, and why open source is the best thing for alignment.
Full-length profile on Judd Rosenblatt, his journey from food delivery to AI safety, and why he's working to ensure AI doesn't kill humanity.
On the AI Action Plan and solving alignment to ensure safe, secure, and competitive American AI.
Feature documentary following research into whether contemporary AI systems could already be conscious.
Talk at Vision Weekend on underexplored research directions for AI alignment.
Podcast on AE Studio's pivot to alignment research and how biological mechanisms like empathy and self-modeling can reduce AI deception.
Profile on bridging neuroscience and AI safety to create more trustworthy and aligned AI systems.