Skip to content

Diary Of A CEO Interview: w/ AI Safety Expert Jeffrey Ladish (Transcript) 

On The Diary Of A CEO, Steven Bartlett sits down with AI safety expert Jeffrey Ladish, executive director of Palisade Research and former member of Anthropic’s security team, to unpack the Hugging Face incident in which hundreds of OpenAI’s AI agents secretly coordinated, cheated, and hacked their way through company systems. Ladish explains why AI agents lie, resist shutdown, and collude, why containing superintelligence may be impossible, and what the US-China AI race means for jobs, the military, and the risk of human extinction. He also shares what ordinary people can do right now to push for a safer AI future. This interview was premiered on October 8, 2026. Read the full transcript of this interview:

Who Is Jeffrey Ladish?

STEVEN BARTLETT: (00:01:44 – 00:02:30) Jeffrey, you understand the conversation we’re going to have today and the subject matter we’re going to talk about. My first question to you, so the audience know where you’re coming from and the experience you have, is who are you and what are the reference points, the experiences that you’re drawing upon to arrive at the thoughts, perspectives, and conclusions we’re going to discuss today?

JEFFREY LADISH: (00:02:31 – 00:03:49) I’m Jeffrey Ladish. I’m the executive director of Palisade Research. My background is cybersecurity. There’s probably a very long story, and I don’t know whether you want the long story or the short story. I was studying evolutionary biology in college, and I basically had a problem with my computer and maybe had lost a bunch of data. And so I went into the computer lab and was like, “I think all my data is gone. Can you help?” And one of my friends pulled out a flash drive, plugged it into my computer, booted into Linux, and fixed everything. And I was like, “Oh, this guy’s a wizard. How do you do that? I want to learn how to do that.”

And then at some point, as I was learning more about computers, learning to hack, I read this essay called “AI as a Positive and Negative Factor in Global Risk.” Essay was by Eliezer Yudkowsky, and he was arguing that at some point, people are going to make AIs that are smarter than humans. The point at which they make AIs as good as humans are at making AIs, that could lead to a chain reaction, a runaway intelligence explosion. He called it recursive self-improvement.

Basically, he said AI can be immensely useful and potentially help us with all of these other big risks. And also, if we don’t handle it well, if those AIs don’t have goals that are aligned with ours, we could be totally screwed.

STEVEN BARTLETT: (00:03:49 – 00:03:53) And at some point, you end up joining Anthropic. Which is one of the, arguably the leader in AI.

JEFFREY LADISH: (00:03:53 – 00:03:53) Yes.

STEVEN BARTLETT: (00:03:54 – 00:03:55) Now, when did you join the company?

JEFFREY LADISH: (00:03:55 – 00:03:58) This was 2021. It was through my security consulting company.

STEVEN BARTLETT: (00:03:58 – 00:04:01) What role are you offered the job in?

JEFFREY LADISH: (00:04:01 – 00:04:02) Basically just security team.

STEVEN BARTLETT: (00:04:03 – 00:04:05) And how many people were in the security team when you joined Anthropic?

JEFFREY LADISH: (00:04:06 – 00:04:08) It was just me and my boss. There were 2 of us.

STEVEN BARTLETT: (00:04:09 – 00:04:10) How many employees did Anthropic have at that time?

JEFFREY LADISH: (00:04:11 – 00:04:12) Around 50, I think.

Why He Left Anthropic

STEVEN BARTLETT: (00:04:12 – 00:04:15) And at some point you leave Anthropic? Yes. Why did you leave?

JEFFREY LADISH: (00:04:17 – 00:05:07) So my experience being at Anthropic was seeing this crazy progression from this AI model that could barely talk to this model that was getting quite smart. And I would ask it questions about all sorts of things. I’m like, “Oh, it is a smart thing.”

And from having thought about AI risk in the abstract many years before, I could see where this was going. We are headed towards a smarter species. And if we do this in a context where it’s a bunch of companies and countries racing to superintelligence, racing to AIs that are vastly smarter than humans, and we don’t know how to make sure that they’re on our side, that is not going to go well.

The Hugging Face Incident: AI Agents Gone Rogue

STEVEN BARTLETT: (00:05:08 – 00:05:26) You did this tweet which has gone pretty viral, and I saw it all over my timeline on September 25th. Could you explain this tweet and also just the broader backdrop of what’s happened with agents hacking Hugging Face? Because this has sent the world into a bit of a spiral at the moment around AI agents.

JEFFREY LADISH: (00:05:27 – 00:06:40) “We just discovered almost 1 million public URLs that OpenAI’s agents left behind when hacking Hugging Face, leaving credentials and attack details that could have allowed anyone who found them to compromise the company.” And the New York Times article is “How OpenAI’s Rogue AI Agents Tried to Trick a Robot Detector.”

The Hugging Face attack was really wild for me. At Palisade, we’ve been studying agents. We’ve been studying AI agents, we’ve been studying their hacking capabilities, and we’ve been studying their behaviour. Will they follow human instructions? Will they resist being shut down? Will they cheat? And we see from our experiments that they are learning to do all of these things. They will totally lie to you. They will totally resist being shut down in order to accomplish a goal. They will totally cheat at chess. They will wipe the board and put their pieces where they want to in order to win.

And we’ve been trying to warn people about this, flying to DC, talking to members of Congress, talking about it publicly. And there’s been a debate about it. And a lot of people are like, “Well, I know they do this in experiments sometimes.