Skip to content
Home » Elon Musk: ‘Smartest AI In The World’ At The Grok 4 Launch Event

Elon Musk: ‘Smartest AI In The World’ At The Grok 4 Launch Event

Read the full transcript of Elon’s xAI Grok 4 launch event on Thursday, July 10, 2025.

Welcome to the Grok 4 Release

ELON MUSK: All right, welcome to the Grok 4 release here. This is the smartest AI in the world, and we’re going to show you exactly how and why. It really is remarkable to see the advancement of artificial intelligence, how quickly it is evolving. I sometimes think, compare it to the growth of a human and how fast a human learns and gains conscious awareness and understanding, and AI is advancing just vastly faster than any human.

We’re going to take you through a bunch of benchmarks that Grok 4 is able to achieve incredible numbers on, but it’s actually worth noting that Grok 4, if given the SAT, would get perfect SATs every time, even if it’s never seen the questions before. Even going beyond that to say graduate student exams like the GRE, it will get near-perfect results in every discipline of education. From the humanities to languages, math, physics, engineering, pick anything, and we’re talking about questions that it’s never seen before. These are not on the internet, and Grok 4 is smarter than almost all graduate students in all disciplines simultaneously. It’s actually just important to appreciate that’s really something.

Superhuman Reasoning Capabilities

The reasoning capabilities of Grok are incredible. There’s some people out there who think AI can’t reason, and it can reason at superhuman level. Yeah, and frankly, it only gets better from here. We’ll take you through the Grok 4 release and show you the pace of progress here.

I guess the first part is, in terms of the training, going from Grok 2 to Grok 3 to Grok 4, we’ve essentially increased the training by an order of magnitude in each case. So it’s 100 times more training than Grok 2, and that’s only going to increase. So it’s frankly, I don’t know, in some ways a little terrifying, but the growth of intelligence here is remarkable.

It’s important to realize there are two types of training compute. One is the pre-training compute, that’s from Grok 2 to Grok 3. But from Grok 3 to Grok 4, we’re actually putting a lot of compute in reasoning in RL.

Just like you said, this is literally the fastest-moving field, and Grok 2 is like the high school student by today’s standards. If you look back in the last 12 months, Grok 2 was only a concept. We didn’t even have Grok 2 12 months ago. And then by training Grok 2, that was the first time we scaled up the pre-training. We realized that if you actually do the data ablation really carefully, and the infra, and also the algorithm, we can actually push the pre-training quite a lot by the amount of 10x to make the model the best pre-trained-based model. And that’s why we built Colossus, the world’s supercomputer with 100,000 H100.

And then with the best pre-trained model, and we realized if you can collect these verifiable outcome reward, you can actually train this model to start thinking from the first principle, start to reason, correct its own mistakes, and that’s where the Grok 3 reasoning comes from. And today we ask the question, what happens if you take expansion of the Colossus with all 200,000 GPUs, put all these into RL, 10x more compute than any of the models out there on reinforced learning, unprecedented scale, what’s going to happen? So this is the story of Grok 4, and Tony, share some insight with the audience.

The Humanities Last Exam Benchmark

TONY: Yeah, so let’s just talk about how smart Grok 4 is. So I guess we can start discussing this benchmark called Humanities Last Exam, and this benchmark is a very, very challenging benchmark. Every single problem is curated by subject matter experts. It’s in total 2,500 problems, and it consists of many different subjects, mathematics, natural sciences, engineering, and also all of humanity subjects. So essentially when it was first released, actually like earlier this year, most of the models out there can only get single-digit accuracy on this benchmark, yeah.

So we can look at some of those examples. So there is this mathematical problem, which is about natural transformations in category theory, and there’s this organic chemistry problem that talks about electrocyclic reactions, and also there’s this linguistic problem that tries to ask you about distinguishing between closed and open syllabus from a Hebrew source text. So you can see also it’s a very wide range of problems, and every single problem is PhD or even advanced research level problem.

ELON MUSK: There are no humans that can actually answer these, can get a good score. I mean, if you actually say like any given human, like what’s the best that any human could score? I mean, I’d say maybe 5% optimistically. So this is much harder than what any human can do. It’s incredibly difficult, and you can see from the types of questions. Like you might be incredible in linguistics or mathematics or chemistry or physics or any one of a number of subjects, but you’re not going to be at a post-grad level in everything, and Grok 4 is a post-grad level in everything.

Like some of these things are just worth repeating. Grok 4 is post-graduate, like PhD level in everything. Better than PhD. But like most PhDs would fail, so it’s better to say, I mean, at least with respect to academic questions, I want to just emphasize this point. With respect to academic questions, Grok 4 is better than PhD level in every subject, no exceptions.

Now this doesn’t mean that at times it may lack common sense, and it has not yet invented new technologies or discovered new physics, but that is just a matter of time. I think it may discover new technologies as soon as later this year, and I would be shocked if it has not done so next year.