Skip to content
Home » Jensen Huang: Nvidia’s Future, Physical AI, Rise of the Agent, Inference Explosion (Transcript)  

Jensen Huang: Nvidia’s Future, Physical AI, Rise of the Agent, Inference Explosion (Transcript)  

Editor’s Notes: In this special episode of the All-In Podcast, the hosts sit down with Nvidia CEO Jensen Huang to discuss the transformative future of artificial intelligence and its impact on the global economy. Huang delves into the evolution of “physical AI,” the rise of autonomous agentic systems, and why he believes we are entering a “million-x” explosion in inference computing. He also provides a unique perspective on navigating geopolitical supply chains, the importance of open-source AI, and how specialized knowledge will become the ultimate moat for future entrepreneurs. (Mar 19, 2026) 

TRANSCRIPT:

Introduction

JASON CALACANIS: Special episode this week, we’ve preempted the weekly show and there’s only three people we preempt the show for: President Trump, Jesus and Jensen. And I’ll let you pick which order we do that. But what an amazing run you’ve had and a great event.

JENSEN HUANG: Every industry is here. Every tech company is here. Every AI company is here. Incredible, incredible, extraordinary.

JASON CALACANIS: And one of the great announcements of the past year has been Groq. When you made the purchase of Groq, did you realize how insufferable Chamath would become?

JENSEN HUANG: I had an inkling that —

JASON CALACANIS: — his friends, we have to deal with him every week.

JENSEN HUANG: I know.

DAVID FRIEDBERG: You had to deal with him for —

JASON CALACANIS: — the six week close.

JENSEN HUANG: I know.

CHAMATH PALIHAPITIYA: It’s like two weeks.

Disaggregated Inference and the AI Factory

JENSEN HUANG: Two weeks. It’s all coming back to me now. It’s making me rather uncomfortable.

The thing is, many of our strategies are presented in broad daylight at GTC years in advance of when we do it. Two and a half years ago, I introduced the operating system of the AI factory, and it’s called Dynamo. Dynamo, as you know, is a piece of instrument, a machine that was created by Siemens to turn essentially water into electricity. And Dynamo powered the factory of the last industrial revolution. So I thought it was the perfect name for the operating system of the next industrial revolution, the factory of that.

And so inside Dynamo, the fundamental technology is disaggregated inference. Jason, I know you’re super technical. Absolutely, I know it.

JASON CALACANIS: I’ll let you take this one. Go ahead and define it for the audience. I don’t want to step on you.

JENSEN HUANG: Yeah, thank you. I knew you wanted to jump in there for a second. But it’s disaggregating inference, which means the pipeline, the processing pipeline of inference is extremely complicated. In fact, it is the most complicated computing problem today. Incredible scale, lots of mathematics of different shapes and sizes.

And we came up with the idea that you would change, you would disaggregate parts of the processing such that some of it can run on some GPUs, and the rest of it can run on different GPUs. And that led to us realizing that maybe even disaggregated computing could make sense, that we could have different heterogeneous nature of computing.

That same sensibility led us to melanize. Today Nvidia’s computing is spread across GPUs, CPUs, switches scale up, switches scale out, switches, networking, processors. And now we’re going to add Groq to that and we’re going to put the right workload on the right chips. We just really evolved from a GPU company to an AI factory company.

CHAMATH PALIHAPITIYA: I think that was probably the biggest takeaway that I had. You’re seeing this fundamental disaggregation where we’ve gone from a GPU and now you have this complexion of all these different options that will eventually exist. The thing that you said on stage was, “I would like the high value inference people to take a listen to this.” And 25% of your data center space you said should be allocated to this Groq LPU GPU combo.

JENSEN HUANG: We should add Groq to about 25% of the Vera Rubins in the data center.

CHAMATH PALIHAPITIYA: So can you tell us about how the industry looks at this idea of now basically creating this next generation form of disaggregated prefill decode disaggregation, and how do you think people will react to it?

JENSEN HUANG: Yeah, and take a step back. At the time that we added this, we went from large language model processing to agentic processing. Now when you’re running an agent, you’re accessing working memory, you’re accessing long term memory, you’re using tools, you’re really beating up on storage really hard. You have agents working with other agents. Some of the agents are very large models, some of them are smaller models, some of them are diffusion models, some of them are autoregressive models. And so there’s all kinds of different types of models inside this data center.

We created Vera Rubin to be able to run this extraordinarily diverse workload. My sense is — and so we added what used to be a one rack company, we now added four more racks. So Nvidia’s TAM, if you will, increased from whatever it was to probably something, call it 33%, 50% higher. Now part of that 33% or 50%, a lot of it’s going to be storage processors. It’s called Bluefield. A lot of it I’m hoping will be Groq processors, and some of it will be CPUs. A lot of it’s going to be networking processors. And so all of this is going to be running basically the computer of the AI revolution called agents — the operating system of modern industry.

The Three Computers: Training, Simulation, and the Edge

DAVID FRIEDBERG: What about embedded applications? So my daughter’s teddy bear at home wants to talk to her. What goes in there? Is it a custom ASIC, or does there end up becoming much more of a broader set of TAM with developing tools that are maybe different for different use cases at the edge and in an embedded application?

JENSEN HUANG: We think that there are three computers in the problem at the largest scale.