Full Transcript

·YouTLDR

Nebius Co-Founder on AI Infrastructure Bubbles | How Price Elastic is Demand for Compute

1:14:23910 summary words · ~5 min readEnglishBy 20VC with Harry StebbingsTranscribed Jun 10, 2026
Analyze another video with Pro30-day money-back guarantee
Summary

The AI infrastructure market is not a bubble but in its infancy, where cheaper unit costs for compute trigger exponential, highly price-elastic demand (Jevons Paradox). To survive hyperscaler competition, independent GPU clouds must vertically integrate downstream into physical datacenters and build upstream software stacks, transitioning from raw bare-metal leasing to managed inference and agentic orchestration.

This video offers a rare, highly strategic look inside the economics, physical bottlenecks, and defensibility playbooks of independent GPU clouds competing against trillion-dollar hyperscalers.

Section summaries

0:00-1:52

Introduction & Context of the AI Capital Race

optional

The host introduces Roman Chernin, co-founder of Nebius, highlighting their rapid growth to a $6.6B market cap. They frame the core debate: whether massive, unprecedented capital expenditure on AI infrastructure represents a market bubble or the start of a transformative tech era.

It offers useful context on Nebius's market position but quickly transitions into the core thesis.

1:52-11:12

Debunking the AI Infrastructure Bubble & Jevons Paradox

watch

Roman explains why he believes we are only at the beginning of enterprise adoption, with coding being the first real use case at scale. He discusses the economics of closed versus open-source models, highlighting how cheaper intelligence increases aggregate consumption (using the DeepSeek launch as a real-world example where Nebius's sales spiked).

Crucial for understanding the microeconomic dynamics of AI compute demand.

11:12-18:40

The Four Layers of the Cloud Infrastructure Product Stack

watch

Roman outlines the four dimensions of Nebius's growth strategy, moving from physical megawatts to managed multi-tenant GPU clouds, managed inference (Token Factory), and speculative agentic orchestration platforms. He notes how this shift expands their addressable customer market from dozens of labs to tens of thousands of developers.

Explains the technical product strategy of modern AI cloud providers.

18:40-28:00

Pricing Elasticity and Customer Portfolio Diversification

watch

Roman discusses why Nebius targets a diversified customer base over ultra-concentrated relationships with hyperscalers like Meta or Microsoft. He addresses the elasticity of compute pricing, explaining how total cost of ownership (TCO) and software-level optimizations are more important to customers than the nominal cost per GPU hour.

Provides key insights into cloud business metrics, margins, and customer lock-in.

28:00-41:04

Managed Inference and the Rise of Niche & Specialized Models

watch

The conversation shifts to the transition from closed ecosystems (like OpenAI) to optimized open-source workloads using Token Factory. Roman explains techniques for making tokens cheaper (distillation, speculative decoding) and forecasts the ongoing rise of domain-specific models in life sciences, cyber defense, and robotics.

Highly relevant for developers and startups looking to optimize inference economics.

41:04-50:24

Enterprise Cold-Starts, Scaling, and Revolute Case Study

watch

Roman details how enterprise customers like Revolute transition to open source. He emphasizes the 'cold start' challenge—where companies must build robust internal evaluations and CI/CD pipelines before seeing exponential scaling. He reiterates that open source and frontier ecosystems will co-exist because of the sheer breadth of untapped workloads.

Gives a concrete playbook for how traditional enterprises adopt and scale AI workloads safely.

50:24-56:00

Sovereign AI, European Model Builders, and Nvidia Relations

optional

Roman analyzes Europe's position in the AI landscape, urging a shift in focus from sovereign power infrastructure to supporting local builder talent (e.g., Mistral, Lovable). He also demystifies the power dynamics of partnering with Nvidia, asserting that deep engineer-to-engineer respect is the true foundation of their relationship.

Interesting geopolitical and partner dynamic context, but secondary to the core business discussion.

56:00-1:05:20

Physical Constraints, Public Resentment, and Space Data Centers

optional

Roman breaks down the physical bottlenecks of datacenters (such as 12-24 month capital lag and permitting issues) and how to manage local community pushback pragmatically. He shares his optimistic perspective on space-based datacenters, arguing that the concentration of smart minds working on the problem makes it plausible mid-term.

Fascinating macro-infrastructure discussion, though highly speculative toward the end.

1:05:20-1:12:48

The Future of Labor, Consolidation Threats, and Execution Philosophy

watch

Roman shares his view on how AI democratizes software building and shifts human premiums to empathy and creativity. He flags market consolidation as Nebius's top threat and reacts to Leopold Aschenbrenner's massive investment position by emphasizing a disciplined, pragmatic focus on execution over market hype.

Summarizes the long-term existential outlook, organizational culture, and human labor implications.

Key points

  • Jevons Paradox in AI Compute Economics — Lowering the unit cost of intelligence (cheaper tokens or highly optimized open-source models) does not cannibalize infrastructure revenues. Instead, it unlocks previously cost-prohibitive enterprise use cases, driving an exponential surge in total aggregate consumption.
  • The Four-Layer AI Cloud Stack Strategy — To build a defensible business, an AI infrastructure provider must transition across four distinct layers: physical capacity (Megawatts of bare metal), managed cloud infrastructure (GPU hours), managed inference (buying optimized tokens via platforms like Token Factory), and agentic orchestration (paying for end-to-end task execution).
  • The Enterprise AI Adoption 'Cold Start' Problem — Large, non-native AI enterprises struggle to transition from closed API models (like OpenAI) to cheaper open-source models because they lack foundational infrastructure for model evaluations, automated testing (CI/CD for AI), and performance benchmarking.
  • Market Consolidation as an Existential Threat — The ultimate threat to independent GPU clouds is not direct competition, but extreme market consolidation. If the world ends up dominated by 3 to 5 closed model 'empires,' specialized infrastructure companies risk being relegated to commodity bare-metal suppliers with zero pricing power.
Every time we got intelligence cheaper... we are not reducing the consumption but we increasing the consumption because we can just solve more complex tasks with the same budget Roman Chernin
In the next 6 months, the capital cannot help. 6 months is too short time. You have what you have, you need to deliver. Roman Chernin

AI-generated from the transcript. May contain errors.

0:00

We are in the capital intensive game and

0:02

we competing with the most capitalized

0:04

companies in the world. Our program this

0:06

year is 2025 billion. Our competitors

0:10

hyperscalers have eight times bigger.

0:13

The AI infrastructure race is on. Capex

0:16

spend has never been greater. At the

0:19

center of this, Nebus. Today I'm joined

0:21

by the co-founder of Nebius, a company

0:23

that has scaled to a $66 billion market

0:27

cap, going head-to-head with some of the

0:29

largest hyperscalers in the world.

0:32

>> In the next 6 months, the capital cannot

0:34

help. 6 months is too short time. You

0:36

have what you have, you need to deliver.

0:37

The main threat for Nebios as a business

0:39

is the world will be too much

0:41

consolidated.

0:42

>> Today, we uncover the AI infrastructure

0:44

bubble and so much more.

0:46

>> It's like a shark. You're alive when you

0:48

move, right? So, we have to move. And

0:50

I'm thrilled to welcome Ronan Chernin,

0:52

>> who has power against Nvidia.

0:54

>> Ready to go.

1:06

Roman, I am so excited for this, dude. I

1:08

think Nebius is one of the most

1:10

unbelievable, incredible stories in

1:13

terms of what we've seen over the past

1:14

few years, but also like, holy [ __ ]

1:17

what an exciting few years we have

1:18

ahead. So thank you so much for agreeing

1:20

to do the show.

1:21

>> Yeah, thank you for inviting and uh glad

1:24

to be here.

1:25

>> Now I would love to start with a

1:27

question that I think is at the top of a

1:30

lot of people's minds which is like

1:31

where are we at in the insertion point

1:33

on AI infrastructure? A lot of people

1:35

are seeing the capital going oh it's a

1:38

bubble and a lot of people are going

1:40

it's just the start. How do you think

1:42

about the we're at an AI infrastructure

1:45

bubble moment right now? No, I I I don't

1:47

believe it's a bubble. Uh uh I mean

1:51

define the bubble. Do I believe that we

1:54

will need tens or hundreds times more uh

1:58

to build? I thoroughly believe uh I

2:02

probably biased. I would probably not

2:05

being in the business that we are doing

2:07

if I wouldn't believe. So I think that

2:10

we are just at the beginning of this

2:12

amazing moment when Jensen calls it like

2:16

useful AI and like we just at the

2:19

beginning of this kind of real adoption

2:22

and honestly we have maybe one use case

2:26

that works out of so many of use cases

2:30

and the one use case that works like

2:32

coding everybody's talking about coding

2:35

started working like maybe few months

2:37

ago. go just so like let's put it in the

2:41

perspective we just few months from the

2:43

moment when we've got maybe first use

2:46

case that's works in like in the scale

2:50

and we start seeing it's applying here

2:54

and there and uh I think we'll see

2:57

obviously like many many many more use

3:00

cases and we will see much more adoption

3:05

uh I think that what we see yet is if

3:09

you take every single company in the

3:11

world uh maybe outside of the fastest

3:14

moving startups before the show we we

3:16

were speaking with like who is moving

3:18

fast enough or not fast enough right so

3:20

maybe there are some exceptions but if

3:23

you take I think if you take any company

3:25

in the world today and you will see

3:28

we'll look at the AI adoption there you

3:32

will actually see that they start using

3:35

AI in a first percent of the volume in

3:40

the first percent of the use cases. So

3:43

if you take any large company even

3:46

pretty advanced technologically

3:48

you will see that they just starting and

3:51

I'm taking from that that we only only

3:53

beginning uh and uh even you even if you

3:57

don't believe in what Musk says about

3:59

everything in the future space and so on

4:02

just practically

4:04

uh from from uh enterprise adoption it's

4:10

it's just the first steps

4:11

>> so we're completely align but it's a

4:13

very boring discussion if I just go I

4:15

agree with you on on everything. Um my

4:18

my question to you on on the back of hey

4:19

we've seen coding now work for the last

4:21

whatever 6 to 12 months. Yes, but there

4:25

is a question that we will move to open

4:27

source models locally hosted because the

4:30

cost will be too significant for some of

4:33

these enterprises to burden and we're

4:35

going to see that shift happen soon. If

4:38

we do that is both damaging to the

4:41

providers open AI and anthropics of the

4:43

world and to anas. Why is that

4:46

perspective wrong? Uh yeah. So first of

4:49

all I think that uh it's not in the

4:53

future it's already in the present. Uh

4:55

again what we see in a lot of examples

4:58

at the moment when our customer or the

5:01

product builder gets to the scale they

5:05

start uh looking uh to the ways to

5:09

improve the economics or accelerate the

5:12

growth and so on and this is the way

5:14

when they most a lot of them start to

5:18

look to alternative models. So the best

5:21

way to build today is obviously to build

5:23

on the frontier models from great

5:26

providers like OpenAI, entropy, Google

5:28

because they actually provide and that's

5:30

true they provide you the best best

5:32

capabilities in the world.

5:35

But then when you figure it out the use

5:37

case, when you start seeing adoption,

5:39

when you see the customer data loop, you

5:43

maybe can find the cheaper or

5:49

even not cheaper but more quality high

5:53

quality way to serve the same use case.

5:55

You don't need maybe you don't need the

5:58

best in the world universal model but

6:00

you can create the specialized model

6:02

that in your particular case will work

6:04

even better and that's the way where you

6:07

need to shift or may consider to shift

6:10

from uh frontier closed models to open

6:14

source. The most important kind of

6:19

cause of those models is not just they

6:21

open source but they are tunable they

6:23

are trainable. So you can take them and

6:25

you can do something you you can

6:27

postrain them and you can create the

6:29

specialized model that in your

6:31

particular case may work better. So

6:33

that's we see over the world over over

6:35

the use cases but why doesn't it hurt uh

6:39

entropic and open AI because in reality

6:44

the they move to the next frontier and

6:48

to the previous uh point that we

6:51

discussed there are so many unsolved

6:53

tasks yet or the tasks that not

6:57

necessarily have the limited budget uh

7:00

uh to be solved And every time we see

7:04

and we saw it like with deepseek one

7:05

year ago and like we continuous see

7:08

seeing it now. Every time we find the

7:11

way to solve some task uh more efficient

7:16

we just start solving more complex

7:18

complex task in the same time and this

7:22

is like continuous journey I believe. So

7:25

you you kind of you always push the

7:27

frontier. You always have the more

7:30

complex tasks to to figure out how to

7:32

solve. When you figure it out how to

7:35

solve them, you can go down and reduce

7:37

the price or improve like the quality.

7:40

But we have so many unsolved tasks that

7:43

entropics and opening and all other

7:45

frontier models still have such a not

7:48

addressed yet time, not addressed yet

7:51

market. they continue kind of

7:53

exponentially grow.

7:54

>> Do you buy that? These companies are

7:57

priced to to perfection in a lot of

7:59

cases at a trillion dollars. If the

8:01

value that they create is eroded and

8:03

they're constantly playing a game of

8:05

leaprogging from value to value to value

8:07

while open source continuously eats

8:09

behind them.

8:12

You got to find a lot of problems

8:14

continuously dude. That's a that's a

8:16

hard life to live. Actually most of the

8:18

people are concerned on other side like

8:20

will we have a strong enough open source

8:22

and strong enough specialized models

8:25

environment to to build this floor. Uh I

8:29

think that we yet at such a early point

8:32

of adoption we have so many unsolved

8:36

problems yet that uh um it's just the

8:41

matter of the total pie and uh I think

8:45

there is uh enough space to solve so

8:47

many tasks in the future that it's

8:50

enough uh enough pi for uh both frontier

8:54

capabilities uh and very tuned

8:58

uh models for specific use cases and all

9:01

the world of these open source or

9:03

specialized models that we can build on

9:05

top of them uh to gain these economics

9:10

advantages and uh performance advantages

9:12

when we know what we need. You said that

9:14

like every time we we have the cheaper

9:16

model is it hearts uh the business and

9:18

my favorite anecdotal story about that

9:20

is I think 15 months ago or so there was

9:24

this deepseek moment if you remember.

9:26

clue. I remember that Nebul stock went

9:28

down 40% in one week or so uh uh in

9:32

February I think it was February or

9:34

March 2024 2025 and anecdotal story the

9:39

same the same exact week we probably had

9:42

the best week in sales. So uh people on

9:46

the market were concerned that market is

9:49

going down and like infrastructure

9:51

companies like Nebio's

9:53

not needed because okay if AI is such so

9:56

much cheaper maybe it's bubble but at

9:59

the same time we never had the best

10:02

commercial week we were pretty early in

10:04

our story but that was the best

10:06

commercial week in the history of the

10:08

company because so many people figured

10:10

out that they can run inference in their

10:14

production workloads uh with DeepSeek

10:17

and economics will work and then like at

10:21

the at the same time like Ktor started

10:23

growing uh I think they were the first

10:26

who really benefited from uh tuning

10:29

those models for coding and so on. Every

10:31

time we got intelligence cheaper

10:34

uh this the same unit of intelligence

10:36

cheaper, we are not reducing the

10:38

consumption but we increasing the

10:40

consumption because we can just solve

10:44

more complex tasks with the same budget

10:46

or we can finally

10:49

uh economically viably solve the tasks

10:52

that we already kind of knew that was

10:54

solvable but economics didn't work and

10:56

we could not scale. So I think it's

10:59

quite uh fascinating what's like this

11:01

economics improvements to observe.

11:04

>> Speaking of kind of Jeav's paradox and

11:06

producing more and that yielding more

11:08

demand where are you not moving fast

11:11

today where you would like to be moving

11:13

faster

11:15

>> everywhere. So when we think about how

11:17

we build a company we talk about it in

11:19

the four dimensions. One dimension is

11:22

capacity like uh how much megawatts,

11:25

gigawatts and the GPUs we deploy. We are

11:28

infrastructure company. We need to be

11:30

large. If you're not large enough,

11:32

nobody needs us to exist. So this is the

11:36

the the the physical world expansion.

11:39

The team is doing amazing job. uh but

11:43

it's never enough and you want to move

11:45

as fast as possible and uh uh there are

11:48

a lot of complications of the real world

11:50

that prevent you to move fast enough

11:53

sometimes

11:54

uh to launch new data center you need to

11:57

go through the entire like supply chain

12:00

regulatory and uh fires and waters uh

12:03

and uh uh everything that happens in the

12:06

real world right so this is one

12:08

dimension

12:09

uh another dimension

12:11

is the product. So you want to move fast

12:16

enough to address new types of the

12:19

workloads, new types of the customers

12:21

that coming to the market. Think about

12:23

it. We started as a industry in this AI

12:29

journey uh from the people who first of

12:32

all built the models right. So there the

12:35

that that was like companies like OpenAI

12:38

in harpscalers large labs and so on and

12:42

what they need from you as an

12:44

infrastructure provider is barely

12:47

compute like just throw infrastructure

12:49

and we see a lot of these large bare

12:51

metal deals on the market and we also do

12:54

them uh but this is only the first layer

12:57

of what we build like scaled physical

13:00

infrastructure that customers like

13:02

Metal, Microsoft in our case can consume

13:04

on a large volumes. This is the first

13:07

layer. The second layer is what we

13:09

called multiden cloud. Uh still

13:14

addressing

13:16

um research heavy teams

13:19

but now we have hundreds uh or thousands

13:24

teams that want to have they don't want

13:27

to deal with the physical

13:28

infrastructure. They want to deal with

13:30

the managed infrastructure. classical

13:32

infrastructure as a service in the cloud

13:34

terms. You have storage, compute,

13:37

networking, virtualized and a good

13:39

environment with API, observability,

13:41

security, everything that normal teams

13:44

expect cloud to have. You log in, you

13:47

get your cluster provisioned and you can

13:50

start training or run inference if you

13:53

need uh and manage your application or

13:57

manage your workflow yourself but have

14:00

infrastructure figure it out for you.

14:02

Right? So still if if the first layer

14:05

speaks in megawatt and literally like if

14:09

you read announcements someone signed a

14:12

large deal with some like Meta or

14:14

Microsoft or OpenAI people speak

14:17

megawatts there uh so it's like you

14:20

deliver the megawatts of compute then

14:22

when you speak about this managed cloud

14:25

people speak GPU hours because this is

14:27

the key unit you sell the efficient

14:31

hours you

14:32

spent on compute with storage with

14:35

complimentary services but you still buy

14:38

managed by com but compute then the next

14:41

layer that we working is managed

14:43

inference when people don't want to go

14:45

in terms of GPU hours they don't want to

14:48

figure out B200s against H200s against

14:52

B300s what is better for particular

14:54

workload they don't want to manage the

14:58

you know VLM or SGAN deploy themselves

15:01

do all the optimizations and here like

15:04

our product called Nebio stocking

15:05

factory this is a managed inference

15:08

platform and again this is the new type

15:10

of the customers mostly people who we

15:13

call them vertical AI companies or

15:15

enterprises so people who actually build

15:18

products they don't do models they build

15:20

progress products on top of the of the

15:23

and this is to your point of specialized

15:25

and open source models when they need to

15:27

shift from entropic for example or

15:29

diversify um the models they use for

15:32

that. So, and again this is the new

15:36

primitive that we provide or the new

15:38

kind of uh new entity that customers

15:41

need. Now we speak in tokens. It's not

15:44

you pay for GPRs,

15:46

you you consume tokens and you can build

15:49

your applications not thinking in terms

15:52

of the clusters underneath.

15:55

And this is where we sit now. But it's

15:58

also I think not the final stage of of

16:01

where we going because now people build

16:04

agentic uh agentic applications agentic

16:07

workflows and when you build uh end to

16:10

end agent you may not even think in

16:13

terms of the particular model and you

16:15

not think you may not think in terms of

16:17

particular number of tokens that you

16:19

want to generate you you want the end

16:22

toend task to be uh efficiently

16:26

executed.

16:27

and provide the expected outcome. And

16:29

then the magic that platform can make is

16:33

actually think for you which model

16:36

better to to use in this particular

16:39

call. uh do you need to go to the

16:42

smarter model or you know you can ask

16:44

two time in the same inference budget

16:46

you can request two models uh lighter

16:50

models and get you know less smart

16:54

tokens and then have the judge model

16:56

that chooses the best result or what

16:59

size of context you should have and so

17:02

on. So this is the next layer when

17:05

developer would maybe not even think in

17:08

terms of particular

17:10

you know types of the tokens but thinks

17:12

in terms of end to end execution of

17:14

their task.

17:15

>> So that's a d layer 4 is a direct

17:17

competitor to open router.

17:19

>> What we would love to bring on that

17:21

level is this the same like what we do

17:24

on the layers below is the optimization

17:28

engine.

17:30

uh you can build your agent in so many

17:34

kind of open source or appropriate tools

17:37

but then when you need to scale it, you

17:41

start thinking about the economics, you

17:44

start thinking about reliability,

17:45

reliable execution like repeatable

17:48

execution. And this is like where it's

17:53

not just like model choice problem. It's

17:56

not just like the outcome problem, but

17:59

it's a system problem. You need to make

18:02

it reliable. You need to make it

18:04

repeatable. And you need to make it

18:05

economically viable. And that's probably

18:08

where Nebios could create the value. The

18:11

same way like we don't tell people how

18:13

to build their applications. We just say

18:15

okay, if you need this model to work for

18:16

you like with this economics, we will

18:18

help you to optimize. The same here. If

18:20

you need this agent end to end run with

18:23

this budget, with this quality, maybe we

18:25

can help you uh to optimize it. And

18:29

again, this is just to make it sure it's

18:32

a kind of a little bit speculative

18:34

thinking about what what's next. It's

18:36

not like what we already have, but this

18:38

is where we think where we see our

18:41

customers evolving and where we think

18:44

that we could create the next kind of uh

18:47

the next layer of the product ordering.

18:49

I love this and I I have all of these

18:51

notes. Um and uh I just want to actually

18:54

go through the four pillars that you

18:55

said there. You said number one,

18:56

capacity.

18:57

>> Yeah.

18:58

>> If you had 10x the capacity today,

19:01

what would be different? Like could you

19:04

sell it overnight?

19:05

>> Yeah. Yeah, it's a good question. Not

19:08

overnight, but we would definitely we we

19:09

definitely have demand for that. uh and

19:12

I think the key question for us it's not

19:16

uh do we have demand or not but how we

19:20

actually build a portfolio of demand

19:23

because you have so many customers on

19:25

this market that you can balance between

19:28

and again to the point of four layers of

19:31

the product. You can sell bare metal,

19:33

you can sell managed customers, managed

19:35

infrastructure, you can sell inference

19:38

and maybe in the future you can sell uh

19:41

some new layers of product. And I think

19:46

what we try to do is to build kind of

19:50

quite diversified portfolio of uh

19:52

customers. We we we believe that the

19:56

higher stack we move the more value

19:58

potentially we can create for the

20:00

customers and actually the higher stack

20:03

we move the bigger population of the

20:06

customers we can serve because again

20:08

like on bare metal level you have maybe

20:09

dozen of the customers in the world that

20:11

you can work with on uh managed

20:14

infrastructure there are hundreds on

20:17

inference there are thousands on a

20:18

gentic there will be tens of thousands

20:21

of new developers that build it Right. I

20:24

Okay. On the kind of customer portfolio,

20:26

I love that for the capacity.

20:28

Absolutely. You want to be big enough

20:30

that you're meaningful, but not too

20:33

large that the business relies on them.

20:36

With that difficult awareness,

20:39

where do you settle on what revenue

20:43

concentration with a meta or a Microsoft

20:45

you're happy with?

20:46

>> It's a great question and I I would say

20:48

it's a it's a main question of our

20:50

business. uh I mean not nebios even but

20:53

the product category and we always told

20:56

it and publicly and uh to our investors

20:59

and to our customers that we believe

21:02

that long-term strategy of Nabios is to

21:05

serve as much diversified portfolio as

21:07

possible. So we we do the best to to to

21:11

have many customers that we work with.

21:14

We build the platform. If you in reality

21:17

again to serve dozen of the customers of

21:20

the world on like of the level of meta

21:22

and Microsoft which super advanced and

21:25

they have their entire software stack,

21:27

they literally need only physical

21:29

infrastructure. They bring with they

21:30

bring everything they have deploy on

21:32

your infrastructure and run right. Uh

21:34

you have a tiny tiny uh additional value

21:38

that you can provide them above the the

21:40

the physical infrastructure. by the way

21:44

to to to satisfy them with what they

21:47

need on physical infrastructure is quite

21:49

a challenge because you can imagine they

21:51

are quite demanding and and and and they

21:53

need the like the most scaled

21:56

infrastructure in the world that exists.

21:58

So uh sometimes people say it's

22:00

commodity but it's not really commodity

22:03

on that scale like nothing commodity

22:05

when it comes to the to the real scale.

22:08

But again to your point uh uh this is

22:12

quite a small population of the

22:14

customers that you can work with and you

22:16

not necessarily need all the full stack

22:19

software to to work with them. So we

22:23

intentionally

22:24

building and from day zero of nebios we

22:28

were building this software stack

22:30

because we thought that it's a much more

22:32

beneficial for us and if I want to be

22:35

pathetic for the world uh to have

22:38

someone who can support customers uh not

22:41

only on this physical infrastructure

22:43

layer but beyond

22:45

>> for the long-term protection of the

22:47

business. Do you not have to build the

22:49

full stack? Because otherwise you become

22:52

the capacity provider to these mega

22:54

players which will make a [ __ ] ton of

22:56

money.

22:56

>> Yeah.

22:57

>> But you're incredibly concentrated and

23:00

very vertically focused.

23:02

>> Yeah, I think I I I think so. And again

23:05

we don't know where the world will end

23:07

up like and in the world of infinitive

23:10

demand uh you you may uh sustain

23:15

even long-term and midterm uh selling

23:18

this like bare

23:21

uh whole bare metal kind of uh

23:24

contracts. Um but the more competition

23:28

let's say you have from the customer

23:30

from demand side uh you you can be picky

23:34

uh even with the customers you you work

23:36

with and work with the customers that um

23:41

appreciate like that value uh the

23:44

platform that we built more and there

23:46

are different customers in the world.

23:48

Someone more obsessed about the price,

23:51

someone more obsessed about the quality,

23:53

someone really want to have much more

23:56

advanced platform because they want to

23:58

concentrate focus on their platform uh

24:01

or product and don't spend time on the

24:04

Yeah. Before before we move to number

24:06

two being product just staying on

24:08

capacity given the insufficient supply

24:11

of capacity today if you doubled pricing

24:16

would you see any change to demand?

24:19

>> Oh it's a it's a it's a difficult

24:21

question. Uh we actually raised prices

24:23

like just% yeah just couple months ago.

24:28

uh uh and we still uh still have fair

24:33

fair kind of pipeline pressure let's say

24:37

uh uh on supply and again uh if we the

24:41

question we we we don't really know

24:43

where is the balance uh and I will tell

24:45

you why uh it's not only us being greedy

24:49

and want to get like as much money and

24:52

like then people in the shortage will

24:54

still will have to pay for some extent

24:56

it works like people needs compute to

24:59

build but then there is a point and

25:03

especially it's less in in training

25:05

because in training it's like oneoff

25:07

cost but if you believe that we're

25:09

moving to inference and inference is a

25:13

uh is the cost of serving the customer

25:16

there is a a a level where economics

25:21

doesn't work and the economics of the

25:25

products of our customers ers if they

25:27

work they can grow and then we can grow

25:29

with them. It's not like just supply

25:32

demand situation and then absolutely

25:35

elastic prices. They are elastic for

25:38

some extent. uh but we also want to be

25:42

meaningful and we want to be thoughtful

25:45

kind of what our customers need and by

25:48

the way it's not only GPU hour cost it's

25:51

all the optimizations you do all the

25:53

real we call it TCO total cost of

25:55

ownership that you in like and this is

25:59

partially why we build the software

26:01

platform and I'm sorry come back to

26:03

product again and again you want to

26:04

speak about capacity but people too much

26:06

obsessed about capacity like capacity is

26:09

important too too much obsessed about

26:11

the nominal price of capacity. You can

26:14

price GPU $3, $4 and $5.

26:19

And depending on the use case and

26:22

depending on on the quality of the

26:25

platform, it can create completely

26:27

different outcomes for the customer in

26:29

real cost.

26:31

How long it works? Well like what what

26:34

is the e effective kind of uninterrupted

26:37

kind of time that you can run there. If

26:40

you talk about inference how much tokens

26:42

you can extract we we see all these

26:45

optimizations that happening that

26:47

changes the price of the tokens in order

26:50

of magnitude. So people so much speak

26:53

about the cost of particular GPU but if

26:56

you do the right thing with the model

26:58

you can change the price like uh in the

27:02

times and this all should work together

27:07

uh uh as a system not just as a again if

27:10

you speak about role infrastructure then

27:12

you can manage only the price but if you

27:14

if you if you build the platform and if

27:17

you provide the high level of the

27:18

service to the customer then you can

27:20

extract much more economics not only

27:23

from the infrastructure cost structure

27:27

right

27:27

>> if we move to that second layer then if

27:29

we move away slightly from capacity to

27:32

GPU hours the product itself

27:34

multi-tenant

27:35

>> what is the main question that you ask

27:38

yourself within that segment if in the

27:40

first capacity it's how much revenue

27:42

concentration we have what is the big

27:44

question in that layer of value

27:46

>> what customer needs it's a normal you

27:49

know you speak with a lot of product

27:50

founders

27:51

uh uh and this is the same like what

27:54

customer needs at the end of the day how

27:57

customers evolve in their needs where is

28:00

the demand moving so it's like we see

28:04

all this transition from training to

28:06

inference we see transition from uh uh

28:09

just using the models to building agents

28:12

uh and we see the transition from mostly

28:16

AI labs uh consuming AI compute to

28:19

enterprises coming in game and all the

28:21

time if we want to be relevant we need

28:23

to follow the changes and this is the

28:26

main question which like we we ask us in

28:28

the in the product like what should what

28:31

customer needs and what is neio's what

28:34

is our value that we need to create

28:36

because again we are small company we

28:38

cannot build everything uh and we need

28:40

to be very precise on what we can do

28:44

better than others and where the value

28:47

that we should focus on uh uh given how

28:50

customers evolving.

28:52

>> What changes are you seeing in customer

28:54

needs that you're not seeing discussed

28:57

much in public?

28:58

>> Everybody's talking about this moving

29:00

from training to inference. I think it's

29:02

just very uh uh 100,000 ft like view

29:06

because this move means actually people

29:11

um build specific products and in those

29:16

products they have their economics they

29:18

have their trajectory of growth and it's

29:22

not just like whatever the same GPU is

29:25

just used uh uh for other purposes I

29:29

think it it brings the new requirements.

29:32

You you need to build your inference

29:34

platform. You need to help your

29:36

customers not only run inference but

29:38

where the model that they inference come

29:40

from. Everybody is taking open source

29:43

models and fine-tune or them. So how do

29:46

we help them and then when they run them

29:48

they generate a lot of data. How do we

29:51

help our customers like when they

29:54

already run their application their

29:55

inference to collect the data to create

29:57

it and then use it to improve uh uh uh

30:01

to improve the model or the application

30:04

that they run. So it's a people like

30:08

this flywheel analogy like uh you you

30:11

you run inference you generate data you

30:14

can observe this data then you can

30:16

improve the model uh that you run and

30:19

kind of continue continue uh um improve

30:23

the quality uh of the end product. So I

30:27

think uh there are a lot of pieces both

30:30

on system level and both on uh um AI

30:35

magic level if you want. Uh and I think

30:38

the the most fascinating m moment for me

30:40

is that I think what we see is that

30:43

barrier to build is going down. So we

30:47

see more and more customers like

30:49

builders coming to the market that not

30:52

necessarily AI researchers or not

30:54

necessarily inference engineers. Uh and

30:59

the value that companies like Nebios can

31:01

create is actually to lower the barrier

31:05

uh to to build AI enabled enabled

31:10

products and AI enabled applications

31:14

that really work and incorporate like

31:18

hide from the developer all the

31:19

complexity of infrastructure, all the

31:22

like

31:24

some complexity of AI like how you tune

31:27

the model or how you optimize the

31:29

inference. It's a lot like research

31:31

heavy area as well and just let people

31:33

focus on their customers and uh and use

31:38

case by the way the same way like they

31:40

do with uh with the closed ecosystems

31:42

like entropics open ais you mentioned

31:45

the word differentiation

31:47

and and one thing that I was discussing

31:49

with my partner before that we have as a

31:51

theme we have to discuss and it's within

31:52

these layers but you've spoken

31:54

extensively about product buildout and

31:56

the importance of building the product

31:57

underneath capacity when People look at

31:59

you versus other Neo clouds, you know,

32:02

we look at you versus a corewave. You

32:04

both run GPUs, you both have Nvidia

32:06

relationships, you both have meta as a

32:08

customer.

32:10

What's the difference? I don't like

32:12

compare with others. The principles we

32:14

build are full stack. We call it full

32:17

stack integration. And you you can think

32:19

about it like full stack down and full

32:21

stack up. Full stack down is we're

32:24

really deep in physical world. We build

32:26

data centers. We built racks and

32:29

servers. We built the platform. And then

32:33

uh and when you control this kind of uh

32:37

things downstream, you can move faster

32:40

and you can squeeze more

32:44

uh cost and provide more economically

32:48

viable solutions for the customers. And

32:50

then your vertical integration upstream

32:54

is actually what we spoke about like

32:56

product and how can you follow the

32:58

customer's needs and customer segments

33:01

and not be limited by the small

33:04

population of the people that just need

33:06

infrastructure but really serve kind of

33:08

enterprises and product companies uh

33:11

with like meet them where they need us.

33:14

And this is like I think what we

33:16

different and then how it how it like

33:21

showing up I would say is again uh less

33:25

concentration in the in the in the in

33:27

the business more diversified customer

33:28

portfolio uh uh we believe long-term uh

33:33

better positioning for going to

33:36

enterprises where we believe eventually

33:39

a lot of demand will come from again now

33:42

most of our segment is working it's AI

33:45

natives working with AI natives but we

33:48

have a huge economics uh huge market of

33:52

enterprises existing companies and

33:55

someone needs to serve them and uh uh

33:59

they will not buy ro compute they will

34:02

need platforms they will need tools they

34:04

will need uh us to respect their legacy

34:08

and being able to work with their more

34:11

complex environment they're not nimble

34:13

they have data to migrate, they have

34:16

systems to integrate and that's that's

34:18

the big game and I think that for us

34:21

it's kind of the main uh direction to

34:25

move.

34:25

>> You mentioned the third layer of the

34:27

four-pillared stack being managed

34:29

inference for people that don't

34:31

understand how do you think about this

34:33

layer and how would you explain it to

34:35

them?

34:36

>> Yeah, very simple. uh you you built your

34:40

product on whatever call your where you

34:42

wipe code.

34:43

>> I I'm I'm actually an open AI in a

34:46

>> code. Okay, good enough. Uh you built

34:50

your great product with OpenAI. Uh you

34:54

you cracked the use case uh and you

34:58

started growing and you have amazing

35:01

traction. The only problem may be that

35:07

uh you don't have enough margin or you

35:12

want to start applying more aggressively

35:14

the data and tune the the behavior of

35:16

the model and you cannot do it in the

35:18

closed ecosystem. So you you you go to

35:21

internet and you read there is there are

35:23

a lot of great open source models that

35:26

on the benchmarks are close to open AI

35:29

and you think oh great it will be 10

35:31

times cheaper inference is cheaper I can

35:35

tune those models I can apply my data

35:37

and my product will be better my growth

35:39

will accelerate. So you you go you you

35:43

take the weights from hugging face you

35:45

take uh some uh engine to run it like

35:49

vlm sg lang something and then it

35:52

doesn't work uh because

35:56

uh you need to to to really extract the

35:59

value you expect you need to do like

36:01

optimizations you need to deploy it in a

36:03

proper way you need not just one GPU

36:08

tokens extraction or one host

36:10

uh setting but you you have the large

36:12

product you run on hundreds of thousand

36:14

GPUs already you need all the

36:16

orchestration you need the caching you

36:18

need uh you need the observability like

36:22

your customers ask you like how does it

36:24

work and so on so forth and by the way

36:26

you had all of that on open AI because

36:29

this is like the production service for

36:30

you it's you don't think about

36:32

infrastructure when you work with open

36:34

just subscribe for the the plan you need

36:36

and you pay for whatever uh end result

36:41

and so that's where you need the product

36:43

like token factory uh you token factory

36:47

gives you the managed inference with the

36:50

open source or specialized models you

36:52

can run existing open source vanilla

36:54

open source model or you can tune the

36:56

model and deploy your own like weights

36:59

and then we'll take care about all the

37:01

rest we'll apply all the optimization

37:03

techniques we'll uh manage the better

37:06

economics for you it will be reliable

37:08

you don't need to think about the next

37:11

100 GPUs where you will find them uh and

37:14

so on so forth. It's service. It's like

37:16

managed managed service.

37:18

>> With Token Factory, you run on 60 open-

37:20

source models. And you said before about

37:23

cutting inference cost by up to 70%

37:25

through optimization.

37:28

Can I ask a dumb question which is how

37:30

do you actually make a token cheaper?

37:33

>> Yeah. So the it's not the magic again.

37:36

you take the model uh the b like some

37:40

baseline model and then you can optimize

37:42

it for particular scenarios that uh that

37:45

you have. So you can do

37:50

actually you can distill the model you

37:51

can make the same like the smaller model

37:53

that works uh uh with the same quality

37:56

you can do spec decoding you can

37:59

optimize caching uh uh and so on so

38:02

forth. So you take the model and out of

38:04

this model you actually build a system

38:06

that in your particular case works with

38:09

your with with your requirements with

38:12

optimized economics. And by the way, one

38:14

of the things that uh also um I think

38:18

important for customers to use managed

38:20

platforms like token factory, the models

38:23

are changing every week, every month

38:26

like right today maybe minimax 3 was

38:31

released and there is ultra that was

38:34

announced released. So, and this happens

38:37

every few weeks and every time the new

38:39

model released, uh, it may work better

38:42

on some benchmarks and maybe not like on

38:46

other benchmarks and so on. And you want

38:48

to have flexibility. You want, uh, you

38:51

want someone to support you on

38:53

experimenting and actually adopting the

38:56

new best models for your use case every

38:59

time they come online. And then like the

39:02

platforms like ours uh actually again

39:05

abstract from you all the work that you

39:08

need to do to actually like change from

39:10

one model to another to benchmark all of

39:12

them and so on. So you you you can be

39:15

sure you can be sure that you will be on

39:17

the frontier like every time something

39:19

new is happening it will be in the

39:22

platform you will be able to test it you

39:24

if it works better for your use case you

39:26

will be able to switch and it all will

39:28

be kind of smooth and transparent for

39:31

you

39:32

>> does the pace of model development

39:34

sustain like you said there I would

39:36

argue respectfully you said every couple

39:39

of weeks I'd say every couple of days

39:41

there's new does that sustain in 5 years

39:45

time are we seeing that level of

39:46

iteration?

39:49

>> Well, I don't know. It's a good chances

39:51

that we'll continue to see a lot of

39:53

niche models uh show up and improved. Uh

39:59

I don't know again I'm a believer that

40:01

we quite far from the wall and we will

40:04

see a lot of like uh models improvement

40:09

uh happening. I think that what we also

40:12

see is

40:14

much more new like modalities and

40:18

specialized models uh coming in game. So

40:21

we speak about this frontier alms but

40:23

there is entire world of life science

40:27

models robotics

40:29

uh world models video models image

40:32

models uh so and they all have their own

40:36

use cases as well and we see more and

40:40

more like small specialized models

40:43

particular use cases coming like very

40:45

much optimized. Just this morning, I

40:47

spoke with a team here in Israel that

40:50

develops uh uh cyber defense uh

40:54

foundational model like the model that

40:57

optimized for to build uh cyber defense

41:01

uh agents and again they don't start

41:04

from the scratch. they they take some of

41:07

the foundational like uh some of the

41:09

open source foundational model but then

41:11

they train it for the particular case

41:14

optimize for the quality and the latency

41:17

that needed in this like cyber defense

41:19

use cases and I think we'll continue to

41:21

see it we'll see a lot of specialized

41:23

bolt post trained models that still need

41:28

uh optimized inference and optimized

41:30

like uh infrastructure around them uh to

41:34

let customer use

41:35

Can I ask going back to token factory on

41:38

token costs and token usage? What are

41:42

you seeing that you don't think other

41:44

people are talking about enough? What

41:46

has shocked you recently? Again I I

41:49

think everybody is speaking the same

41:50

thing like how fast it's growing uh uh

41:54

when we see this uh uh trajectories of

41:56

some companies like entropic and cursor

41:59

and cognition in coding and now we see

42:03

start seeing in other verticals as well

42:05

uh some uh healthcare examples some uh

42:10

financial uh use cases I think uh I

42:13

think it's like quite amazing what it

42:16

what's interesting is to see how like

42:19

noni startups are moving. So we I can

42:22

give an example. We we have the customer

42:24

of uh revolute uh and uh when we started

42:29

working with them I think 99%

42:34

of their budget inference budget was in

42:36

uh uh closed models in open AI

42:40

uh and they started to crack some of the

42:42

use cases and some of them didn't work

42:45

for them economically. So they

42:47

practically couldn't replace the humans

42:50

or cannot enhance the humans uh in the

42:54

in the use cases they wanted to address

42:56

and they started moving to open source

42:58

models but it didn't move fast for them

43:02

because they had to spend time on

43:05

building the entire engine internally in

43:08

the company and first of all they were

43:11

focusing on evaluations. So, and I think

43:13

this is something that people

43:14

underestimate

43:16

how important to build kind of the

43:19

foundation for improvements and

43:22

experimentation engine. Uh when you

43:26

understand as a company, as a team what

43:28

is good for you because again like you

43:31

you you you close some use case it works

43:34

but then you want to change the model.

43:37

How do you know you don't uh you don't

43:40

ruin the quality? you need to have like

43:43

metrics, you need to have a valve

43:44

mechanism, you have you need to have

43:46

this CI/CD process uh established for EI

43:50

development.

43:52

And I think that what we see a lot of

43:55

customers like Revolute, they have this

44:00

foundational investments that need to do

44:03

in the understanding of how to evolve

44:05

the models, how to actually safe safely

44:09

integrate them in their production

44:12

processes.

44:13

But when they solved these foundational

44:16

problems, they start growing

44:18

exponentially.

44:20

And I wouldn't

44:22

uh underestimate

44:25

how fast those customers can grow when

44:28

they build the system that let them ship

44:30

fast. And ship fast means they know how

44:34

to evolve. They know how to make a

44:35

decisions.

44:37

And this is something that we see across

44:40

a lot of customers. They have this you

44:43

can call it foundational investments or

44:45

cold start problem. How to start

44:47

shipping. But when they solve it, they

44:51

start to grow exponentially and they can

44:53

use different models. They can build

44:55

much more products inside the company

44:57

and so on so forth. And I think this is

45:00

this is something that kind of when you

45:02

look from outside you kind of oh they

45:05

are not growing they start small they

45:07

take time and so on. But this is in a in

45:10

a if the company has a strong team they

45:15

build this foundation and then they

45:17

start growing exponentially and I think

45:19

we'll see a lot of explosive growth

45:22

in enterprises in the digital like in

45:25

the cloud companies in cloudnative

45:28

companies like revolute Shopify pro

45:32

booking.com when they solve this cold

45:34

start problem they build the system how

45:37

to ship and then they will grow like in

45:40

their AI adoption like crazy.

45:42

>> How much more do you think Revolute will

45:44

pay you in 3 years time?

45:46

>> I don't know. I don't know. No, I don't

45:48

want to speak about that. No, but I I

45:50

can say that like they in total I think

45:54

they they grow times like they grow like

45:58

this. We we all see this AI companies

46:00

reporting IR growth, right? For them

46:03

it's not IR, it's like their budget. But

46:06

I think that the most advanced companies

46:09

their AI budget and it's not like this

46:11

fake or not fake like this more all this

46:15

maxaxing kind of race uh we see it like

46:20

how they do it in the in the production

46:22

workload. So they they grow the same

46:24

pace like this uh AI native companies

46:27

reporting they are growing their AI

46:30

whatever uh consumption equal to their

46:33

IR. So the companies like Revolute,

46:35

they're growing the same exponential uh

46:37

trajectory.

46:38

>> So I always push back on people who

46:41

proclaimed that open source would be a

46:43

credible threat to the largest model

46:44

providers because I said listen the

46:46

biggest enterprises want reliability.

46:48

They want security and most of all they

46:50

want ease. They don't want to be

46:52

tinkering around with all the

46:54

architecture and [ __ ] beneath the

46:56

surface. What you're telling me is

46:59

you're able to be all of that to allow

47:03

them to pipe away from those providers

47:05

and have a cheaper better experience

47:07

because you take away the plumbing.

47:09

Correct.

47:11

>> Yes. But I again I think it's not about

47:14

my point is closed models with open

47:18

source models. It's not about like

47:20

reliable or not reliable. Again the work

47:23

of the companies like Nabio to make

47:25

possible like as you say not think about

47:27

plumbing if you want to use alternative

47:29

models but I think it's about

47:32

capabilities again I think that closed

47:35

source models like frontier models are

47:37

great and they will become even better

47:40

and they will solve so many problems

47:43

that we don't solve yet

47:46

and we we have such a diversity of the

47:49

use cases we want to solve.

47:52

that there will be market for the

47:55

smartest models of the world, the

47:58

fastest models of the world, the in

48:01

between models of the world. Smart

48:04

enough but cheap enough and you as a

48:07

customer will be able to just pick the

48:10

right, you know, the the right source of

48:13

token

48:15

uh for each particular uh task. And back

48:19

to the agentic uh layer point maybe it's

48:24

even won't be the customer kind of task

48:27

to choose the like which model to call

48:30

now it will be the engine that knows uh

48:34

all the capabilities like all the models

48:36

underneath and then uh when you go to

48:41

open AI and you do the research you

48:44

don't think in terms of how many loops

48:48

you want it to make. You don't think in

48:50

terms uh when it should go to level lm

48:53

and when it should go to search. You

48:55

don't think should it now call like

48:59

which prompt to to call. Right? It's

49:02

it's happening. You just you give a task

49:06

there is an engine

49:08

uh the reasoning uh engine that decides

49:12

how to run this task and you got the

49:15

result. So I think that a lot of

49:17

enterprise cases a lot of these agentic

49:20

tasks will be solved in the same way

49:22

when it's not you as a developer that

49:25

focusing on customer need will need to

49:28

kind of orchestrate all these tokens and

49:30

models and then we will need all the

49:33

models the smartest one for the most

49:36

complex kind of intelligence

49:40

and the fast models that can do like

49:44

quick iterations.

49:45

And again we don't speak even about all

49:47

the modalities and like what we'll need

49:50

in the physical AI world and so on. So I

49:54

think again my point we will have enough

49:56

of PI for

49:59

different models and what we need to do

50:03

as a as a infrastructure company uh is

50:07

just help for extent we can to make

50:12

developers comfortable how they use all

50:15

these opport all these capabilities that

50:17

models provides because as you as you

50:21

rightly said it's not about like model

50:24

capabilities. It's uh it's not only

50:26

about model capabilities. It's about

50:28

like not plumbing, getting them working,

50:31

getting them optimized, getting them

50:34

reliable. When we look at the explosion

50:36

of models and the specialization of

50:37

models like you said there and how many

50:40

will be built and the depth across

50:42

different use cases

50:44

sadly the one thing that is quite clear

50:47

is that Europe does not have anywhere

50:49

near the model buildout that we've seen

50:51

both in the US and in China. How

50:54

important do you think it is that

50:56

nations have their own sovereign models?

50:59

It looks like the world is divided or uh

51:04

we we can we we may not like it. H and I

51:08

think that having good enough

51:10

foundational models

51:13

uh available for the big parts of the

51:16

world is important. And I think here in

51:19

Europe uh we or at least at this part of

51:24

the world we should think uh how we

51:28

have enough capabilities available

51:32

uh here and I think that we had a lot of

51:35

conversations over the last couple of

51:37

years in like about the serverity and so

51:40

on all this like sovereign AI agenda and

51:44

I think it was too much concentrated

51:47

around like again megawatts and and

51:50

power rather than on what we have on the

51:55

build builder layer right and I think

51:59

that megawatt will come uh I think that

52:03

it's it's uh what what we in Nebios

52:06

always told is we will build

52:07

infrastructure the companies like us

52:09

will build infrastructure if we have

52:10

demand and demand is coming from the

52:12

builders and I think that uh what we

52:17

need to care about here is to have more

52:20

great companies like lovables, black

52:23

forest labs, I don't know, mist drives

52:26

of the world and we have enough people

52:28

that invest in research, have enough

52:30

people that invest in the products uh

52:33

and then they will create enough of

52:35

demand and there will be enough of

52:37

flywheel again to have a good enough

52:39

models if we need. So I think this is

52:41

something that like we should care

52:43

about. Where is the most interesting

52:45

area to invest today? Okay, I'm giving

52:48

you four options.

52:50

Infrastructure,

52:52

horizontal model, vertical model,

52:55

application layer.

52:58

Uh I mean we built infrastructure. So uh

53:02

uh uh we are quite happy here. I think

53:04

it's a good place to be in the current

53:06

world. I think that uh even though like

53:10

we we for some extent we are building

53:13

kind of the easiest part not in a way uh

53:16

it's complex execution but we kind of

53:20

know what's needed and our customers

53:22

help us to understand what's needed. I

53:25

think the most amazing people in this

53:27

industry are those who take a risk to go

53:29

and build uh enduser products in my view

53:34

and they actually drive the most of uh

53:37

uh a most of growth here like people who

53:41

take a risk like the real like the real

53:43

risk of building something people would

53:45

need or not need. Uh I think this is the

53:48

most the heroes uh of our like AI

53:52

journey. Speaking of heroes of AI

53:54

journeys, before I do a show, I go and

53:56

speak to I'm very fortunate. Now, you

53:58

mentioned earlier I've interviewed some

54:00

some big people. I go and speak to some

54:02

of those big people. A theme that did

54:04

come up when I was speaking to them was

54:06

the relationship with Nvidia. And is a

54:09

marriage a marriage if one has more

54:11

power than the other? How do you think

54:14

about the the power dynamics in a

54:17

relationship with Nvidia when they have

54:19

so much power? We look at this in a very

54:22

simple manner. We just need to build

54:24

what we build. Uh we need to build uh

54:27

our product. We need to tell our story

54:30

and then uh the rest will complement it.

54:34

Uh I think what is the most fascinating

54:38

uh Nvidia is still for big extent is an

54:42

engineers driven company and I think the

54:45

best thing you can do to get respect

54:48

from Nvidia it's my read uh they may

54:52

have a different uh uh point of view but

54:56

if engineers in Nvidia respect uh your

54:59

engineers you will have the right

55:02

foundation for relations let's Okay. And

55:05

I think that uh we managed to prove uh

55:09

again and again that we

55:14

know what we build and we have a strong

55:16

engineering team and I think that they

55:20

see it and they respect it and we have a

55:23

lot of like engineers to engineers

55:24

relations on physical like on a on a

55:27

hardware level on the software layer on

55:29

the inference platform layer and the

55:32

better engineers in Nvidia think about

55:35

you, the better

55:38

uh relations and partnership uh I think

55:41

it's enables and uh and again we may be

55:45

maybe wrong thinking this way but uh but

55:49

but but that that's what we see like we

55:52

can do and uh we we we just focus on

55:56

being reasonable and being kind of

55:59

focused on the long-term value. It

56:02

sounds like fluffy. Everybody say it.

56:04

But uh just do do your [ __ ] job at

56:09

the at the end of the day, right?

56:11

>> I'm going to title this Roman. Just do

56:13

your [ __ ] job.

56:15

>> No. What what else we can do? I mean,

56:17

it's not we are we we we we are in such

56:20

a race and we just can do we can do the

56:23

best to do our work better. I think

56:25

that's that that Yeah.

56:27

>> Just do your [ __ ] job. I Jose I know

56:30

it's funny. I like it. Huh? But like

56:33

what's the hardest part of just doing

56:35

your [ __ ] job today?

56:37

>> Four dimensions. Uh build scale, build

56:41

product,

56:42

work with customers. It's actually like

56:44

two dimensions. We discussed like scale

56:46

and product. The thought is customers.

56:49

We are in the field business. We cloud

56:52

is the we we like to say that cloud is

56:54

post sales business. When you sell, you

56:56

sell the promise and then the c you need

56:59

to satisfy the customer and working with

57:02

the customers, covering the customers,

57:04

having this strong customer engineer

57:07

like customerf facing engineering team,

57:09

FDE team. This is the third dimension.

57:12

Go talk to your customers. Make sure

57:15

that they know you, that you know them.

57:17

This is the third dimension. And the

57:19

fourth, the most boring but also the

57:22

most exciting is the capital. We are in

57:25

the capital intensive game and we

57:28

competing with the most capitalized

57:31

companies in the world.

57:32

>> If if I gave you unlimited budget, what

57:35

would you do differently?

57:37

>> Uh build faster. That's that's very

57:41

easy.

57:41

>> Build what faster?

57:43

>> Yeah, data centers and fulfill them with

57:45

GPUs. Like just build faster. Our copics

57:49

program this year is 2025 billion. our

57:53

competitors hyperscalers have like 10

57:56

times like eight times bigger. If I

58:00

would have like uh 10 times bigger

58:03

capital, I would just build more data

58:06

centers and fulfill them with GPUs

58:08

faster and uh serve more customers.

58:11

That's what we started with like what

58:12

would I do if I had like 10 times more

58:15

supply? I would have I would move

58:17

faster. Gavin Baker said, I think quite

58:20

intelligently, that permitting and

58:22

regulation and the delayed buildout of

58:25

data centers has actually helped because

58:28

if I enabled you to build 10x the data

58:31

centers today, it would actually create

58:34

the glut. Yeah, it's it's actually a

58:36

great question and uh

58:39

and like our investors sometimes ask us

58:42

like what is the main bottleneck and the

58:44

main bottleneck again it's all it's it's

58:46

it's everything but if you you need to

58:49

look at this from the time time span

58:51

perspective again in the six in the next

58:54

6 months the capital cannot help like

58:57

you 6 months is too short time you you

59:00

have what you have you need to deliver

59:03

then in the next 12 months You can

59:05

accelerate something but again it's more

59:08

like capacity constraints and in the

59:11

next 12 months we can accelerate uh with

59:15

the capital or with execution something

59:17

but but then in 24 months you definitely

59:20

can unlock so many things and you can we

59:23

are not building one data center. It's

59:25

also important to understand we are

59:27

building the portfolio the portfolio of

59:29

capacity and the more execution power we

59:33

have the more capital we have we can do

59:37

the things in parallel we can unlock

59:40

like that's why we do how we do we

59:44

secure power and land then we build data

59:47

centers then we fulfill them with GPUs

59:50

every next stage requires more capital

59:53

but we do as much as possible in advance

59:56

to make sure that when we will be on the

59:58

next stage we already have power secured

1:00:01

when we will have enough capital to

1:00:03

deploy in GPUs we will have data centers

1:00:05

that up and running so it's like phases

1:00:08

of investments and again the bottlenecks

1:00:12

are different on the different time span

1:00:14

perspective so obviously if you have

1:00:16

more capital you can move faster not in

1:00:19

6 months but in whatever 18 24 months

1:00:22

for sure

1:00:23

>> can I ask you when you think about the

1:00:25

the the data center build out there that

1:00:26

we're seeing more and more public angst

1:00:29

towards AI. Eric Schmidt's getting booed

1:00:31

off stage. Um not because of the content

1:00:35

but because of the AI inventions. Um and

1:00:38

we're seeing like public resentment

1:00:40

towards data centers. I think 40 out of

1:00:42

100 now are not being built when they go

1:00:45

through planning and approvals.

1:00:47

How do you think about and reflect on

1:00:49

that internally?

1:00:51

This is the environment we need to work

1:00:53

in. So again there are two sides of the

1:00:57

thing. One is how we think pragmatically

1:01:00

as a business. Uh that's what I said we

1:01:03

we think about it as a portfolio of the

1:01:05

projects. We need to make sure that we

1:01:07

are like overs subscribed if you want

1:01:11

and if one data center will be delayed

1:01:14

we will still deliver enough capacity to

1:01:16

our customers and most of the customers

1:01:20

they are not locked in one physical

1:01:22

location. And they just like it's a

1:01:24

cloud uh we we can build uh in different

1:01:28

places and then bring the workloads

1:01:30

where we have capacity and but this is

1:01:33

the pragmatical side of the things. Then

1:01:36

what we obviously see that communities

1:01:38

and uh the local authorities require the

1:01:42

companies like us to work closely with

1:01:45

them and explain and uh show what what

1:01:50

what we do and work with them on their

1:01:52

concerns and like address them. This is

1:01:55

the reality. I mean uh you can compare

1:01:58

it uh when Uber uh started growing and

1:02:04

in many places there was the push back

1:02:06

right so oh what's happening it's

1:02:08

something new we it's moving too fast we

1:02:11

didn't didn't expect it to move so fast

1:02:13

and so on and I think that you you you

1:02:17

go and work and you explain and it's a

1:02:20

it's a it's just a part of your of your

1:02:23

duty to engage and work with the new

1:02:27

communities that become dependent on you

1:02:30

and they have concerns and sometimes

1:02:33

they just they have concerns because

1:02:35

they not educated enough. Sometimes they

1:02:37

have rational concerns that you can

1:02:39

address and

1:02:42

the same do your job.

1:02:46

>> Do you think you've done a good job at

1:02:48

it so far?

1:02:49

>> We come from the place we we we always

1:02:51

we always think that we didn't do

1:02:53

enough. Uh I think that we we got quite

1:02:57

a progress uh in the in the places where

1:03:00

we when we started building. Um

1:03:05

historically

1:03:07

uh we had more experience in Europe. Uh

1:03:11

we now like probably 70 75% of the new

1:03:16

capacity that we built midterm is in US.

1:03:18

So we built a lot of presence uh on the

1:03:21

ground and in like to communicate with

1:03:24

those local communities in US and we try

1:03:26

to do the best job. Yeah. We we need to

1:03:28

do better always but we we are moving.

1:03:31

>> Can you help me on another one? We

1:03:33

laughed earlier when we said about

1:03:34

space. Data centers on planet earth is a

1:03:37

very difficult logistical buildout. Data

1:03:40

centers in space. I love technology. I'm

1:03:43

an optimist. I hope it

1:03:46

is that [ __ ] nuts.

1:03:49

>> I think everything we see is [ __ ]

1:03:51

nuts.

1:03:53

No, so many smart my my my view is very

1:03:56

simple. Uh so many smart people now

1:03:59

working to make it happen. So most

1:04:02

likely

1:04:04

uh I I may be less pessimistic that

1:04:07

we'll see I I don't know what is there

1:04:09

like we'll build more in space than on

1:04:12

earth in 3 years. my view I I'm humble

1:04:16

enough to say that so many smart people

1:04:18

are trying to solve uh this uh this task

1:04:23

and uh bring compute to the space that

1:04:26

why wouldn't I believe it will happen

1:04:28

and I think there are a lot of

1:04:30

challenges still like a lot of like a

1:04:33

lot of things to figure out

1:04:35

but if someone would said say us that uh

1:04:40

even 3 years ago that we will build like

1:04:42

multi- gigawatt data centers and it it

1:04:45

will be like large interconnected

1:04:47

compute clusters. Would you believe I I

1:04:50

I didn't think like that and it's we are

1:04:53

here it's it's routine.

1:04:55

>> Um I want to do a quick fire with you.

1:04:57

So I say a short statement you give me

1:04:59

your immediate thoughts. What job does

1:05:02

not exist today that you think will be

1:05:05

very common in 5 years time? One thing

1:05:08

that obviously happening is we

1:05:11

democratizing what people like called

1:05:14

being developer right now each of us

1:05:18

can be a developer and like what what I

1:05:22

mean being developer is to convert the

1:05:24

idea in some digital digital asset. So

1:05:28

and I hope that again we have to be

1:05:31

optimist here and I hope that uh this

1:05:35

democratizing of building like letting

1:05:39

each of us being builder will open up so

1:05:42

many opportunities and like that we even

1:05:45

don't imagine yet when we will give like

1:05:48

millions of new people tens of millions

1:05:50

of new people's ability just to convert

1:05:53

their idea into something that works

1:05:55

very easily. we will see a lot of new

1:05:59

businesses and a lot of new ideas kind

1:06:01

of just uh un like coming in life and

1:06:05

they will create a lot of new works that

1:06:07

we don't even think exist. So it's like

1:06:12

second you know uh uh second orital of

1:06:16

uh uh of all this kind of democratizing

1:06:19

of the building also what is challenging

1:06:22

and what will need to be changed and I

1:06:25

think it's like as risky as opportunity

1:06:27

as risk as an opportunity is how the

1:06:29

education will change because

1:06:33

uh now when everybody has access to

1:06:35

intelligence what should people learn

1:06:38

you definitely don't need them to learn

1:06:41

the facts. Everything is available. Like

1:06:44

all the knowledge is kind of available.

1:06:46

Like how do you really like train people

1:06:49

to think when they don't need to think

1:06:52

so much? How to teach people to

1:06:55

continuously change like many

1:06:57

professions will be not stable? How do

1:07:00

you how do you help people to find

1:07:02

themselves in the changing environment

1:07:05

and actually like

1:07:08

think and learn the new concepts

1:07:12

constantly. I think this is this is

1:07:14

something that very like a lot of gives

1:07:17

a lot of new opportunities but also like

1:07:19

creates a lot of risks.

1:07:20

>> You you mentioned you know you have um

1:07:22

two teenage uh daughters. Uh what do you

1:07:25

advise them that they're entering the

1:07:27

workforce in the next 10 years. What do

1:07:29

you advise them?

1:07:30

>> No, I what I literally tell them is I

1:07:34

think two things will be needed. I don't

1:07:36

know what will be needed, but I'm sure

1:07:37

that two things will be needed. One is

1:07:40

like being able to communicate with the

1:07:43

people with empathy with emphatic

1:07:45

communications. So like understand

1:07:47

humans like communicate with humans and

1:07:50

being empathic. And the second is uh uh

1:07:53

creativity like uh uh all the the art. I

1:07:58

hope that the art in in in a way will be

1:08:01

will exist. So I think that all the hard

1:08:04

skills that I thought 10 years ago will

1:08:06

be needed when I thought that the most

1:08:08

important thing they need to learn is

1:08:10

math and uh engineering now I'm far from

1:08:13

this belief and I'm quite happy they

1:08:16

much more in the soft skills than than I

1:08:20

was when I was a kid and uh I again like

1:08:24

understand like being able to

1:08:26

communicate with humans understand the

1:08:28

humans and be emphatic to the humans and

1:08:31

have this creativity uh idea like being

1:08:34

able to try new things and like be

1:08:36

creative. I think this two if you can

1:08:39

help your kids to develop those uh I I

1:08:43

think they will in 10 years they will be

1:08:45

in demand.

1:08:46

>> There's a question of how do you teach

1:08:47

creativity um but I completely agree

1:08:50

with you. The big finish this sentence.

1:08:53

The biggest threat to Nabius is not

1:08:56

competition but dot dot dot

1:08:59

>> uh but consolidation in general. Yeah. I

1:09:04

think that the main threat for Nebios as

1:09:06

a business is the world will be too much

1:09:08

consolidated again like like we

1:09:10

discussed we try to be diversified like

1:09:12

we try to solve like problems of

1:09:15

different customers and have different

1:09:17

customers on different layers. If you'll

1:09:19

end up in the world where I don't know

1:09:21

three five superm models, super

1:09:23

companies, super empires control the

1:09:25

world then nebios or companies like

1:09:27

nebios will be needed only to help them

1:09:29

maybe serve their needs on physical

1:09:32

layer. Um so I think that the in general

1:09:36

the consolidation is our main threat.

1:09:39

The the more world democratized the more

1:09:42

world diversified the more we need as a

1:09:45

business. Do you think that's likely?

1:09:48

We're seeing the we're seeing the

1:09:49

concentration of value to fewer and

1:09:51

fewer players. We're seeing the opposite

1:09:53

of diversification.

1:09:55

>> I hope it will not happen. As a

1:09:57

business, I think that it's better for

1:09:59

us as a humans as well. Uh for you and

1:10:02

me, the world will be become like remain

1:10:06

uh quite diversified in a different

1:10:08

manners. And uh uh I'm optimistic here.

1:10:12

I think that there are so many people

1:10:14

that

1:10:16

want to build something indep

1:10:19

independently

1:10:21

let's say like there is a lot of people

1:10:23

with the need to try things and build

1:10:27

new things that it's organically creates

1:10:31

this pressure and organically creates

1:10:34

more diversified world so hopefully will

1:10:37

will remain

1:10:38

>> penultimate one Leo Ashen Brener is a a

1:10:41

famous investor right now has huge cult

1:10:44

following. Um he recently disclosed a

1:10:47

very large position for him. 5.3% of the

1:10:51

company I think it's 15% of his

1:10:53

portfolio.

1:10:54

>> How do you guys sit internally? Are you

1:10:56

like yeah go Leo? I wouldn't say that we

1:11:00

didn't me like notice it obviously like

1:11:04

everybody noticed it and like the the

1:11:06

the the stock jumped and like it was a

1:11:08

big news in the around. Uh again I think

1:11:12

that we take it as a justification of

1:11:16

what we do. Uh and then you you got this

1:11:20

justification you say yourself okay

1:11:22

those people they give you a credit

1:11:25

that you will execute. It's uh uh I I I

1:11:29

I come back again and again to what we

1:11:32

do is post sale business. Every time we

1:11:35

sign a deal, every time someone invests

1:11:37

in us, they give us a credit and

1:11:40

opportunity to deliver. Then go back to

1:11:43

your job and deliver. And I think that

1:11:46

we are in a such a market where

1:11:49

emotional market as well that you should

1:11:53

keep keep yourself like down to the

1:11:57

ground. Remember that all this growth uh

1:12:00

all these credits that customers give

1:12:03

you. It's opportunity to deliver. Go to

1:12:06

do your job. Uh at

1:12:08

>> you're such an Israeli. Americans would

1:12:11

be like yeah go. You're like

1:12:14

I I think I'm Russian in this way. Like

1:12:17

Russians always know that uh things like

1:12:20

you need to you need to look in the uh

1:12:23

very pragmatically and uh you know

1:12:26

Russians always with this like faces

1:12:29

like always expect something will happen

1:12:31

and you you need to be you need to be

1:12:34

ready uh you need to be ready. So no I I

1:12:37

I I I I I I think that uh it's really

1:12:41

important part that comes uh from our

1:12:43

CEO also uh uh and founder Ki you wake

1:12:47

up and it's it's a new customer new day

1:12:51

you need to deliver nothing is

1:12:53

guaranteed just you need to you need to

1:12:55

concentrate on the work and I know how

1:12:57

much effort team is putting on things to

1:13:01

work and how much depends on every day's

1:13:06

dedication and how much how fast market

1:13:10

is moving and to stay relevant you need

1:13:13

to continue moving in the same pace or

1:13:15

try to move in the same pace with the

1:13:18

market and uh again you I think that on

1:13:21

a romantic note I I would say that we

1:13:24

could celebrate a little bit more but we

1:13:27

just don't have time to use opportunity

1:13:29

actually to say kudos to the team I I

1:13:31

don't think we celebrate enough and I I

1:13:33

I think that We I think it's right we

1:13:36

are not relaxed but I think we could

1:13:38

celebrate a little bit more uh and just

1:13:42

give the team like more uh more respect

1:13:46

and like how much uh how much is done

1:13:50

and it was not easy and it's still not

1:13:52

easy uh and it will not be easy but yeah

1:13:56

never stop uh we we cannot stop like you

1:13:58

you it's like

1:14:01

you it's like a shark you're alive when

1:14:03

you move Right. So, uh, this famous

1:14:06

thing. So, we have to move.

1:14:09

>> On that note, I cannot thank you enough

1:14:11

for joining me and for putting up with

1:14:12

my very meandering questions. You've

1:14:14

been fantastic, Roman. So, really huge

1:14:16

thank you.

1:14:17

>> Thank you. And, uh, too kind to me.

Continue with YouTLDR

Analyze another video with Pro

Process a new video, search every timestamp, compare sources, and keep the result in your library.

Get Pro — $12/month30-day money-back guarantee

More transcripts

Explore other videos transcribed with YouTLDR.