Intelligent Founder AI
Intelligent Founder AI Podcast
Ep.018 - Nvidia and the Compute Wars: What It Means If You Are Building Right Now
0:00
-13:01

Ep.018 - Nvidia and the Compute Wars: What It Means If You Are Building Right Now

The 9-episode Build/Buy/Rent Infrastructure Series Lessons and What choices founders need to make next.

Over the past eight episodes, we set out to make one idea practical:

building with AI is no longer just about choosing a model or writing a prompt.

It is about making a sequence of technical, commercial and operational decisions that determine whether an AI product becomes a durable business.

We explored the questions founders now have to answer in real time:

where AI genuinely creates an advantage,

how to choose between building and buying, how to think about agents and automation,

what it takes to move from prototype to production, and

why infrastructure choices quietly become product strategy.

We did not approach these topics as a catalogue of tools. We approached them as decisions and as trade-offs.

This is the final episode of this first run. It closes where many AI businesses ultimately arrive: at compute.

Intelligent Founder AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

Compute can sound like a technical footnote.

In practice, it shapes model choice, unit economics, latency, deployment geography, data governance, reliability, fundraising narratives and the speed at which a team can experiment. If the previous episodes were about deciding what to build with AI and how to turn it into a useful product, this episode is about the physical and economic system underneath those choices.

What changed?

When I first prepared this episode, the story looked relatively straightforward. Nvidia’s data-centre revenue had reached $51.2 billion in a single quarter, Blackwell demand appeared intense, and GPU rental pricing seemed likely to become steadily cheaper as more capacity arrived.

The first observation has become even more significant.

Nvidia subsequently reported $75.2 billion in data-centre revenue for its fiscal first quarter of 2027, up 21% quarter on quarter and 92% year on year. This is not simply a company earnings statistic. It shows the scale at which accelerated compute has become foundational infrastructure for the AI economy.

The second observation also holds:

demand for advanced AI systems remains very strong. But the useful founder lesson is not to treat any single backlog figure or “sold out” headline as a permanent fact. Capacity differs by configuration, region, customer relationship, commitment length and workload profile.

Supply chains change.

customer priorities change;

a chip that is scarce for frontier training may be entirely unnecessary for a production inference workload.

The third observation needs more nuance.

Compute has not followed a smooth, one-way path towards lower prices.

Older-generation hardware, marketplace supply, reserved commitments and specialist clouds can create excellent economics. At the same time, premium capacity can tighten quickly when large training runs, agentic products or new model releases absorb supply. The relevant question is therefore not, “Will GPU prices fall?” It is, “What is the least expensive reliable way to meet our specific quality, latency, privacy and volume requirements?”

Share

The real Nvidia moat.

Nvidia’s position is about more than high-performance GPUs. It is a full operating environment: CUDA, libraries, frameworks, optimized kernels, deployment tooling, networking, systems expertise and a deep global base of engineers who already know how to use it.

That matters because switching compute platforms is not the same as changing a cloud region. A model may technically run elsewhere while requiring new optimisation work, different kernels, revised serving infrastructure, a changed observability stack and a fresh performance-validation cycle. The strongest lock-in often lives in engineering time and accumulated operational knowledge, not a contractual restriction.

But dominance does not mean every founder should default to the newest Nvidia hardware. It means Nvidia is the benchmark against which other choices should be evaluated.

Competition is becoming useful

AMD remains the most important broad GPU alternative, while custom silicon from the major cloud platforms has become relevant for teams operating within their ecosystems.

Google TPUs can be compelling for selected training and inference workloads.

AWS Trainium and Inferentia can offer meaningful savings for compatible models, although savings are not automatic: they depend on model support, migration effort, deployment scale and actual utilisation. AWS has published customer examples of around 50% cost reduction after migration to Inferentia, which should be treated as evidence of potential, not a universal pricing promise

The practical consequence is healthy competition, but not effortless portability.

Founders should not build around a theoretical ability to run identically everywhere. They should build practical optionality: the ability to test a second credible path before dependency becomes expensive.

That means separating application logic from infrastructure-specific code where possible, maintaining reproducible evaluation suites, recording real cost and latency data, and avoiding optimizations that only make sense for a single supplier until there is a demonstrated payoff.

Share Intelligent Founder AI

Local AI is real!

I initially referred to “RTX Spark,” which was incorrect. Nvidia DGX Spark, a compact desktop system built around the Grace Blackwell platform, with 128 GB of unified memory is the new product Nvidia is positioning for local AI development and says it can support models up to 200B parameters, with up to one petaFLOP of FP4 AI performance.

those specifications do not translate directly into every model’s real-world inference speed but the broader point stands.

Local AI is becoming a genuine part of the development and deployment landscape. For many founders, its value will not be replacing cloud infrastructure. Its value will be faster private experimentation, lower friction during development, work with sensitive data, edge deployment prototypes and the ability to validate a workflow before committing to recurring cloud spend.

For regulated, industrial or data-sensitive applications, this is especially important. The architecture may not be “cloud versus edge.” It may be local development, cloud-scale training, and targeted edge or on-premises inference in production.

The build-rent-buy staircase

A durable compute strategy is not a single decision. It is a staircase.

At the bottom is rent: use APIs, managed model services, or on-demand compute while testing whether users care. This keeps capital commitment low and lets the team change direction quickly.

The next step is optimize: select smaller or more efficient models, use batching, caching, routing and quantization, and measure whether latency, quality and cost actually meet the product requirement. Most teams can gain more here than by chasing the next hardware release.

After that comes commit: reserved instances, capacity agreements, dedicated infrastructure or a managed deployment partner. Move here only when demand is sufficiently predictable that the economic benefit outweighs the loss of flexibility.

Finally, for a limited set of high-volume, sensitive or latency-critical workloads, there is own or operate: dedicated hardware, on-premises systems or edge deployment. This can be compelling, but only once the operational burden and lifecycle costs are understood.

The mistake is to climb this staircase too early. The opposite mistake is to remain at the expensive, unmeasured rental stage after usage has become stable and large enough to justify change.

What the series concludes.

Across these eight episodes, the recurring message has been simple: the technology moves quickly, but the disciplines of building an enduring business change much more slowly.

  1. Start from a sharp customer problem rather than a model capability.

  2. Test the workflow before scaling the architecture.

  3. Treat data quality, evaluation and reliability as product work.

  4. Be realistic about integration and human adoption.

  5. Measure unit economics early. And preserve enough technical flexibility that a supplier, model or platform change does not force a rewrite of the business.

The compute wars do not remove those responsibilities. They make them more important.

For founders, the opportunity is not to predict which accelerator wins.

It is to build a product that can benefit from competition without being trapped by it. Use the hardware that delivers the required outcome today. Keep one credible alternative visible. Let evidence, and not hype, benchmark headlines or fear of missing out, determine when to change your stack.

AI may or may not company-building easy, but with a clearer view of the choices will make it more possible.

Listen to the full episode here, in Substack app, or Apple, Spotify / youtube.

Thanks for reading Intelligent Founder AI! This post is public so feel free to share it.

Share

New On the Quantopinion and AIU

Quantum Opinion
What Will Make Quantum Computing Useful?
My previous essay began with quantum roadmaps. the plans that banks, technology companies and governments are creating to understand where quantum computing may fit into their future…
Read more

AI Unfiltered
The Internet Doesn't Belong to Humans Anymore!
Elon Musk saw a Cloudflare data point and turned it into a much bigger prediction: that AI agents will soon outnumber humans online, while Starlink could carry more than 90% of global IP traffic…
Read more

Leave a comment

Share

Share Intelligent Founder AI

Discussion about this episode

User's avatar

Ready for more?