The NVIDIA AI Factory: Roadmaps, Risks, and Future Dominance
NVIDIA 360 Update -From GTC 2026, where this essay started to now, and what has changed.
I did this deep dive around GTC 2026 (mid-March), in the immediate wake of Jensen Huang’s post-conference podcast blitz i.e. Ben Thompson, All-In, and a string of long-form interviews that reframed NVIDIA as a vertically integrated “AI factory” company rather than a GPU vendor. The I believe I did another update in april but din’t get a chance to publish then. At that time, the whole thesis seem to have rested on the Vera Rubin platform announcement, the DSX AI factory reference designs, the emerging agentic stack (NemoClaw/OpenClaw), and a handful of “what we missed” additions covering networking dominance, DGX Spark, BioNeMo, Omni-verse, energy/grid strategy, and the telco AI Grid pivot.
Four months on, the essay needed updating in three ways.
First, two full quarters of actual financial results have landed, replacing forecasts with hard numbers.
Second, the roadmap has hardened since, from “expected” to “confirmed and shipping,” with Vera Rubin now in full production. and,
Third, a genuinely new competitive and structural picture that has since emerged: custom silicon, not AMD, is now the more interesting threat, and NVIDIA seems to be restructuring its own reporting to reflect a broader physical-AI and edge ambition than the original report I did captured.
ANd, Few other points were about the circular-financing risk running through NVIDIA’s own investments in its customers, and the antitrust/regulatory pressure building across three continents, that I noticed, so I am adding them all and extending the map into everything that has happened since, including the telco story in full, and giving it bit more contested picture that I didn’t look in to earlier.
Table of contents
What NVIDIA actually delivered since GTC (the numbers)
The roadmap, confirmed: Rubin, Rubin Ultra, Feynman
The competitive picture: who is really challenging NVIDIA
The new businesses nobody was watching (networking, edge, life sciences)
The telco angle in full: partners, holdouts, and the fight over AI-RAN silicon
The quantum and robotics angle: where your own work intersects
Circular deals and the financing risk underneath the boom
Anti-trust and regulatory scrutiny across three continents
What has changed for the worse: China and margin corrections
The founder playbook: five plays and a decision framework
What is likely next
What NVIDIA actually delivered since GTC (the numbers)
The clearest signal that the “AI factory” thesis was not hype is the Q1 FY2027 print, reported May 20 for the quarter ended April 26, 2026. Revenue hit a record $81.6 billion, up 85% year-over-year, with Data Center revenue of $75.2 billion, up 92%.
Net income surged 211% to $58.3 billion, and gross margin held around 74.9%, directly reversing the margin-compression narrative that circulated at GTC time. For full fiscal 2026, revenue reached $215.9 billion (up 65% year-over-year) with Data Center revenue of $193.7 billion, now representing 90% of total revenue, up from 78% two years earlier. The board also authorized an $80 billion buyback and lifted the quarterly dividend fivefold, from $0.01 to $0.25 per share.
Jensen Huang stated at GTC that the company had $1 trillion in committed orders through 2027 and that NVIDIA “is going to be short” on supply, thats a claim now largely borne out by the pace of bookings and the buyback confidence.
The roadmap, confirmed: Rubin, Rubin Ultra, Feynman
At GTC-time, Vera Rubin was an announcement.
It is now a fact on the ground:
NVIDIA confirmed Vera Rubin entered full production at GTC Taipei on June 1, 2026, with partner shipments scheduled for the second half of 2026. The annual cadence has also been made explicit and re-confirmed as recently as July.
Blackwell now,
Vera Rubin H2 2026,
Rubin Ultra H2 2027, and
Feynman in 2028, with
a new Rosa CPU replacing Vera on that later timeline.
Rubin Ultra will use an NVL576 rack configuration (four compute chiplets per GPU, 1TB of HBM4E memory) delivering roughly four times the compute of the Rubin NVL144 generation, paired with Groq LP35 LPUs supporting the NVFP4 data format.
Feynman in 2028 is described as a “qualitative leap” -
that’s introducing 3D die stacking, custom HBM beyond HBM4e, the new Rosa CPU (designed to cut CPU development cycles from four years to two), and co-packaged optics for NVLink switches that will let racks scale to 576 or even 1,152 GPU packages, addressing the physical bandwidth limits of copper interconnects.
This optics shift is arguably the most underrated technical detail in the entire roadmap - it is NVIDIA’s answer to the physical ceiling on scaling further with copper.
The Network Is Now the AI Computer.
Right now, the AI story you see online is all about “agents,” “copilots,” and “one‑person unicorns.” It looks like magic on a laptop screen. But behind every slick demo there’s something much more boring 😑 🥱 yet much more important going on and that is
The competitive picture: who is really challenging NVIDIA
I wanted to look at the competition a bit here.
As of mid-2026 is that AMD is a credible “second source” but not an existential threat, while custom silicon from hyperscalers is the more structurally important story.
NVIDIA holds roughly 80% of the AI accelerator market, with FY2026 Data Center revenue of $193.7 billion against AMD’s estimated $7–8 billion in AI GPU revenue (5–7% share). AMD’s relative position has grown fast - from under 1% to 5–7% in three years, but the absolute gap is widening, not narrowing, because NVIDIA’s base is so much larger.
The more interesting threat however is custom ASICs:
Broadcom’s AI ASIC revenue alone exceeded $20 billion in FY2025, and custom silicon shipments from cloud providers are projected to grow 44.6% in 2026 versus 16.1% GPU shipment growth.
Analysts increasingly describe the market settling into a three-tier structure by 2028:
NVIDIA retaining 60–75%, AMD reaching 10–15%, and custom silicon capturing 15–25%, concentrated mostly in cloud-locked inference workloads. and NVIDIA’s response to this? is to deepen its moat past silicon:
a $2 billion investment in CoreWeave, open-sourcing its Dynamo inference software and Nemotron models, and licensing Groq technology directly into its own roadmap rather than competing against it.
The new businesses nobody was watching!
Several verticals that i thought were footnotes are now material. specially NVIDIA’s networking segment that’s built on the 2020 Mellanox acquisition generated over $31 billion for the full fiscal year and $11 billion in a single quarter (up 263% year-over-year), making Jensen Huang’s claim that NVIDIA is “the largest networking company in the world” defensible rather than promotional.
NVIDIA has also formally restructured its reporting into two market platforms -
Data Center (split into Hyperscale and ACIE: AI Clouds, Industrial, and Enterprise) and
Edge Computing (PCs, robotics, automotive, AI-RAN base stations).
On the consumer/prosumer edge, NVIDIA unveiled RTX Spark at Computex in June, extending the $3,999 DGX Spark “personal AI supercomputer” concept into a mainstream laptop and compact-desktop tier bundling CUDA, RTX, DLSS and TensorRT for “personal AI agents.”
Automotive has also expanded well beyond the original scope, with new partnerships across Hyundai/Kia, Uber’s full-stack autonomous vehicle stack, BYD, Geely, Isuzu, and Nissan on the DRIVE Hyperion platform, alongside a new Halos OS for AI-driven automotive safety.
The telco angle in full: partners, holdouts, and the fight over AI-RAN silicon
The telco strategy is larger and more contested than a single AI Grid announcement suggests, and the partner ecosystem has grown well beyond the six operators (AT&T, T-Mobile, Comcast, Spectrum, Indosat, Maxis) named at GTC.
At Mobile World Congress in early March, NVIDIA secured a formal commitment from a much broader coalition - Booz Allen, BT Group, Cisco, Deutsche Telekom, Ericsson, MITRE, Nokia, SK Telecom, SoftBank, and T-Mobile - to build 6G on open, AI-native platforms, positioning NVIDIA at the center of a genuinely global standards push rather than a US-only story. NVIDIA is also a founding member of the AI-RAN Alliance, which has grown past 130 participating companies, and separately launched the AI-Native Wireless Networks (AI-WIN) project with Booz Allen, Cisco, T-Mobile, MITRE, and the ODC to build an all-American AI-RAN stack aimed specifically at accelerating the path to 6G ahead of other blocs.







