Capital War Series Issue #5 Nvidia and the Toll Road of Intelligence
From Selling GPUs to Defining the AI Factory
How CUDA, Networking, and Rack-Scale Systems Turned a GPU Company Into AI’s Control Plane
In July 2026, Nvidia and a coalition of Japanese industrial companies announced plans for a 140-megawatt AI factory built around 27,500 Rubin GPUs, 13,750 Vera CPUs, Vera Rubin NVL72 racks, Spectrum-X networking, and Nvidia’s DSX data-center architecture. The project — developed with Noetra Corp., a consortium backed by SoftBank, NEC, Sony, and Honda — will support Japan’s government-backed FRONTia initiative to build foundation models for robotics, manufacturing, logistics, and other forms of physical AI. Nvidia is calling it the world’s first national infrastructure project of its kind.
Japan isn’t simply buying chips; it’s buying an entire architecture, and that distinction explains how Nvidia became the central company of the AI capital cycle. The GPU made Nvidia indispensable, but its position no longer rests on the GPU alone. It extends through CUDA, optimized software libraries, high-speed interconnects, Ethernet and InfiniBand networking, complete rack-scale systems, deployment software, and increasingly the operational design of the data center itself.
In Issue #1, we argued that AI is not primarily a software race — it is a contest over the physical infrastructure that intelligence depends on. Issues #2 and #3 examined two constraints below the accelerator: high-bandwidth memory and advanced semiconductor fabrication. Issue #4 moved outward to the power systems required to operate the machines. This issue asks who has captured the greatest economic value from that stack.
The answer is Nvidia — not because it owns every critical input, but because it has become the company that coordinates them. Nvidia no longer sells a component; it sells the fastest, safest route from capital expenditure to working intelligence. That route has become the toll road of AI.
The GPU Was the Entrance
The GPU remains the visible center of Nvidia’s position. Modern AI models require enormous volumes of parallel computation, and graphics processors were unusually well suited to performing it. But raw processing power was never sufficient. A useful AI accelerator must be supplied with data from high-bandwidth memory. It must communicate with hundreds or thousands of other accelerators. Its workloads must be divided across those processors without losing too much time to data movement. Compilers, drivers, libraries, frameworks, and orchestration software must all work together, and the complete cluster must remain operational despite the failure of individual components.
The value of a GPU therefore depends on the system around it, and Nvidia understood this earlier than most of the industry. CUDA, introduced in 2006, allowed developers to use Nvidia GPUs for general-purpose parallel computing. Over time, the company added mathematical libraries, deep-learning primitives, compilers, debuggers, performance tools, model-serving software, and support across scientific and industrial applications.
The GPU created the installed base; software increased its value; more developers produced more applications, which created more demand for hardware, which gave Nvidia more resources to improve the software again. The resulting advantage compounds. A competitor can design a processor with attractive specifications, but it is far harder to reproduce two decades of code, tools, documentation, optimized libraries, trained engineers, and production experience. The chip was the entrance to Nvidia’s platform; CUDA is the road behind it.
CUDA and the Cost of Leaving
CUDA is often described as Nvidia’s moat, but that description can mislead. CUDA is not primarily a licensing business — Nvidia does not collect a royalty each time a developer executes a CUDA instruction, much of its developer software is distributed freely, and the leading AI frameworks sit above the hardware layer. Its economic function is different: it makes Nvidia hardware easier to use, and harder to replace.
A company that has built its infrastructure around CUDA has accumulated more than source code. Its engineers understand Nvidia’s tools. Its models have been tested on Nvidia systems. Its deployment pipelines rely on Nvidia libraries. Its performance assumptions, monitoring, and debugging procedures all reflect years of production experience on Nvidia hardware. Leaving that environment requires code migration, performance tuning, model validation, infrastructure changes, and employee retraining — so a rival accelerator does not merely have to be cheaper. It must be cheaper by enough to offset the cost and risk of switching.
That premium matters most at the frontier, where models and workloads change quickly and general-purpose GPUs let researchers modify architectures, precision formats, kernels, and training methods without waiting on a new purpose-built chip. The moat, in other words, is not a single proprietary technology — it is the accumulated work customers would have to repeat somewhere else.
Nvidia says more than half of its engineers now work on software, a striking allocation for a company still commonly described as a chipmaker. Its most durable product may not be any particular GPU generation. It may be the expectation that software written for Nvidia today will remain useful on Nvidia’s next system.
Networking Became the Second Tollbooth
As AI systems grew, the relevant unit of computation changed. The performance of one GPU still mattered, but the performance of the cluster began to matter more, because thousands of accelerators had to exchange model parameters and intermediate results at extremely high speed — any delay in communication left expensive processors waiting instead of computing.
Nvidia’s 2020 acquisition of Mellanox gave the company control over a critical part of that problem. NVLink connects GPUs inside tightly coupled compute domains; NVLink Switch extends those connections across larger configurations; InfiniBand and Spectrum-X Ethernet connect racks into clusters; ConnectX network adapters and BlueField data-processing units manage the movement, isolation, and processing of data across the system. Together these let Nvidia optimize compute and communication as a single problem, which matters because the customer isn’t ultimately buying theoretical processor performance — they’re buying completed training runs and generated tokens. A cheaper accelerator can become the more expensive system if the cluster is difficult to deploy, achieves lower utilization, or requires more engineering work to produce the same output.
Nvidia’s latest financial results show how important this layer has become. In the quarter ended April 26, 2026, data-center networking revenue reached a record $14.8 billion — up 199% from a year earlier and 35% from the previous quarter — while data-center compute produced $60.4 billion. Networking is no longer an accessory attached to the GPU business; it is a major business in its own right. The Mellanox acquisition did more than add a product category. It moved the boundary of Nvidia’s platform from the server to the data center.
From Chip to Rack
Blackwell moved that boundary again. The GB200 NVL72 combined 72 Blackwell GPUs and 36 Grace CPUs inside a liquid-cooled rack connected by NVLink. Blackwell Ultra extended the system, and Vera Rubin now integrates Rubin GPUs, Vera CPUs, ConnectX networking, BlueField DPUs, NVLink 6, Spectrum-X Ethernet, and Nvidia’s software environment into a single unit. The product is no longer a processor installed inside someone else’s machine — increasingly, it is the machine itself.
This changes both the technology and the economics. Building a large AI cluster requires thousands of decisions about power delivery, cooling, cabling, networking, storage, component compatibility, workload scheduling, and physical layout, and each decision introduces potential delays or failures. By offering validated rack designs and data-center reference architectures, Nvidia converts that coordination problem into something closer to a standardized purchase — and customers pay a premium because the system promises to shrink the time between committing capital and producing intelligence.
The Japan project illustrates the result. The planned AI factory will use Nvidia’s processors, racks, networking, and DSX architecture as a common foundation; Nvidia is effectively supplying the blueprint through which a national industrial strategy will be executed. This is the deeper meaning of the “AI factory” language Nvidia now uses: the company wants to define not only the machines performing the computation but the architecture of the factory surrounding them. The more of that architecture Nvidia controls, the larger its share of the customer’s capital budget — and the harder it becomes to replace any single layer without replacing all of them.
The Economics of Control
Nvidia’s financial results increasingly resemble those of a platform operating inside a capital-intensive industry. Nvidia generated $81.6 billion of revenue in its fiscal first quarter of 2027, up 85% from the previous year, with data-center revenue reaching $75.2 billion — more than 90% of the company’s total — at a GAAP gross margin of 74.9%. For the following quarter, Nvidia forecast $91 billion of revenue while assuming no data-center compute revenue from China at all. Those figures show a company extracting software-like margins from the largest physical infrastructure buildout of the current cycle.
But the margin story isn’t one-directional, and it’s worth sitting with the exception rather than skipping past it. For fiscal 2026, Nvidia’s gross margin fell from 75.0% to 71.1%, driven partly by a $4.5 billion charge tied to H20 inventory and purchase obligations after U.S. export restrictions upended the market for that China-focused product, and partly by the company’s shift from Hopper HGX offerings toward more complete Blackwell data-center systems. That second cause is the more durable one. Selling complete systems expands Nvidia’s revenue opportunity, but it also transfers more component cost, integration work, inventory risk, and manufacturing exposure onto Nvidia’s own balance sheet — costs the company didn’t carry when it was selling a chip into someone else’s design. A dollar of rack revenue and a dollar of accelerator revenue are not the same dollar. The latest quarterly margin has climbed back to roughly 75%, but the earlier dip is a preview of what happens whenever a product transition, a supply shock, or an export rule lands: the wider the toll road becomes, the more of the road Nvidia itself is responsible for maintaining. That is the trade Nvidia has chosen — more revenue per customer, in exchange for absorbing risks it used to pass upstream.
The Most Powerful Fabless Company
Nvidia controls the architecture, but not the industrial foundation beneath it. The company relies on outside foundries, including TSMC and Samsung, to manufacture its chips; buys memory from SK Hynix, Micron, and Samsung; depends on sophisticated packaging such as TSMC’s CoWoS; and uses contract manufacturers to assemble and test its boards, servers, and racks. That gives Nvidia a capital-light model — it can direct extraordinary resources toward design, software, and ecosystem development without spending tens of billions of dollars building its own leading-edge fabs.
It also makes Nvidia dependent on nearly every chokepoint covered in the previous issues of this series. A shortage of HBM can restrict shipments. A packaging bottleneck can delay entire systems. A fabrication problem can disrupt a product cycle. Power and cooling constraints can prevent customers from installing equipment they’ve already bought. Nvidia is therefore both the most powerful company in the AI stack and one of its most dependent — its advantage comes from coordinating the stack, not from owning it.
The Detours Are Being Built
Every toll road eventually creates an incentive to find another route, and Nvidia’s largest customers are also the companies with the strongest reason — and the greatest financial ability — to reduce their dependence on it. Google has spent years developing TPUs. Amazon is expanding Trainium and the Neuron software stack. Microsoft is developing its own accelerators. AMD is building complete rack-scale systems around its Instinct GPUs, EPYC CPUs, Pensando networking, and ROCm software.
Meta is now accelerating its own effort. According to a July 2026 internal memo reported by Reuters, the company plans to begin manufacturing an in-house AI chip known as Iris in September, part of a four-generation Meta Training and Inference Accelerator roadmap that will add a new chip roughly every six months through 2027. Meta expects to operate seven gigawatts of computing infrastructure in 2026 and fourteen in 2027; at that scale, even a narrow custom accelerator produces real savings. Meta is designing the chip with Broadcom and manufacturing it through TSMC. Iris isn’t meant to eliminate Meta’s Nvidia purchases — it’s designed to supplement them. But partial substitution is enough to matter.
Nvidia’s moat is strongest when customers need flexibility: frontier training, rapidly changing model architectures, scientific computing, and workloads that require a broad software ecosystem. It’s more vulnerable when workloads become stable, repetitive, and large enough to justify specialized silicon. Inference is therefore both Nvidia’s largest opportunity and its most likely point of pressure. As AI applications generate more tokens, the cost of serving them becomes increasingly important, and hyperscalers can move predictable internal workloads to proprietary accelerators while reserving Nvidia systems for the most complex or flexible tasks. The toll road doesn’t have to be abandoned for traffic to shift at the edges.
The Geopolitical Bypass
China represents a different kind of detour — one Washington is now managing in pieces rather than closing off entirely. In July, a U.S. official confirmed that a small number of Nvidia H200 processors had begun shipping to approved Chinese customers, with several companies — including a unit connected to ZTE, alongside Alibaba, Tencent, and ByteDance — cleared to purchase the chips, though actual shipments remain limited. Read against the FY2026 H20 charge, the pattern is telling: Washington shut off one China-facing product abruptly, at a cost to Nvidia of $4.5 billion in a single quarter, and is now reopening a narrower channel through a different one, calibrated and reversible. That’s a more manageable risk for Nvidia than a blanket ban, but it’s also a risk Nvidia doesn’t control — the next policy shift could tighten the H200 channel as easily as it opened it, and Nvidia’s current guidance already assumes zero China compute revenue as the safer planning baseline.
Meanwhile, Huawei unveiled its Atlas 950 SuperPoD, designed to connect thousands of domestic Ascend processors through high-speed interconnects, and DeepSeek’s V4 model has reportedly been adapted to run entirely on Huawei-based clusters. U.S. controls may limit China’s access to Nvidia in the short term, but they also give China a strategic reason to finance an alternative hardware and software ecosystem that could eventually compete outside its domestic market too. For most companies, access to Nvidia is a commercial decision. For entire countries, it is increasingly a geopolitical one — and geopolitical decisions don’t reverse as easily as commercial ones once an alternative ecosystem is built and running.
The Open Question
Nvidia’s moat was never that no one could design another accelerator. It’s that customers must replace an entire working system — compute, software, libraries, networking, deployment tools, trained engineers, production experience — and the alternative has to be cheaper or better by enough to justify the cost of leaving. That remains an exceptionally high bar. It is not, however, a permanent one. Open software can reduce migration costs. Custom silicon can absorb standardized workloads. Competitors can adopt rack-scale architectures of their own. Governments can subsidize domestic ecosystems. And Nvidia’s own prices and margins give its largest customers a constant, compounding reason to keep trying.
Nvidia will almost certainly lose some accelerator share as the market expands and fragments; that much is already visible in Meta’s Iris roadmap and Huawei’s SuperPoD push. The harder question is whether Nvidia can keep redefining the architecture of AI systems — GPU to networking, networking to racks, racks to national infrastructure — faster than the companies paying its tolls can finish building their own way around it. Right now, Nvidia is still winning that race by widening the road before its customers can finish the detour. Whether it can keep doing that for another five years, against customers with Meta’s balance sheet and countries with China’s patience, is the question this series will keep coming back to.
