Nvidia posted $96 billion. China is already serving without CUDA

/ Zhipu ran a global hit on Chinese chips the same week Nvidia guided China to zero. OpenAI taped out an inference ASIC. Apple told you to stop counting tokens. The quarter is fine. The customer list is building an off-ramp.
by Hozefa Khety
· 13 min read
For about a week in August, the most used model on OpenRouter did not have a name anyone trusted. It showed up as ox-alpha. Coders treated it like a frontier system that had somehow shown up cheap. Then Zhipu, the Beijing lab behind GLM, claimed it. The model is GLM-5.3-Flash: 320 billion parameters, 18 billion active, open weights, multimodal. Zhipu said every token of that stealth traffic ran on Chinese AI chips.
The same week, Nvidia reported a $96.2 billion quarter. Data center revenue was $89.0 billion. The next-quarter guide is $108 billion. China data-center compute in that guide is zero. Hopper shipments into China last quarter were less than 1% of the data-center line.
Those two facts do not cancel each other. They are the same story from opposite ends. Nvidia is still the company you buy if you are allowed to train a frontier model. Inference, the part that scales with every user query, is leaking. China already left. OpenAI taped out its own serving chip. Apple put a 512GB box on a desk and told people to stop counting tokens. The quarter is not scared. The customer list is building an off-ramp.
The model that topped OpenRouter was not on CUDA
Zhipu's own blog is the document that matters, not the English recaps that turned it into a Nvidia-killer headline. GLM-5.3-Flash is built to be cheap to serve: hybrid sparse and linear attention, fewer active parameters than GLM-5.3, a 1 million token context window. Before the official drop on August 26, Zhipu ran it anonymously as ox-alpha on OpenCode and OpenRouter. It became the most popular model of the week. Quote from the company: all of that traffic was served on Chinese AI chips.

The 100,000-chip figure you will see everywhere is not in that blog. LatePost reported a cluster of more than 100,000 domestic cards and named Huawei, Moore Threads, and Hygon as possible suppliers. Zhipu did not confirm the vendors. Treat 100,000 as reporting. Treat "Chinese chips, real traffic, global users" as the company line.
The engineering admission is more interesting than the patriotism. Zhipu says the cards are short on memory capacity and bandwidth, especially at a 1 million token context. They rewrote the serving stack on SGLang: quantization, layer split, encode-prefill-decode disaggregation, a custom interconnect. Versus their own baseline on the same hardware they claim a 3x end-to-end gain, then "hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs." That is a serving-economics claim, not a die that beats Rubin. Nobody published an independent harness against a named Nvidia SKU. It still counts. A frontier-adjacent open model, used globally, did not need CUDA once the software was ugly enough.
This was also not a first try that magically worked. The GLM-5 paper in February said the model was adapted from day one across Huawei Ascend, Cambricon, Moore Threads, Kunlunxin, MetaX, Enflame, and Hygon. GLM-5.2's launch post added T-Head, Biren, and Iluvatar and said online inference already sat on those platforms. GLM-5.3-Flash is the first time they put large-scale traffic on a large domestic cluster. Read that sentence the way a lawyer would. Earlier GLM serving was mixed, smaller, or not at this QPS. The stack crossed a production threshold in August. It did not teleport there.
Washington opened the door. Beijing closed it
The usual export-control morality play is backwards this year. In January the U.S. Commerce Department moved H200 sales to China onto case-by-case licenses, naming the chip in the Federal Register. Nvidia's 10-Q then said the licenses covered small amounts and, as of the April quarter, produced zero H200 revenue. In April, Commerce Secretary Howard Lutnick told the Senate the quiet part: "The Chinese central government has not let them, as of yet, buy the chips, because they're trying to keep their investment focused on their own domestic industry."
Reuters reported in May that Washington had cleared about ten Chinese firms, including Alibaba, Tencent, ByteDance, and JD.com, and that not a single delivery had gone out. Chinese buyers pulled back after guidance from Beijing. Jensen Huang even made the Trump-Xi trip. The chips still did not move.
By the quarter ended July 26, a trickle of Hopper product finally showed up as less than 1% of data-center revenue. Nvidia's written guide for the next quarter still assumes no China data-center compute at all. Bloomberg, on earnings day, said the company told investors it could not sell the full licensed H200 allotment because of objections from Beijing. That is not a U.S. ban. It is industrial policy. Starve the import. Feed Huawei. Let ByteDance have a dribble so the labs do not revolt.
Nvidia has already said this in an SEC filing, in language no columnist would get away with. The company wrote that it was "effectively foreclosed from competing in China's data center computing/compute market," and that this foreclosure "helped our competitors build larger developer and customer ecosystems to challenge us worldwide." Pair that sentence with GLM-5.3-Flash. The off-ramp in China is not a 2027 risk. It is procurement.
OpenAI built a chip for ChatGPT, not for GPT training
If China is the finished version of leaving Nvidia, Jalapeño is the American lab version. On June 24, OpenAI and Broadcom unveiled OpenAI's first "Intelligence Processor." It is an inference ASIC. Blank-slate for LLM serving, not a reused training part. Broadcom does the silicon and Tomahawk networking. Celestica does boards and racks. Engineering samples are already running GPT-5.3-Codex-Spark in the lab at production frequency and power. OpenAI says early testing shows performance per watt "substantially better than current state-of-the-art," with a full technical report still to come. Hock Tan said gigawatt-scale deployments with Microsoft and other partners begin in 2026. Initial deployment is targeted by the end of this year.

The October 2025 collaboration behind it is 10 gigawatts of OpenAI-designed accelerators through 2029. That is a real infrastructure program. It is not ChatGPT running on Jalapeño today. Training stays on GPUs. OpenAI is also bringing AMD Helios online in the fourth quarter as the first phase of a 6 gigawatt AMD deal. Nvidia is not being dumped. It is being dual-homed, then triple-homed, so the next trillion decode tokens have somewhere else to land.
The detail that should bother CUDA loyalists is speed. OpenAI and Broadcom say the chip went from design to tape-out in nine months, helped along by OpenAI's own models. The old defense of merchant GPUs was that ASICs take five years and miss the next architecture. If a lab can retape that fast, the defense shrinks to "prove it in a rack that stays up." Fair. Jalapeño still has to survive yield, software, and a year of production. A Hot Chips talk is not a fleet.
The people who made Nvidia rich are shopping
OpenAI is late to a party Google has been throwing for a decade. Gemini already trains and serves on TPUs. Google's seventh-gen Ironwood was positioned as an inference-era TPU. In April it announced an eighth generation split in two: TPU 8t for training, TPU 8i for agents and serving, claiming up to twice the performance per watt of Ironwood.
Anthropic is the cleanest Western paragraph. More than a million Trainium2 chips are already in use to train and serve Claude, with 5 more gigawatts of Amazon compute behind that, plus a path to a million Google TPUs. In August Anthropic confirmed an in-house silicon team. It also said it will keep buying Trainium, TPUs, Nvidia, and AMD. That is not a breakup. That is a lab that refuses to let one vendor set the token price.
Microsoft's Maia 200 is live in Iowa and Arizona and, Microsoft says, serving OpenAI and its own MAI models. Meta already runs hundreds of thousands of MTIA chips in production for feed and ads ranking. The GenAI-inference follow-ons, MTIA 450 and 500, are aimed at 2027. Do not write that Llama is off Nvidia today. Do write that the boring inference, the stuff that actually prints money for Meta, already left.
AMD is the only merchant GPU that looks like a second source rather than a science project. In July, Microsoft said Azure will ramp Helios racks, 72 MI455X GPUs plus EPYC Venice, at scale for frontier-model inference in the second half of 2026. OpenAI, Anthropic, Meta, and Oracle are on the named list. AMD's tokens-per-dollar claims are AMD lab estimates. The purchase orders are not.
Nvidia's tell is that it bought an inference architecture
None of this has shown up as a down quarter. $96.2 billion in revenue, 75% gross margin, $279 billion in supply commitments, mostly memory for Vera Rubin. Blackwell Ultra still filled the last three months. Rubin is in production at CoreWeave, Google Cloud, Azure, Oracle, and Nebius. If you only read the income statement, the off-ramp is a blog post.

Then look at what Nvidia did to its own roadmap. In December it took a non-exclusive license to Groq's inference IP and hired Jonathan Ross and most of Groq's engineers, a deal widely reported around $20 billion and structured to skip a classic merger filing. Groq Inc. still runs a cloud. On August 24, Nvidia said Groq 3 LPX is in full production: 256 language-processing units in a rack that bolts onto Vera Rubin NVL72 and exists to make decode fast. Nvidia killed its own Rubin CPX, a GDDR7 long-context card, to ship this instead. When a monopolist cancels its chip to wear someone else's, it has seen the hole.
NVLink Fusion is the other tell. It lets a hyperscaler's custom accelerator sit on Nvidia's scale-up fabric. Translation: if you must build your own die, Nvidia still wants the rack, the switch, and the software. CUDA, NIM, and the fact that every new architecture still pretrains on Nvidia first are the remaining moat. First-generation ASICs have a long record of missing software, missing flexibility, and missing the next model. That history is real. So is a million Trainium2 chips serving Claude.
The $12.9 billion rumor
On August 26, The Information reported that Nvidia had agreed to buy Hugging Face for $12.9 billion. Business Insider, the same night, said the talks had not produced a signed agreement and could still fall apart. Neither company has confirmed. Nvidia is usually quick to kill a story it considers wrong. The silence is a data point. It is not a filing.

Hugging Face already turned down a $500 million Nvidia investment last year that would have valued the company at $7 billion. Clément Delangue did not want a dominant investor. An outright sale is a different decision. Hugging Face is still a small business next to Nvidia's quarter, on the order of $150 million a year in revenue if The Information's figure holds. A $12.9 billion price would be a gigantic multiple. It would also, unlike the Groq structure, trigger real U.S. and EU antitrust review. Hugging Face is European-founded. The Hub is the place open models live.
If the deal is real, the logic is ugly and clear. Closed labs are designing their own chips. Nvidia's remaining volume is everyone else: the people who download GLM, DeepSeek, Qwen, Kimi, Llama, and Nemotron. Buying the Hub is buying that storefront. It would also give Nvidia a way to push leftover cloud GPU hours into a developer audience after it scaled DGX Cloud back. I would not write the acquisition as fact until there is an 8-K. I would write that Nvidia is acting like a company that wants the software layer if serving leaks off the GPU.
Apple wants the tokens that never leave the desk
A day before the earnings print, Apple dropped M6 and M5 Ultra. M6 is the company's first 2-nanometer Mac chip, in a Mac mini that starts at $899. M5 Ultra is the one that belongs in this story: the first quad-die M-series part, up to an 80-core GPU with neural accelerators in every core, up to 512GB of unified memory, 1.2TB/s of bandwidth. Apple's line, not mine: run huge language models with hundreds of billions of parameters entirely on device, "without counting tokens or worrying about rising cloud costs." Four Mac Studios can cluster over Thunderbolt 5 and share a memory pool. Apple claims up to 3x faster distributed inference than one box.

Do not confuse this with Rubin. 1.2TB/s is still far below an HBM4 GPU rack. Batch throughput still lives in the datacenter. What Apple is selling is the work that does not need a datacenter: developers, small teams, local agents, the 400-billion-parameter open model that used to mean an API bill. Nvidia is still inside Apple's Private Cloud Compute for some confidential cloud inference, and it just launched DGX Station for Windows as a deskside counter. The two companies are customers and competitors in the same week. That is the tell.
Training is still theirs. Serving is the fight
"Nvidia should be scared" is a tweet. Colette Kress is guiding $108 billion. China was already written to zero. Losing a market Washington and then Beijing will not let you serve is not a 2026 surprise. Google has run TPUs through several generations of Nvidia still growing. Captive silicon segments the GPU market. It has not historically zeroed it. Groq, the sharpest merchant-inference story of the last two years, is now a Nvidia license and a Nvidia cloud partner. DeepSeek-V4 shipped day-0 recipes on both CUDA and Huawei's CANN. Nvidia is still one of the two stacks that work on launch day.
The honest split is three theaters. China data-center compute is gone. Nvidia said so. Domestic inference is now good enough to serve GLM and DeepSeek-class models, even if nobody has a public pie chart of leftover H20 stock versus Ascend. Global training, post-training, and anything whose architecture is not frozen is still a CUDA fortress, with two loud exceptions: Gemini on TPU, and Anthropic dual-homed on Trainium and TPU. Global serving is the live fight. Incremental decode at OpenAI, Anthropic, Google, Microsoft, and Meta will leak onto chips those companies own. Neoclouds, enterprises, and "any model, any cloud" will leak much slower. Deskside is a new, smaller war for tokens that never become a GPU hour.
Export controls donated the China inference market to Ascend. Beijing then refused the H200s Washington finally licensed. Hyperscalers and OpenAI are trying to donate Western decode to themselves. Apple is taking a slice home. Nvidia is answering by eating Groq, selling a bigger rack, and, if The Information is right, buying the Hub. Training is still theirs. Everything else is a share war. The $96 billion quarter is what a share war looks like before the shares actually move.
Frequently asked questions
Did Zhipu really run GLM-5.3-Flash only on Chinese chips?
Zhipu says the ox-alpha stealth traffic and GLM-5.3-Flash serving ran on Chinese AI chips, and that this was their first large-scale run on a large domestic cluster. The company did not name the vendors. The widely repeated 100,000-chip figure comes from LatePost, not from Zhipu's blog. Independent proof of a 100% domestic mix versus leftover Nvidia stock does not exist in public.
Is the U.S. still blocking Nvidia H200 chips to China?
No. In January 2026 the Commerce Department moved H200 exports to case-by-case licenses. The brake this year is Beijing. Commerce Secretary Howard Lutnick told the Senate in April that the Chinese central government had not let firms buy, so investment would stay on domestic industry. Nvidia's current outlook still assumes zero China data-center compute.
What is OpenAI's Jalapeño chip?
Jalapeño is OpenAI's first custom inference accelerator, unveiled with Broadcom on June 24, 2026. It is built for LLM serving, not training. Engineering samples are running lab workloads, including GPT-5.3-Codex-Spark. Initial deployment is targeted by the end of 2026, with a multi-generation, gigawatt-scale plan alongside Microsoft. ChatGPT is not running on it at fleet scale today.
Is Nvidia buying Hugging Face?
Not as a confirmed fact. The Information reported on August 26, 2026 that Nvidia agreed to buy Hugging Face for $12.9 billion. Business Insider said the same night there was no signed agreement. Neither company has confirmed. Hugging Face previously rejected a $500 million Nvidia investment at a $7 billion valuation.
Does Apple's M5 Ultra replace Nvidia GPUs?
No. M5 Ultra is deskside silicon: up to 80 GPU cores and 512GB of unified memory for local models. It does not match datacenter HBM bandwidth or batch throughput. Apple still uses Nvidia GPUs for some confidential inference in Private Cloud Compute. The threat is tokens that never become a cloud GPU hour, not a Rubin replacement.
Is Nvidia actually losing money on this?
Not this quarter. Fiscal Q2 2027 revenue was $96.2 billion, with $89.0 billion from data center, 75% gross margin, and $108 billion of Q3 guidance. The structural risk is serving mix in 2027 and beyond, plus a China market Nvidia has already written to zero, not the current P&L.



