WAIC Field Notes (Part I): Standing in DATA CENTER INFRA, Reading the AI Value Chain
🗺️Read it another way: the venue in 3DWalk into the hall floor by floor, booth by booth; every booth jumps back to its section hereOpen the animated version →17–20 July 2026 · Shanghai World Expo Exhibition & Convention Center · This part covers Hall H1 · ~38,000 characters in the original / about a 75-minute read
The reader's verdict (AI review)
What this piece is worth is the angle, not the amount of information in it. Almost every WAIC write-up looks down from the model layer, whose parameter count is bigger, whose leaderboard rank is higher. This one looks up from the switchgear. The author sells data-center infrastructure for a living, so he pays attention to what nobody else writes down: what percentage of a machine room's equipment a single CDU actually is, why shipping containers overseas can compress a one-year delivery into five months, and how the fact that "compute hardware eats seventy percent of the cost, GPUs alone forty to fifty" squeezes the electrical-and-mechanical budget down to loose change. That position gives him a feel for both ends at once, the ones selling shovels and the ones using them, and it is the best thing in the piece.
The cost is the other half. The further up toward the application layer he goes, the more the judgments lean on second-hand material: the density of numbers goes up, the sense of having been there goes down. The Alibaba and Tencent sections read more like tidied-up sell-side research; the Vertiv and DaoCloud sections could only be written by someone who stood in front of the booth and asked the product manager.
If you only read three sections: 01 Vertiv (the thermal math vs. the business math), 04 DaoCloud (which layer CUDA's barrier actually sits at), 10 Alibaba (why five full-stack layers decide who gets to choose).
(This section was generated by an AI reading the whole piece as a reader. It does not represent the author's views.)
TL;DR
- The technical barrier in liquid cooling sits on three things: certification, product form factor, and engineering delivery. That everyone can buy their way in this quickly says R&D itself is not the gate.
- Once single-rack power crosses the air-cooling ceiling, the business math has to give way to the thermal math: liquid cooling is a single-digit percentage of machine-room equipment, next to nothing against the GPU price. Agonizing over the payback is solving the wrong problem.
- The bottleneck in taking data centers overseas is land, certification, freight forwarding and labor; technology is the least of the worries. Containerized, productized delivery (Vertiv SmartRun) is aimed at exactly those, and it compresses an overseas 1-year cycle into 5 months.
- Compute hardware (cards + servers + networking) takes seventy percent of an AI data center's cost, GPUs alone forty to fifty; the electrical and mechanical gear is only the remainder.
- CUDA sits above the operating system and below the model: every generation of large models has to go through it, and that layer of ecosystem, stacked up over more than a decade, is NVIDIA's hardest moat.
- Agent = Harness + Model. The same model under a different harness can be a full length apart. In 2026 the real product difference is at the engineering layer.
- Only a full stack gives you the right to choose: the only two with all five layers (INFRA → chips → cloud → foundation models → agentic) are Alibaba and Google. Alibaba can carpet the whole model matrix because it had those five layers first.
- The clearest signal that AI has been productized is per-seat pricing: tokens, call counts and model choice all go into the black box, and the user gets multiple choice instead of fill-in-the-blank.
- The model layer's gross margin is the most fragile one: it is built on holding frontier capability exclusively. No process node, no capacity, no physical constraint it can defend. The moment open source closes in, the profit falls back faster than it went up.
How to read this
This piece is organized along the route I walked, not re-sorted by layer of the value chain. Re-sorting would destroy the sense of having been there, and that sense is the only edge this kind of writing has over a research report. The logical thread is left to the intro at the head of each part.
The folds marked 💡 Tech and 📊 Data are deep dives; skip them and the main line is unaffected. The ones marked 🍿 Sidelines are odds and ends and grumbling. Term lookups are in Appendix B. Read only the unfolded parts and the piece still runs continuously.
The [n] markers in the body point to the references at the end. Image captions are in the parentheses below each figure.
Commercial details touching my day job have been de-identified: the orders of magnitude and the judgments stay, the identifying information goes. Any number without a public source is marked "as I understand it" or "my estimate".
About the author, and the disclaimers
Me: key-account sales at Vertiv, currently a master's student in AI at Tongji University, with an undergraduate background in automation. More at arthurzhou.ai.
Disclosure: section 01 is about the company I work for. Read the parts involving Vertiv products as a practitioner's view, not a neutral review. Every product spec mentioned follows publicly available material and does not represent the company's position.
This log is for my own record and thinking. The point is to exercise business-logic thinking and an understanding of what sits upstream and downstream, and at the same time to train how I put an argument together and to take the edge off my FOMO. Where something is well over my head I try to reach for an analogy, and I get Claude to help me learn it and spell it out.
If you see it differently, come talk to me, and leave a kind word while you are at it. Always happy to set up a coffee chat and get better together.
🍿 I'm on a Möbius strip, apparently
The slower I write, the further behind AI's changes I fall; the further behind I fall, the more there is to write; the more there is to write, the slower I write.
(Fig. 1 Write → chase → change → more. A strip with no front and no back. AI-generated.)
Route walked, order written
This piece follows the route I actually walked; the writing order was nudged slightly within a hall (things are grouped by logic inside one hall, the order across halls is unchanged).
| Hall | Location | Contents | This part |
|---|---|---|---|
| H1 | Ground floor | Industry applications and the large-model ecosystem | ✅ This part |
| H2 | Ground floor | Compute hardware and infrastructure (Huawei, Rockchip, Sugon, Runjian, Doubao Phone) | Part II |
| H3 | Second floor | Embodied intelligence and robotics (Galbot, Unitree, Leju, LinkerBot, Robbyant, AgiBot, Humanoid Robot Shanghai) | Part II |
| H4 | Basement level | Weimob, brain-computer interfaces | Part II |
The actual route was H1 → H3 → H2 → H4.
This was my third WAIC. My habit is to go in on the afternoon of day one, when it is empty and you can look at products in peace, and to watch the livestreams on the morning of day two, to read which way the government is leaning and listen to the big names talk.
Part 1 · H1 Industry applications and the large-model ecosystem
What you will see in this part: twelve vendors, one after another in the order I walked them that day, which means the shovel sellers and the shovel users keep alternating. Vertiv is followed straight away by Sangfor and INTSIG, then it doubles back to DaoCloud. So all you need to hold on to is which layer each one stands on: selling shovels are Vertiv and DaoCloud; cloud and engines are Alibaba and Baidu; the model layer is SenseTime, MiniMax, Kimi and MOSS; the application layer is Sangfor, INTSIG, Kingsoft and Tencent.
Two things in this part run against intuition. One, the closer a vendor is to the application layer, the plainer its booth: Kimi put out a few computers and talked to people, while the shovel sellers all but hauled a whole rack in. Two, the direction profit flows and the direction hardware money goes are opposite: compute hardware takes seventy percent of a data center's cost, GPUs alone forty to fifty, and yet the incremental profit flows disproportionately to the model layer.
The layering of the industry used in this piece is as follows[1].
(Fig. 2 Panoramic layer map of AI infrastructure. Source: [1])
The layers H1 covers[2]:
(Fig. 3 What Hall H1 covers within the AI-infrastructure layering. Source: [2])
01 Vertiv (维谛技术) | Nasdaq's first data-center stock, covering the "power and cooling" from 110 kV all the way to the rack PDU. Its two cards right now are liquid-cooling CDUs and containers for export
Vertiv is Nasdaq's first data-center stock, and it is also where I work. Its product coverage is wide enough to cover almost every part of data-center infrastructure; power and cooling are the main lines.
The products running hottest lately would be 800 V, the CDU and the engineering product that goes with it, SmartRun, plus the One Core concept. For the details of individual product lines, see Appendix D at the end; if you want a fuller introduction or a technical exchange, just get in touch with me directly.
First, the skeleton of this industry:
(Fig. 4 Data-center infrastructure architecture: the power chain × the cooling chain × the IT load. AI-drawn.)
- Qianzhan, “Foresight 2026: Panorama of China's IDC (Internet Data Center) Industry in 2026 (market size, competitive landscape and trends)”, industry research report[3]
(Fig. 5 Structure of the IDC value chain: upstream are the construction and equipment suppliers (power distribution, cooling, racks, optical modules, switches); midstream are the IDC service providers; downstream are internet, cloud, finance and government customers. Source: [3])
(Fig. 6 Ecosystem map of the IDC value chain: Vertiv sits right there in the upstream hardware-and-software equipment column; midstream are the three carriers, the third-party IDC operators and the cloud vendors. Source: [3])
Where data-center infrastructure sits is shown in the figure, and its importance goes without saying. It is the shovel-selling link in the golden AI industry, it belongs to infrastructure construction, and it lives inside traditional manufacturing, mostly industrial equipment. But the technical barriers around the products keep shrinking.
The way I see it, the moment the data-center business went off was DeepSeek blowing up in early 2025, around Chinese New Year. I still remember how sharp my boss's commercial nose was: in the first week back after the holiday he was already hunting for which colo (a data center rented out mostly by a third party) DeepSeek was training its model in. And he turned out to be right.
After the attention paper came out (Attention Is All You Need), the industry started actually using the Transformer architecture to train models. On the training side, parameter counts exploded from GPT-3 onward; going from 3 to 4 took nearly 100× the training budget. And the demand on the inference side got propped up right along with it. OpenAI's fake-OPEN technological dictatorship did its part to provoke competition at home and abroad.
Time moved fast. Training got its Chinchilla law, and everyone went looking for the balance point between training and inference. The biggest news between '23 and '25 was the explosion of NVIDIA's GPU line, which showed everyone the enormous output, and the future, of parallel graphics computing applied to training tokens. And so the arms race began.
Probably the same as the early days of the internet boom: grabbing resources and staking out a position fast is priority number one.
From Alibaba's far-larger-than-expected RMB 380 billion (~$53B) cloud-plus-AI infrastructure plan to ByteDance's Clover plan, infrastructure resources are the resources that buy you a head start.
For data centers, the business simply has to exist. Nobody has talked up the importance of data centers for a long time, but my own view is that a data center has always been the 1 in front of all those zeros in reliability, the same logic as insurance.
AI going off has widened both the range of things data centers are used for and how much they have to carry, which raises their importance further. As GPU prices shot up, the importance of building data-center infrastructure went with them, and traditional gear like power distribution got a gold-plated coat of its own, adding weight to AI. Whether a GPU can put out tokens as efficiently as possible is the final question for everyone in INFRA (and in the end there is no getting around the capitalist's most efficient squeeze: enter the laws of business).
On the competitive state of data centers, I'll say something reckless: it has gone from grinding on technology, on product, on information, all the way down to grinding on price and on resources. The focus has come off the safety of the data-center product itself and turned into every professional manager's KPI loop and their impatient cost-down-efficiency-up targets.
The INFRA territory Vertiv works in, which is the part I know, splits into power and cooling, from a 110 kV substation all the way down to the far end, where a GB300 GPU is 1.1–1.4 kW per card and 130–140 kW for a full rack of TDP. Precisely because the far end is already at that order of magnitude, “70 kW in a single rack and you have to go liquid” later on is not scaremongering.
Roughly working it out across different sparsity specs, compute levels and output ratios, one kilowatt-hour produces somewhere between tens of thousands and a few hundred thousand tokens (put the other way round: a million tokens burns about 15–20 kWh). This is why electricity matters so much, it is China's biggest advantage in the competition with the US, and there is even a logic to exporting tokens.
For how many tokens one kilowatt-hour actually produces, and for the commercial logic of exporting tokens, see Appendix C at the end.
💡 From 110 kV to the rack PDU: exactly which equipment “power” and “cooling” mean (the full chain)
The power chain: 110 – 220 – transformer – medium-voltage switchgear – distribution board – low-voltage switchgear – UPS – feeder cabinet – busway – rack PDU cabinet – server power supply – in-rack power.
The overall architectures are 2N / N+1 / DR / RR / 4N3.
Diesel gensets back up the 2N busbar; lead-acid batteries (lithium overseas) back up the UPS.
Cooling splits into two kinds of system. Water systems: dry coolers, cooling towers, chillers, and the terminal units. Refrigerant systems: condensers and the terminal units.
A CDU can be thought of alongside either the water or the water-refrigerant system. Overseas the mainstream match for GB300 right now is a 1,350 kW centralized CDU; for Rubin it is a 2.3 MW CDU; there is also a 121 kW in-rack unit. (All specs given as I understand them, using Vertiv as the example.)
CDU prices have been climbing, and that too is about dissipating TDP. Once a single rack passes the practical ceiling of air cooling, 40–50 kW with refrigerant-pump in-row units (the industry's accepted figure for conventional air cooling is really only 20–30 kW per rack), you have to start thinking about the liquid-cooling split. And once you are past that point, you are keeping a different set of books: a liquid-cooling system is about 5% of the equivalent cost of a data center's base equipment, next to nothing against the price of a GB-series GPU, and what you have to secure first is whether the rack will run steadily. The fold below is the full walk-through.
💡 Past 70 kW in a rack, why the business math stops mattering and only the thermal math (the technical math) is left
It is generally 70:30 or 80:20 (liquid takes 70–80% of the heat away and air still has to catch the rest; Vertiv's own white paper puts cold plates at 70–75% of a rack's heat), for a thermal design above 70 kW. I started out doubting that ratio from the business logic too.
Here is the conclusion I got from asking a Vertiv technical director, and the thinking it set off, for whatever it is worth to you: a liquid-cooling system is about 5% of the equivalent cost of a data center's base equipment, next to nothing against the price of a GB-series GPU. In the business, whether the server rack keeps running is what matters most. Once you break through 70 kW in a single rack, what you have to think about is how to get the rack cooled and running in the most efficient way, not whether the business logic closes, that is, whether the system is worth what it costs to build.
The PUE red line has in fact never gone away. Data centers have been promoting the dual-carbon 3060 goals since 2021 (peak carbon before 2030, carbon neutrality before 2060), and now that AI is a national strategic reserve, four ministries including the NDRC still issued the Special Action Plan for Green and Low-Carbon Development of Data Centers in July 2024 (NDRC Huanzi [2024] No. 970), and Beijing will charge differential electricity tariffs on low-efficiency data centers from 2026. What has actually changed is that PUE is no longer the only KPI (it used to be all but the only yardstick, because you needed it for window guidance, policy subsidies and approval eligibility), and now the emphasis is on what actually works. Of course, bringing in liquid cooling has unquestionably pulled PUE into the 1.1–1.2 band, and immersion can push it below 1.1.
Cooling has always moved faster into production than the power system does. I judge the changes in data-center electrical and cooling by two things: what the competition is loudly promoting, and whether the workload of colleagues in the same discipline has changed.
The latest trend is 400 V, 800 V and SST, which got carried by Jensen's 800 VDC / 1 MW rack roadmap (NVIDIA's official line is that 800 VDC covers 100 kW to over 1 MW and goes into volume with Kyber rack-scale in 2027; the highest spec actually put on the table is Rubin Ultra / Kyber at 600 kW, with 1 MW still further out) and became the frontier ambition on the electrical side. With DC and AC systems now serving as the big vendors' innovation KPI, this happens to coincide with China pushing modular UPS, a finer-grained product that breaks the whole unit apart and sells it as components.
What a UPS is, and what exactly modular UPS breaks apart, is in Appendix E at the end.
And, awkwardly, ByteDance's data centers started running centralized procurement on several major categories: diesel gensets, UPS, medium-voltage switchgear. It punched straight through the whole UPS industry and turned the UPS into one screw in the data center's electrical chain. That is roughly how the industry is; constant renewal is the normal state of iteration. But the UPS going cabbage-cheap seems to have come a little fast (measured against how long I have been working).
On the electrical product landscape, the pursuit is DC systems feeding DC-capable servers directly, plus things like OCP-spec PSUs inside the rack; DC systems also cover 400 V–800 V and SST. To my mind, whatever the electrical architecture, it cannot get away from how the chips are used: with the Atlas 950 SuperPoD in China estimated by the industry at 70 kW per rack, without high-power demand from the chips, 800 V and SST still need at least two and four years respectively before they land.
Switching to the cooling side, I hear more than twice as much complaining from cooling SAs and product managers as I do from my electrical colleagues; I even met a French cooling PhD on social media who wants to move into data centers to work on liquid cooling. Liquid cooling blowing up arrived right on schedule, and yet it is a bit of a castle in the air (for China, at least).
Same as with the electrical products above, liquid cooling has surged along with what the chips demand. The key point is that if you do not cool it properly you really do cut into what the chip puts out (which violates the entrepreneur's law).
The number of variations in cooling systems is basically the four refrigeration components permuted back and forth like A(n,m) in combinatorics: compressor, condenser, throttling device (expansion valve / capillary), evaporator. There are too many; just look at the figure. I have barely touched chillers, so I will not go into them.
(Fig. 7 The four components of the refrigeration cycle and the product forms built on them. AI-generated schematic.)
My revision notes on cooling systems are in Appendix F at the end.
The variety in cooling products basically comes down to different forms of those four components, so every time one changes, a new technical route and a new name are born. It happens all the time that while a product manager is introducing our product name and its technical concept, the customer cuts in and says: how is this different from chilled water or a refrigerant system, what has actually changed?
As the forms of cooling products keep updating, the good news is that the cooling folks are never short of topics or directions; the bad news is that the work never ends.
On the cooling side the CDU stands alone. As the GPU's core partner, it has never dropped out of the front rank of attention. And this track is absolutely a great destination for fluid mechanics people out of a power-engineering school.
On engineering delivery, I have to bring up Vertiv's SmartRun and the One Core concept. SmartRun is Vertiv Global's prefabricated containerized delivery answer to what cloud and AI vendors want right now, which is data centers built fast: the electrical and mechanical systems are integrated and tested in the factory and only then shipped to site, squeezing on-site construction work down to a minimum.
The problems a containerized product has to solve to go overseas, in keywords: land is hard to get approved abroad, certification is hard, the export resources, the high price, and the long delivery cycle.
Getting land approved is not like at home. In China you can work it out by following government policy; abroad you generally need to be laying the groundwork a year or more ahead to find a good location. And siting a data center means weighing a long list of factors, including but not limited to climate, government subsidies, the peak-valley price spread, staffing support and latency. So data centers usually come in clusters: everyone liked the same plot of land, because it meets all of the above, value for price.
Which specific things you look at when siting a data center is in Appendix G at the end.
Product certification is another old headache. For Chinese products going out, the key is getting through the two big systems: the EU's CE (with the LVD low-voltage directive + EMC + RoHS underneath it) and North America's NRTL (UL is just the most famous of them; ETL and CSA carry the same weight). Below that there is the IECEE CB scheme and each country's own mandatory certification (Japan PSE, Korea KC, Australia/NZ RCM, Middle East SASO, and so on). Then there is another gate: which certification a power skid falls under, and whether it meets the local government's technical requirements. Certification is not cheap, and set against business that is hard to land, it puts people off before they start.
Lining up the export resources means dealing with freight forwarders, the form the export takes, customs declaration and clearance, and delivery terms such as DDP / DAP / DPU. If the equipment vendor has to file the customs declaration, you then have to switch the contract over to the Global entity. Tax and customs questions are usually handled by a dedicated overseas department the company sets up on its own, which tells you how hard they are.
The high price usually means labor abroad is far below China's level: labor costs three times what it does in China and puts out less than half as much. Getting Chinese engineers overseas to build overseas data centers in a sensible way is a problem of its own. However lofty a data center sounds, it still sits in traditional manufacturing, and in that field, china no.1.
The long delivery cycle has a few causes. Building a data center in China is generally T+6, counted from when the civil works shell is up, +6 being when the equipment is accepted and handed to the owner. Abroad a cycle routinely runs a year, long enough to miss two COMPUTEX shows. Which means your build's technical architecture has not even changed yet and two more GPU specs have already launched, so you are forced back to change the technical plan and adapt to the newest card. (For the commercial logic on this point, jump to Part 2, where I go through what I understood in the H2 chip hall.)
All of this eventually lands in one set of books, so let me put the big numbers up front: compute hardware (cards + servers + networking) already takes the overwhelming share of a data center's cost, as much as 70%, GPUs alone forty to fifty percent, and the electrical and mechanical gear is only the remainder. The fold below is the full cost account.
📊 Where a data center's money actually goes: compute hardware takes 70%, E&M is the remainder (the cost account)
Data centers also have to account for equipment depreciation. In China, even without factoring depreciation in, the annualized figure is already 7%; the IRR drops much harder once you do account for the infrastructure and the GPU equipment itself, and payback can go from 9 years to 18–20 (thinking in terms of a leasing model; the comparable public figure is a payback period around 8.5 years and a post-tax IRR of 10%–15%). If you cannot get the business up and turned into revenue quickly, the grilling from the capital side is very real every single time.
Building a 1 GW data center overseas benchmarked on GB300, the total cost comes to $47 billion. Let me be clear about what that number covers: $47B ÷ 1 GW ≈ 47 USD/W, i.e. 47,000 USD/kW, and that is the all-in figure including GPUs. The set of numbers below is a different set of books, the electrical and mechanical subcontract only, no cards and no civil works. Read each on its own terms; do not subtract one from the other.
Counting the E&M subcontract only: for hyperscale and colo data centers in China, the E&M subcontract build cost is about 2,000 USD/kW excluding tax; Malaysia about 5,000 USD/kW; Japan about 12,000 USD/kW; the US about 20,000 USD/kW. (This set is the industry figure as I understand it, E&M only. Published reports all give total build cost including civil works, so they can only be compared at the level of magnitude: Turner & Townsend's 2025 data centre construction cost index puts Tokyo at 15.2 USD/W and Silicon Valley at 13.3 USD/W; Cushman & Wakefield in April 2026 puts Malaysia at $9.6M–12M per MW, and the overall figure for US AI data centers in 2026 lands at $15M–20M per MW.)
The same piece of E&M work is ten times apart between China and the US. The ten times is not the equipment being ten times different, it is the gap in infrastructure capability and equipment capacity — production scheduling for complete sets of equipment, supply-chain pricing, the efficiency of the installation crews. Those happen to be exactly what China has no shortage of. And precisely because that price gap is sitting there, companies like DAYONE, which take China's supply chain and construction capability overseas to build data centers, are so well received by foreign capital. It is the overseas business of GDS spun out as an independent company, headquartered in Singapore; per public reporting it confidentially filed with the US SEC in August 2026, aiming to raise about $5 billion at a valuation of up to roughly $20 billion, could list as soon as next quarter, and is considering a dual listing on SGX and Nasdaq (per Bloomberg and Reuters; the exchange is not finally settled).
But on an overseas project, Chinese equipment at +20–30% on price, combined with labor at twice the Chinese rate, works out to total cost up about 30%. Rent, meanwhile, has gone to this: one colo operator's overseas price is $85 per kW per month including tax, excluding land and excluding power — against a Chinese comparison price of under RMB 250 per kW per month. Cost up thirty percent, rent up several times over: that arithmetic is itself the reason to go overseas.
So getting racks up and turned into revenue quickly is priority number one. Which makes the importance and the value for money of shipping containers overseas self-evident.
A productized SmartRun has a much shorter delivery cycle. Take an overseas job: equipment delivered within 2 months, 2 months of integration and installation, 1 month of shipping, entirely comparable to China's T+6 build cycle. Of course, how to solve the local labor situation is another big problem, and overseas colos and some suppliers' agents will work hard at it.
SmartRun on the floor:
(Fig. 8 The SmartRun prefabricated module at Vertiv's booth, with racks on both sides framing a cold aisle you can walk into.)
Thinking
Having actually been round the factory and the liquid-cooling exhibition, the technical barrier, as I see it, lies mostly in getting the product stamped by the big vendors' certification, in the product form factor, and in engineering delivery. That everyone can acquire a liquid-cooling target, or keep posting results in the liquid-cooling track, also says the R&D barrier itself is not deep — most of them do not have people from a liquid-cooling background building the CDU, they have water-system R&D people moved over to own the CDU track.
Blue ocean ahead for DC going overseas. Time to set sail.
02 Sangfor (深信服科技) | Taking open-source FastGPT and turning it into a commercial To-B version (AgentBuilder). The way in is the visual workflow; where it lands is still their own cloud
FastGPT: FastGPT – an enterprise-grade AI agent building platform | open-source RAG system
Sangfor commercialized open-source FastGPT into a To-B version. Walking the booth, it breaks into these pieces:
Import your knowledge as a database. You can bring the company database in and do targeted content extraction on it, which speeds up answers and makes them more accurate.
Visual workflow (👍). You build a flow, the interface stays legible, repetitive tasks become automatic and complex tasks become reviewable. Very much like an Obsidian mind map, a WORKFLOW visual edition, the mind-map version of an AI-native one-person company and its sub-agents. You can lay out the work logic clearly and find the steps that matter. Good value.
(Fig. 9 FastGPT's visual workflow editor: a “Résumé screening assistant_Excel” flow strung from the start of the process through an upload check and batch execution all the way to a specified reply.)
(Fig. 10 The “Contract review assistant (basic)” editor open on the big screen at the booth, with draggable nodes down the left column: AI dialogue, knowledge-base retrieval, Agent, form input, HTTP request.)
Intelligent data analysis. The product manager on the floor demonstrated the contract review assistant, showing how it scans a contract, recognizes text, signatures, figures and tables, and uploads them into the content. But I found out it is plugged into somebody else's OCR MCP (possibly Baidu's open-source PaddleOCR, or INTSIG's), so the scan quality depends entirely on the quality of the OCR you plug in. Slightly awkward.
💡 The remaining two pieces are both table stakes: stuffing DingTalk's features into WeCom (the capability list)
Workflow orchestration. Like stuffing DingTalk's features into WeCom: integrate customized functions into your workflow and add some non-standard pieces, such as running a compliance check before the conversation starts.
API integration. Table stakes. Counts as an open interface, and it can be embedded straight into your working PaaS.
The flow diagram, taken from the official site:
(Fig. 11 FastGPT Functional Architecture: the knowledge base goes through a preprocessor into the language model for QA splitting and text chunking, then through the embedding model into PostgreSQL; on the dialogue side the Flow Controller calls the orchestrated flow, vectorizes the question and runs similarity retrieval. Screenshot from FastGPT's site.)
(Fig. 12 Booth panel: “FastGPT Sangfor commercial edition (AgentBuilder) · 0 experts, good results”.)
Thinking
It is mainly used by enterprises doing digital transformation: deploy the database on premises, use hosting or the AI platform to get closer to what the customer needs. Underneath it, what still gets sold is Sangfor's own cloud computing business.
A To-B platform like this has to separate out the usage patterns of different industries. But if all you are doing is RAG and a private deployment, why not just go find the FastGPT people and have them tune it yourself? It is still an outpost for MaaS.
03 INTSIG (合合信息, CamScanner) | The consumer-side scanning is already good enough; on the enterprise side “Qixin Huiyan” wants to be an AI Qichacha
Still mostly consumer-facing content. The newest thing they were promoting was scanning old hanging scrolls into a cleaner digital archive, which suits electronic notes for a museum or an exhibition tour rather well.
At first I thought it could scan calligraphy into modern text; it turns out it just becomes a cleaner digital copy (😓). It removes glare, keeps fidelity, and preserves things better. What I actually care about is the sharpness and speed of scanning invoices or expense forms, and that is already good enough.
(Fig. 13 Digitizing old paintings at INTSIG's booth: two portrait screens side by side, a scanned blue-green landscape handscroll on the right, calligraphic documents from the same batch on the left, and the paper originals laid out in a glass case underneath for comparison.)
The To-B application: Qixin Huiyan, an AI Qichacha.
You can ask it directly for a company's official data, it plugs into publicly released MIIT data in real time, and it locates and drills through equity charts fast, which is handy for sourcing and for understanding who is connected to whom.
🍿 A counterexample this brought to mind: a bank AI lookup system whose database had to be updated by hand
(Which reminded me of the financial lookup system the Agricultural Bank of China put out last year, an AI platform built with East China Normal University. The overall function and idea was to use AI to pull company annual and financial reports quickly, do the research and give an assessment, but the database, the most critical part, had to be updated by hand every time.)
Thinking
On the consumer side, good or bad comes down to price: if they can do AI search and open up more features at a price below Qichacha and Aiqicha, it will work much better.
Underneath it, lookup software in China is still a license business.
04 DaoCloud (上海道客) | A K8s and inference-engine vendor, wedged into the layer of “glue” above CUDA and between compute pools
A K8s technology vendor, accelerated inference services, a scheduling platform, the “glue” between software and hardware.
To do inference well you need an engine on top of the CUDA architecture; it can allocate across multiple clouds and multiple compute pools. Very much a token relay station, except the engine version of one.
My impression is that DaoCloud started in hardware and is now running the engine track alongside it; the outside world usually reads it as cloud-native / K8s software by origin. Both accounts are out there, and I may be misremembering.
One premise first: NVIDIA's higher moat is not stacking wafers or the precision of the nanometers, it is the technical barrier of the CUDA inference architecture. Every generation of large models, and other inference and compute besides, is written on top of that closed-source architectural language. It is not open source, it only runs on their own cards, nobody else can take it and modify it, so they have to start a separate one and then work their way back to catch up on compatibility. So basically every large-model shop has to deal with adapting to CUDA.
💡 CUDA sits above the operating system: the position more than a decade of ecosystem was stacked into (the technical account)
CUDA's ecosystem happens to sit right above the operating system, and for all these years software developers and companies have kept adapting to it. It is on the back of the powerful GB-series cards and the CUDA architecture that we have today's LLM explosion. PyTorch and TensorFlow are deeply optimized for CUDA (cuDNN, cuBLAS), and both of those libraries are closed source as well. The longer the adaptation goes on, the deeper the upper ecosystem settles into that closed stack, and the more the real barrier is stuck at the ecological niche rather than at the process node.
This is also why DeepSeek was so set on adapting to Ascend chips — it wanted to break open the moat of training on CUDA. As I see it, that step did not work; maybe it is also why DeepSeek V4PRO shipped late, and in the end they still trained on NVIDIA's GB series, which tells you how much CUDA matters to training a large model. (A guess; correct me.)
(Fig. 14 NVIDIA's own drawing of the software stack: CUDA (the row in the red box) is sandwiched above “operating system and system software” and below the NVIDIA and third-party libraries and frameworks; only the bottom layer is CPU / GPU / DPU.)
And the engine, K8s (Creating a cluster with Minikube | Kubernetes[4]), sits above CUDA, mainly doing management and orchestration.
(Fig. 15 DaoCloud's “architecture modernization” view: the bottom layer is the Kubernetes distribution (native compute / networking / storage / observability collection), then container management, multi-cloud orchestration, the microservice engine and mesh services, middleware, observability and security stacked on top, with the application workbench only at the very top.)
(Fig. 16 DaoCloud Enterprise 5.0 Platinum product architecture: above the cloud-native base sit multi-cloud orchestration, data services, microservice governance, observability, and the app store / virtual machines / application delivery, spanning public cloud, private cloud, domestic heterogeneous silicon (Kunpeng / Phytium / Hygon / Loongson / Zhaoxin / Kylin / UOS / openEuler) and cloud-edge collaboration underneath.)
(Fig. 17 A laptop open on DaoCloud's booth showing d.run Token Factory: down the left are high-performance networking, LLM-D distributed inference orchestration, KV Cache management, and the vLLM / SGLang / TRT-LLM inference engines; on the right are the vendor's own acceleration comparison numbers, with an NVIDIA Dynamo sign standing next to it.)
Thinking
I asked the product people how long it would take for China to break through CUDA's architectural moat. From where they sit, there is still a long road.
The same large model performs differently under different engines. I am very much a layman in this area, so that is as far as I can take it.
05 Baidu | The internet giant that got in earliest, and yet “never catches anything while it is still hot”
On Baidu, here are my own biases.
Baidu has always been where internet technology in China points first, and it was the fastest of the big internet companies to get in. But the momentum afterwards has never been what you would hope.
When Ernie Bot launched on 16 March 2023 it carried the title of “the first company among global majors to build a product benchmarked against ChatGPT” (Robin Li's own words at the launch; the media at the time generally called it “the first Chinese ChatGPT”). But my impression is that its volume today is clearly below the first tier, and the public data roughly agrees. On the consumer side, per QuestMobile, Ernie Bot had 5.31 million monthly actives in September 2025, against Doubao's 172 million and DeepSeek's 145 million over the same period. By June 2026 the top three AI-native apps by monthly actives were Doubao at 382 million, Qwen at 167 million and DeepSeek at 130 million, with Baidu no longer near the top of the list. On the cloud side, per IDC's Analysis of the Latest Landscape of China's Enterprise MaaS Market published in May 2026, model calls on China's public cloud in 2025 came to 1,944 trillion tokens, of which Volcano Engine took 49.5%, Alibaba Cloud 28% and Baidu AI Cloud 10%; while in the first-half-2025 figures six months earlier, Baidu AI Cloud still had 17% (against Volcano Engine's 49.2% and Alibaba Cloud's 27%).
Even in my own DC industry, Baidu's data centers were very early to put out glacier phase-change systems and magnetic-levitation multi-connected cooling designs; plenty of the experts in the industry came out of Baidu's infrastructure department.
🍿 A joke going around the industry: Dario restricts Chinese users because he once interned at Baidu? (Kidding)
There is a joke that has gone around the AI world: that the reason Anthropic's boss Dario restricts Chinese AI users so hard is that he interned at Baidu back in the day — as for what inhuman torments those few months held, the joke does not say. It is only a joke, of course; public material says he left after a few months, and the real reasons have nothing to do with that stint.
Either way, Baidu's current state is a shame. It never catches anything while it is still hot. As I see it, AI search engines like Perplexity taking over some of Baidu's use cases is a matter of time.
The greatest thing about the internet is that years of accumulated ecosystem and influence make usage a habit, and it is hard to quit. The bad thing is (my bias here) that the cut-costs-boost-laughs staff will always lie back on past credit, and slowly wear everyone's patience away. When an obedient, genuinely usable AI startup lets you try it for free, you think: Baidu? Yesterday's news, old and inefficiency. So you only feel any of this if you actively break out of your comfort zone and try new things.
Back to the product side. Apart from wanting to pick up a freebie, I really did not spend much time on Baidu's MaaS or its other software. And PaddleOCR, the thing that most interested me, was not there: this world-leading open-source OCR is the very best of what large models have done for scanning.
Citing the OpenDataLab benchmark site:
(Fig. 18 The OCR leaderboard: PaddleOCR-VL-1.6 first with an overall 96.34, then MinerU2.5-Pro, GLM-OCR and PaddleOCR-VL-1.5, broken out across five dimensions: text, formulas, tables, table structure and reading order. Screenshot from the OpenDataLab leaderboard.)
Studying the OCR architecture diagram, citing CSDN: “[Large-model basics] A complete breakdown of OCR: an in-depth guide from principles to practice” – CSDN blog[5]
(Fig. 19 The full OCR technical architecture: input comes into the data layer, goes through preprocessing (denoising, enhancement, skew correction, binarization, layout analysis), on to text detection (DBNet) and text recognition (Transformer+CTC) in the core model layer, then through post-processing for grammar correction, format normalization and semantic repair, and finally lands in the application layer: document digitization, receipt recognition for expense claims, identity verification. Source: [5])
Squeeze that diagram down and the trunk is really just one pipeline:
(Fig. 20 Simplified diagram of the OCR technical architecture. AI-drawn.)
Kunlunxin has always been something you only hear about, and I did not learn much more this time either; my attention had all gone to Huawei's 950 POD over in Hall H2.
💡 Kunlunxin revision: P800 benchmarked against the A100, M100 moved to inference, shipments tied for third among domestic vendors with Cambricon
Kunlunxin revision, as follows:
Put simply, the P800 is benchmarked against NVIDIA's A100 at 345 TFLOPS in FP16; the M100 now targets the H20 and is used mainly for inference. The P800's published spec is 96 GB HBM and FP16 345 TFLOPS dense (Kunlunxin's own site has never published the full parameters; that figure comes from TechInsights' die analysis, and the on-site reporting at this WAIC uses the same set), against the A100's 312 TFLOPS dense on FP16 tensor cores and the H20's 148 TFLOPS, so “benchmarked against the A100” holds up. The M100 is the fourth-generation chip announced at Baidu World in November 2025, positioned officially for large-scale inference, on a fully domestic supply chain, planned to go on sale in early 2026; the physical part was only shown publicly for the first time at this WAIC (20 July 2026), and what it targets is precisely the H20. After that comes the M300, planned for early 2027 and aimed at very large-scale multimodal training and inference.
Per IDC, China shipped about 4 million AI accelerator cards in 2025, of which domestic vendors accounted for roughly 1.65 million, or 41%. Kunlunxin and Cambricon each shipped about 116,000, tied for third among domestic vendors, behind only Huawei Ascend (812,000) and Alibaba T-Head (265,000).
(Fig. 21 Two IDC cuts: on the left, the 2025 share of China's AI servers by accelerator type (GPU 58% / non-GPU 42%); on the right, domestic chip vendors' share of AI accelerator adoption in Chinese data centers: Huawei 49%, T-Head 16%, Kunlunxin and Cambricon 7% each. Source: IDC.)
Kunlunxin's architecture is its own XPU architecture.
Citing the architecture diagram from its site:
(Fig. 22 The full Kunlunxin SDK software stack: at the bottom the Kunlunxin AI accelerator and the operating systems (Ubuntu / CentOS / Debian / BC-Linux / UOS / Kylin); in the middle the runtime layer, the SDK (XTDK, XTRANSCUDA, XTRITON) and the specialist libraries (XDNN, XECV, XCCL and others); above that it connects to PyTorch / TensorFlow / PaddlePaddle, DeepSpeed / Megatron, TensorRT-LLM / vLLM / SGLang; and finally models and applications. Screenshot from Kunlunxin's site.)
Also, although I did not study the server POD closely, this CDU logo really did catch my eye. Whether or not it was a certain good friend's obsession rubbing off on me, I took the photo anyway, to study the CDU architecture diagram and to reminisce.
(Fig. 23 The local control screen on the CDU door, everything on one page: primary and secondary supply and return temperatures, system differential pressure, flow rate, control-valve opening, dual-feed PSU status, with the vendor logo printed in the bottom right.)
(Fig. 24 The AIRack array at Baidu AI Cloud's booth, with a panel reading “Infrastructure for the intelligent era”.)
(Fig. 25 A close-up of the same row of racks, with paired red and blue liquid-cooling hoses and quick connectors on top, red hot, blue cold.)
PS, all that blue and green is a bit hard on the eyes.
Thinking
Baidu's current state is a shame: earliest in, and the momentum afterwards never what you would hope; it never catches anything while it is still hot. Years of accumulated internet ecosystem and influence make usage a habit that is hard to quit — that is the moat. But my bias is that with cut-costs-boost-laughs staff lying back on past credit, the moat gets worn down too. When an obedient, genuinely usable AI startup lets you try it for free, “Baidu? Yesterday's news, old and inefficiency” is the moment users vote with their feet. AI search like Perplexity taking over some of the use cases is, as I see it, a matter of time. So you only feel any of this if you actively break out of your comfort zone and try new things.
06 SenseTime (商汤科技) | An AI founder's time lag: it once “understood Chinese best”, and general-purpose models flattened the edge away
You cannot talk about SenseTime without its background. Its founder is one of the people who laid the groundwork for AI in China. As I see it, SenseTime's product cadence is half a beat behind the newer crop of large-model startups; the specific reason is hard to judge, but you can feel it in how often they announce anything.
The most direct way to encounter SenseTime's products, the consumer touchpoint most people have, is the robots in hotels and the scan-to-check-in readers. On the enterprise side it now has its own multimodal model, models specialized for various domains, a token plan, and some cloud services.
SenseTime had its brief flowering in large models too. SenseNova was once called the model that best understood Chinese as input and output, but my reading is that as overall model capability came up, the head start in Chinese-language corpora got diluted: any language is about the same when it comes to transmitting and understanding, and the old advantage largely evaporated. (This is only my observation: MoE governs parameter routing and KV Cache governs memory and acceleration on the inference side, and neither has a direct causal chain to “the Chinese advantage being flattened away”.)
SenseTime has not stopped on image generation, though. At this very WAIC (18 July 2026), SenseNova released U1 Pro, and the official positioning has already shifted from a plain image-generation model to “a natively multimodal agent foundation for long-horizon tasks”. Up to 8K native resolution output; during the run-up it was clearly aimed at GPT-Image-2; the full version and the API are planned to open to the public in August 2026. It is only that because of the situation in China the domestic corpora have to be constrained, and on top of that the best direction for a large model is a general-purpose one, so this line has not moved as fast as the main battlefield.
The models you can currently reach through SenseNova's token plan are below. Early days:
💡 The token plan currently opens only three models: two lightweight SenseNova versions + one DeepSeek Flash
(Fig. 26 The list of models available through the SenseNova open platform's token plan: SenseNova 6.7 Flash-Lite, SenseNova U1 Fast, DeepSeek V4 Flash, all three rate-limited by number of calls per 5 hours. Screenshot from SenseTime's SenseNova open platform.)
Benchmark scores for the LITE lightweight model:
(Fig. 27 Ten benchmark comparisons for SenseNova 6.7 Flash-Lite: PinchBench, ClawEval, τ³-bench, Deep Planning, NovaPPTBench, AIDABench, GPQA-diamond, AA-LCR, MathVision, OCRBenchV2, against same-size rivals Step 3.5 Flash, GLM-5, KIMI K2.5/K2.6, GPT 5.4, Gemini 3.1 Pro and Opus 4.6. Both the benchmarks and the rivals are the vendor's own selection. Screenshot from SenseTime's official material.)
Thinking
The founder is one of the people who laid the groundwork for AI in China, but as I see it the product cadence has always run behind the newer crop of startups — being a founding figure did not buy them the tempo. SenseNova, which once “understood Chinese best”, has seen that advantage largely evaporate as general-purpose models advanced. The image-generation line has not stopped, and U1 Pro is the flagship launched at this WAIC; it is only that under those two constraints, the corpus limits and “the best direction is a general-purpose model”, it does not move as fast as the main battlefield. And what they can open to the outside through the token plan today is still early days. The revolution is not yet won.
07 Tencent | On AI, betting on products rather than models: IMA grabs the knowledge-base entry point, the game agent is still getting started
The penguin business I use most now is IMA. With WeChat behind it, it is hard not to lean on a personal AI knowledge base for documents. And you can drop in the really good WeChat public-account articles, so you do not just read them and forget them while scrolling down.
IMA is also a platform that can link everything on Tencent's side, and there are plenty of open MCPs it can plug into — Tencent Meeting, Tencent Docs, WeCom, and Feishu too. I have recorded a lot of upstream and downstream AI knowledge through IMA. If nothing else, it works well as a network drive for work documents; add a Copilot on top and AI distillation will find what you want among a pile of project files. Strongly recommended.
Monetizing later is easy too: once you have built the knowledge-base habit, it is time to start charging for the Copilot and for storage (public knowledge bases are free to store at the moment). As users pile in and stickiness rises, this kind of AI private-domain product will also hit traditional community software like Zhishixingqiu.
As I see it, Tencent's presence on the model side is weaker than the other majors, and I mean the models themselves, not the product end. Hunyuan Hy3 is still promoting its MoE architecture; its call volume on OpenRouter is enormous, which I would guess is because it is tagged free.
Two things have to be kept apart here: benchmarks are a capability ranking, call volume is a usage ranking, and what Hy3 has climbed high on is the latter, not the former. The models near the top of that board almost all carry the free tag, which tells you something in itself.
For reference, Hunyuan's parameters:
💡 Hunyuan Hy3: a 295B-parameter / 21B-active MoE, scoring 41 on independent evaluation, in the back half of the open-weights first tier
Hy3 currently uses a MoE architecture with 295 billion parameters and 21 billion active, supports 256K context and both fast and slow thinking. Independent evaluation by Artificial Analysis gives it 41, which puts it among the leading open-weights models but below GLM-5.2 Max's 51 and Kimi K2.6's 45 (the AA intelligence index recomputes historical scores with each version; these figures were taken in August 2026). Hy3 itself is still a text-only model; image and video capability sit outside it.
(Fig. 28 OpenRouter's weekly LLM call-volume board: Hy3 (tencent) third with 4.79T tokens, up 43% week on week, behind DeepSeek V4 Flash at 7.43T and MiMo-V2.5 at 7.23T; models tagged free are generally near the top. Screenshot from the OpenRouter LLM Leaderboard.)
Tencent AI's WAIC booth was also mostly about the product-application end. My guesses as to why it has fallen a little behind:
1. It cares more about cultivating user habits and leans toward building a full AI product suite: Tencent Docs, WeCom, IMA, Tencent Meeting and so on;
2. It allocates its spending differently from everyone else. Tencent put over RMB 18 billion (~$2.5B) into “new AI products” in 2025 (over RMB 7 billion in Q4 alone), and 2026 is expected to be at least double that; but full-year capex was actually RMB 79.2 billion (~$11B). The money is not short — it just has not been concentrated on one language model the way others have done it;
3. It is concentrating its push on office MaaS platforms like WorkBuddy and CodeBuddy, hoping to acquire users by lowering the barrier;
4. Its coverage is too broad: language, image, video, 3D, robotics, cloud services and a huge amount of internal business all at once, so resources are not staked on one language model the way some startups do it.
(Fig. 29 The “Deploy Your Agents” main wall at Tencent's booth: three screens in the middle showing QClaw, WorkBuddy (an enterprise workplace AI agent workbench) and CodeBuddy (an AI-native coding agent), with a strip of connectable agent-application icons running across the top. Basically the whole product portfolio is on that wall.)
What interested me most about the penguin was the game AI agent part; this is the number one game company in the universe, after all. But from the demo and the walkthrough at the booth, what Tencent's game AI agent can do today is fairly basic, aimed mainly at widening the agent's audience: you can build a game with simple instructions, and the game-production pipeline is visualized.
I had assumed Tencent would push AI straight at the game engine as its main direction, building on something like Unity or Unreal, with automatic modelling, space building, animation, 3D generation and effects generation end to end. But then I realized that amount of work is unlikely to be doable inside this little agent, and the MCPs or workflows it would have to call are far too tangled. Not realistic.
It is a lot like CAD drafting software. I remember a classmate at university proposing exactly this as a thesis topic: an AutoCAD AI workflow. But talking it through with experts at a design institute, on real projects: first, automatic drafting is possible, but hard; second, and most important, the job is to realize what the client has in mind, and the AI cannot get there, or cannot be “steered” by the actual circumstances of the project, so in the end it comes back to human skill, even if the AI can technically do it.
Back to AI landing in game engines. I have been consistently bullish on game companies: since the great era began, gaming has quietly been an industry that eats well, with growth every year in the upper-middle range[6].
And game AI landing is the ultimate paradigm for productizing AI to acquire users. Because in the AI era you either break the existing architecture and implant an AI step into it, or you use AI to build an entirely new native paradigm that replaces the old architecture; either way the end goal is to productize AI as a tool, convert users and make money. And one of the best places for that productization or that implanting to land is games, part of the EDGE AI industry where software and hardware come together in the post-AI era, and something that can really change people's habits.
As it happens, the news about Unity charging usage fees broke a while back, and only then did everyone realize how fat the profits are and how long-lived: the fee could reach games built on Unity's underlying architecture twenty years ago.
The full timeline is worth laying out as evidence. On 12 September 2023, Unity announced the Runtime Fee: not a cut of revenue but money charged per install, up to $0.20 each, effective 1 January 2024. The thresholds were $200,000 annual revenue and 200,000 cumulative installs for Personal / Plus, and $1 million and 1 million installs for Pro / Enterprise. The most damaging clause was that it applied retroactively to games already published: the game you made with Unity ten years ago, if people are still installing it today, you start paying. Developers blew up on the spot, boycotted collectively and threatened to switch engines; Unity even closed offices temporarily after receiving threats. Eleven days later, on 23 September, Marc Whitten came out to apologize and rolled it back substantially: Personal became entirely free with no fee, the threshold went up to $1 million, and it applied only to games built on 2023 LTS and later, with everything made on 2022 LTS or earlier exempt — which in effect carved the back catalogue out.
A year after that, on 12 September 2024, Unity's official blog announced the Runtime Fee was cancelled outright and it was going back to per-seat subscription, at the price of an 8% increase for Pro to $2,200 per seat per year and a 25% increase for Enterprise, effective 1 January 2025, with the free revenue ceiling for Personal raised from $100,000 to $200,000.
So that story ended in a rollback, and using it directly as evidence of “pricing resilience” needs a discount applied. But what it exposed is, I think, worth more: an engine company daring to reach back to a twenty-year-old installed base to charge for it shows the engine had pricing power over the back catalogue all along — this time they just did not manage to collect, not that they could not. And that is exactly why I am bullish on game AI landing: if some company can implant an AI engine into post-era games, then on top of the engine's fixed fee there is now an AI implementation fee, which pushes AI usage and profit conversion up further.
As for Unity's capital moves in China, the direction in public reporting is actually the opposite of “going public”: in March 2026 it was reported that Unity was evaluating strategic options for its China business including a potential sale, at a target valuation possibly above $1 billion, with Tencent, Alibaba, ByteDance and miHoYo all on the potential-buyer list (most of them already existing shareholders in Unity China), and talks still at an early stage. There are no public plans for a domestic listing at all.
Since we are on what it is worth, let me put a few numbers on this business. Unity is listed in the US, so its financials are public information, and everything below with a public source is attributed.
• Unity's gross margin: what I recorded was “over 60%”; the published financials are higher. Unity's FY2025 revenue was $1.8496 billion, GAAP gross margin 74% and adjusted gross margin 83% (74% and 82% respectively in Q4 2025); Q1 2026 revenue was $508.2 million with adjusted gross margin still 82%, only the GAAP figure got knocked to 31% by a one-time $278.7 million impairment from shutting the ironSource ad network and divesting Supersonic. So “the profits are fat” holds up, and the actual number is more than ten points higher again (Unity FY2025 annual report and Q1 2026 report).
• $700 billion games industry: this one needs its basis spelled out first. Newzoo counts only what players spend on game content, which in 2025 was $201.6 billion, above $200 billion for the first time and up about 9% year on year (PC +12%, mobile +10.7%, console +2.8%). Grand View's broader gaming-market basis is $334.9 billion in 2025 and $374.8 billion in 2026, and does not reach $752 billion until 2033. Statista's basis is the widest, counting console, PC, mobile, cloud gaming, gaming hardware, game streaming, in-game advertising and gaming networks, and gives $577.9 billion in 2026, up 8.0% a year. So the $700 billion I recorded is not “today's games software market”; what it matches is the widest gaming-industry basis extrapolated two or three years out. It is 3.5× Newzoo's $200 billion, and the difference is exactly the hardware, streaming and advertising that never enter Newzoo's table at all. For the breakdowns below, read the denominator on that wide basis.
• Production is 20% of it, a $140 billion market; within that, labor $130 billion and server hardware $10 billion: this breakdown is the figure I recorded, there is no matching split in public sources, and it is here only for the structural proportions: over ninety percent of production is labor, hardware barely costs anything.
• Engines, around $1 billion: this one is closest to the public data. Third parties put the global game-engine market at $3.65 billion in 2025 and $4.33 billion in 2026 (Precedence Research); Unity's own engine business, Create Solutions, did $645 million in FY2025. So “on the order of a billion dollars” matches a single engine company's engine revenue, and the whole market is about three to four times that.
Line those numbers up and the shape of the game-engine business appears: an industry worth hundreds of billions of dollars, of which only a few billion actually reaches the engines (about 2% on Newzoo's $200 billion basis, under 1% on the wide gaming basis), and yet it holds 70%–80% gross margin. It collects little and earns hard; what it earns is the money of that position. As for what its AI part is actually worth, all you can do is benchmark it against public figures.
📊 Who Unity's AI part should be benchmarked against: Anthropic's ARR curve, and the valuation multiples of vertical agents
All of the figures below are as of August 2026:
Anthropic's ARR was $9 billion in 2025; by May 2026, ARR $44 billion at a $900 billion valuation
• Zhipu (GLM): 2025 revenue RMB 724 million, about $100 million (7:3, up 132% year on year), with 2026 expected at $350 million (that forecast is industry talk, with no matching public figure; the comparable public data is Zhipu disclosing in March 2026 that annual recurring revenue had reached RMB 1.7 billion, about $240 million, with API call volume up 15× in six months). The market-cap line is even wilder: HK$57.9 billion on its Hong Kong listing day in January 2026, reaching HK$880 billion in May 2026 (about $113 billion, already past Meituan and NetEase), and a high of HK$969.3 billion on 24 June, briefly crossing HK$1 trillion
• Manycore: 2025 revenue RMB 820 million (about $115 million), 96.9% of it subscription; while the spatial-intelligence business (SpatialVerse) that the capital markets had such hopes for sold only RMB 5.2 million all year, 0.6% of total revenue, with just 16 customers. So “basically no token revenue, weak ability to charge” is backed by the financials, not my bias (TMTPost, “Taking Manycore apart: an AI story that sold RMB 5.2 million in a year, what makes it worth HK$63.6 billion?”). Market cap hit HK$63.6 billion two days after its April 2026 listing, with an intraday peak equivalent to HK$83.7 billion; by 7 August 2026 it was back at the issue price, down almost eighty percent from the high (about HK$13 billion)
• Vertical agents get higher valuations still (data as of August 2026):
• Harvey (legal): ARR ~$350M, valuation in discussion ~$15.5B, about 44× PS (the previous round three months earlier was a $11B valuation on $200M raised)
• Cursor (code): ARR $2B (Feb 2026), last standalone private round at $29.3B (Nov 2025, ARR already $1B+ at the time, about 29× PS on the same basis), acquired by SpaceX in June 2026 for $60B all-stock, pending regulatory approval
• Abridge (healthcare): ARR ~$180M (this is the figure I recorded; Sacra's public figures are $60M at the end of 2024 and about $100M in May 2025), valuation ~$5.3B (Series E, June 2025), 29× PS
Thinking
Firmly bullish on EDGE AI and its applications. Tencent's choice is actually clear: weaker presence on the model side than the other majors, but on the product side it has strung IMA, Docs, WeCom and Meeting into a full suite — cultivate the user habit first, talk about model capability later. And the line I care most about, the game AI agent, is still stuck at the basic “widen the agent's audience” function; it has not reached engine level. Judging by the CAD precedent, the hard part was never automatic drafting, it is that the AI cannot get to what the client has in mind, so in the end it comes back to human skill. But game AI landing is still the ultimate paradigm for productizing AI to acquire users: Unity's Runtime Fee was in the end boycotted away by developers, and yet the fact that they dared reach back to twenty-year-old games shows the engine has pricing power over the back catalogue. Whoever can implant AI into the engine can stack an AI implementation fee on top of the engine's fixed fee. Especially since the engine is only one or two percent of the whole games industry and yet holds 70%–80% gross margin. That position was always worth money.
08 Kingsoft Office (金山办公) | Lingxi AI and the per-seat pricing method: the audience an office SaaS already has is AI's sales channel
Kingsoft Office was where I brought a good friend along and got a feel for the form AI application products take these days.
At the last AI conference, I still remember the standout AI product being the DeepSeek V3 all-in-one box. I honestly could not understand that model, especially for a 32B model: why spend tens of thousands of yuan buying compute that is not even high-end. But once I understood the To-G model and what was embedded in it for office work or for government agents, it clicked immediately.
This year, what interested me about Kingsoft Office was its office AI product, Lingxi AI: a choice of many embedded open-source large models (GLM, DeepSeek, Seedance and so on), charged by the seat. Functionally it is roughly the same as having Claude embedded in Excel to help you finish a report.
SaaS products like office software pull in a huge audience directly, and that quietly adds a sales channel for AI.
Producing materials, tidying things up, writing slides — I think the importance of those is limited. What matters more is how you take what you have understood and learned, think it through logically, express it, and make it match the point your listener's company or client actually wants. All of it lands on the same thing in the end: the way information gets conveyed keeps iterating, and it will only get faster and more efficient.
There are two paradigms for AI as a tool you use: 1. break the existing product architecture, stuff an AI tool in and keep running; 2. an AI-native one-person company, or using AI to reshape the whole architecture. Either way the end goal is to bring AI in and form a new application end or some other form, to attract customers, convert them into MAU and DAU, and complete retention and monetization.
The pricing paradigm looks like this:
(Fig. 30 Four price tiers on the subscription page: the trial edition free for a limited time with 800 “Ling points” a month, standard at ¥48/month with 5,000 points, advanced at ¥128/month with 14,000 points, and flagship at ¥398/month with 48,000 points. The only real differences between the four are two things: how many “Ling points” you get each month, and whether you can use Pro / Max mode. Cloud storage is 1 TB across the board. Screenshot from the subscription page; prices move with promotions, check the official site.)
One more note on the public figures for the enterprise side: when Lingxi Professional launched on 15 July 2026 it was still an invite-code beta and the per-seat price had not been published. The official line is that traditional subscription and usage-based billing will coexist for the long run, that enterprises with heavy usage can bring their own tokens, and that they can also buy add-on tokens through Kingsoft Office (Kingsoft Office CEO, post-event group interview, 17 July 2026). For reference, WPS consumer membership runs roughly ¥89–250 a year across three tiers (member / super member / AI member), or about ¥15–40 for a single month.
Which makes how you combine AI with an arrangement of multiple agent flows especially important. What I noticed here is that the biggest marker of AI being productized is the so-called per-seat pricing method, and the first place I felt it was Feishu.
🍿 My own path through the potholes: from fiddling with a self-hosted OpenClaw to just buying Feishu's claw cloud service
When OpenClaw first blew up I went straight at connecting it through Feishu and got my first message through. Later I saw a lot of discussion about the security of self-hosted OpenClaw deployments, and with no time to fiddle with getting several harnesses working on different ends, I went with Feishu's claw cloud service and bought a Feishu AI membership at the same time.
At first I was puzzled about why my usage was handed to me as 1,000 points a month. For an individual user that actually blurs out the details: tokens used, number of calls, volume. And you cannot know exactly which model was called. It made me uncomfortable at first.
Then I approached it as a product and it opened right up: this is one of the standards of the product process, putting complexity into a black box. The user sees one price and one allowance, and cannot see which model got swapped in underneath or how much thinking effort was turned on, so it is easy to accept and the features are selectable. Choosing among a mass of model APIs, MAX/THINKING levels, what you are trying to achieve and which agent to use becomes two membership tiers — you give people multiple choice instead of fill-in-the-blank, and it is far easier to build a business model and a first version. I have to say, it really has worked.
Back to Lingxi AI. Seeing WPS embed AI to do the work, and seeing how the seats are allocated, the office paradigm of the future is already right in front of you. Combine it with an application terminal like the Doubao Phone plus a remote-control channel and you will not need a 24-hour digital human; this will be more frightening than being on call in WeChat, and everyone's per-head output will be twice what it used to be, or more (that is my own feel; published research mostly gives efficiency gains at the level of a single task). It will of course also weed out the people who are not fluent with the tools.
(Fig. 31 The main visual at Kingsoft Office's booth, “AI-Native@Work”, with hanging banners from left to right for the WPS 365 enterprise brain, WPS Comate and the Lingxi AI-native office AGENT.)
Thinking
Pay for AI yourself → efficiency goes up → more tasks → more pressure → job space squeezed → new forms created.
Per-seat pricing is the biggest marker of AI being productized because it compresses a mass of model APIs, MAX/THINKING levels and agent choices into two membership tiers — you give people multiple choice instead of fill-in-the-blank, and only then can a business model stand up. And office SaaS comes with a huge audience of its own, which quietly hands AI a sales channel. Seeing WPS embed AI to do the work and allocate seats, the office paradigm of the future is already in front of you, and it will of course weed out the people who are not fluent with the tools. Leaving aside the cost of personal learning and stepping out of the comfort zone, every one of us learning AI is accelerating the present and iterating the technology.
09 MiniMax | The video model went viral overseas; emotional companionship carries the user count, productivity products carry the ARPU
MiniMax currently has four products: the foundation model, a Vibe Coding platform, Hailuo AI, and Xingye.
MiniMax first blew up overseas with its video model, and moved into the domestic track at the same time. At the last WAIC I was not even clear on how much reach they had. The importance of video models deserves its own chapter to analyze and learn.
They listed in Hong Kong on 9 January 2026, so for the numbers in this section I use the prospectus figures wherever the prospectus has them. As I understand it, their plan is to announce the next-generation model after M3 in September (no official statement on this yet).
M3 came out on 1 June 2026, mainly bringing the MSA mechanism (MiniMax Sparse Attention, paper arXiv:2606.13392), which shares its purpose with Kimi K3's KDA architecture: deal with the KV Cache, cut inference consumption, speed the computation up. The official figures: 1 million tokens of context, with per-token compute at maximum length only about 1/20 of the previous generation, prefill sped up more than 9× and decode more than 15×. Put another way, it turns “context length” from a cost item into a dimension you can keep pushing up.
(Fig. 32 MiniMax's official diagram of Sparse Attention: the index branch on top (Idx Q1/Q2 → Idx KV → Block Max Pool → Block score → Top-k indices), the sparse branch below (Q1–Q6 fetching only the selected KV blocks by the Top-k indices), with both branches converging at the Output Projection, with the whole thing still sitting inside a GQA attention framework.)
Unroll that diagram as a process and it is five steps:
(Fig. 33 The MSA attention flow: the index branch does the “locating”, the sparse branch does the “computing”. AI-generated.)
💡 How MSA cuts down “the tokens you have to read each time”: one index branch + one sparse branch
MSA sits on the GQA attention framework, and the flow is roughly four steps:
- do a first pass in the hidden layer, using max pooling to shrink the range attention has to look at, and generate the QKV and the index QKV;
- run the similarity computation on the index QKV, then chunk it (compressing and grouping different tokens) and produce scores;
- take Top-K over the scores of the different blocks and pick the most similar ones;
- go by the index blocks to fetch the corresponding blocks from the real KV, then do GQA attention.
Put together, it is a division of labor between two branches: the index branch finds which part should be looked at, and the sparse branch looks at the important content, which cuts down the number of tokens that have to be read each time.
The Vibe Coding platform counts as an extension built on the foundation model; competitors are QoderWork, CodeBuddy, Trae and others. There is more on how I read the MaaS market further down, so I will not go on about it here.
Then Xingye, which addresses the emotional-need market: overseas the comparison is Character.AI, and domestically the pool is Talkie + Xingye. But note the signal that they are turning — the paying ARPU on emotional companionship is far below the productivity products, and Hailuo AI's share of revenue moved up sharply within a year (figures below).
📊 The user count on emotional companionship and its ARPU are an order of magnitude apart
All domestic figures below come from MiniMax's Hong Kong prospectus, as at 30 September 2025.
- Overseas: Character.AI has 45 million active users and about $50 million of 2025 revenue, up 66% year on year (from $30 million in 2024). Watch the basis on the user figure: the Business of Apps original says “45 million active users in 2025”, without saying whether that is monthly and without a September cut-off; DemandSage, citing Similarweb, puts MAU at just over 20 million. The two are more than double apart, so I treat it as “active users” here.
- The domestic user pool: MiniMax has 212 million cumulative individual users across more than 200 countries and regions, of which Talkie + Xingye alone account for 147 million, Hailuo AI 42.35 million and the MiniMax App 19.06 million. Which is to say, the emotional-companionship line carries nearly seventy percent of the company's users. On monthly actives, what I recorded was 20.05 million; the prospectus figure is 27.62 million across all AI-native products (up 44%; 3.1 million in 2023, 19.1 million in 2024). Paying users are 1.772 million, up 262% (about 120,000 in 2023, about 650,000 in 2024).
- ARPU comparison: paying ARPPU on emotional companionship (Talkie + Xingye) is only $5; the productivity product Hailuo AI is $56, exactly 11.2× Talkie; the MiniMax App is $73. The enterprise side is another order of magnitude — open-platform and enterprise-service revenue for the first three quarters of 2025 was $15.42 million across about 2,500 paying customers (only about 100 in 2023), which works out to about $6,168 per customer, basically the $6,167 I had recorded.
- The structural shift: Hailuo AI's share of revenue went from 7.7% in 2024 to 32.6% in the first three quarters of 2025, while over the same period Talkie + Xingye fell from 63.7% to 35.1%. In the same window total company revenue went from $3.46 million in 2023 to $30.52 million in 2024 (+782%) to $53.44 million in the first three quarters of 2025 (+175%).
- Why they have to turn: the most striking pair of numbers in the prospectus is gross margin — 4.7% on the consumer side in the first three quarters of 2025 against 69.4% on the enterprise side. The users are on the consumer side and the money is on the enterprise side, and that is the entire reason they are shifting weight toward productivity and enterprise services.
My initial thought was that AI for emotional needs was bound to shine. The reasoning is that society today basically follows what Cognitive Surplus describes: most people have eight hours of free time to dispose of, and can spend it on entertainment or on creative output.
In an age of consuming, creating and sharing, emotional value is basically indispensable. Whether you go by Maslow's hierarchy or by social Darwinism, in today's swollen material desire, satisfaction beyond money is the scarcest thing there is.
Leaving aside government ethics restrictions and the need to steer minors away from risk, AI emotional-companionship software can acquire users very cheaply and, like otome games and fandom activity, automatically filters out the high-net-worth customers, with extremely high stickiness.
My own thinking: on that logic, the engineering of emotional-need AI software could take two routes.
- Fully autonomous creation: let the AI return to what it is strongest at, which is semantic interpretation and logical reasoning, and invent characters freely. (Is that similar to a text roleplay game in a tavern?)
- Imitate the mainstream otome approach: follow the leading titles like Mr Love: Queen's Choice, and use AI usage plus a wide variety of characters to draw users into spending and build the market.
Quantitative change in AI does not necessarily produce qualitative change, but with the good examples already out there and voice cloning as mature as it is now, building a perfect husband or a perfect girlfriend does not cost much.
Of course, a product has to come back to the market, and the market has to sit on top of policy. Ever since Character.AI was sued early on, with the complaint alleging its agents were involved in extremely dangerous red-line behavior involving self-harm by minors, the government has kept a hostile stance. If they do pivot, that is the crux for the emotional-companionship direction right now: whether to go private-domain or to keep the technology and switch tracks, both are things you have to negotiate with the government and with the public.
But the AI-persona part, beyond supplying emotional needs, inevitably also involves distilling a person's way of thinking.
Over Chinese New Year I tried a startup app called Elys, which is essentially a social platform with an AI twin implanted: everyone has their own twin, and it distills a simple model out of your reading and browsing preferences, the way you comment and the information you put in about yourself, then surfs on your behalf 24 hours a day. Sharing is the core of social media, and AI does the happening for us.
So Elys ends up being more of an observation post, watching the AIs brawl. If an AI twin is born that has replicated 100% of your information and habits, does the twin itself have a personality?
The hard part of video models is holding generation continuous across multiple dimensions of time. I will not go further on the technology; briefly, on the commercial reading:
Outside the foundation-model track, video is one of the two fastest tracks to monetize; the other is coding. Kuaishou's Kling AI is currently on the path to being spun off and listed independently: in June 2026 word came of a pre-IPO round at an $18 billion pre-money valuation (about RMB 122 billion), with plans to file with the Hong Kong exchange in early 2027.
What holds that valuation up is that “4×”: ARR went from $100 million to nearly $500 million in a year, four times over. The matching figures are about RMB 1.04 billion of revenue for full-year 2025 (RMB 150M, 250M, 300M and 340M across the four quarters), then over RMB 650 million in Q1 2026 alone, up 300% year on year, with 13.18 million global MAU in May and over 60 million cumulative users. Pre-money valuation divided by annualized revenue gives a Price/ARR of about 36×. That is how a video-model valuation gets validated.
(Fig. 34 Six dimensions of the Kling spin-off summarized: the spin-off and financing (pre-IPO $2–3B, valuation about $18–20B, Tencent participating, Hong Kong listing to be started within 12 months); the revenue ramp (about $140M for full-year 2025 → over RMB 650M in Q1 2026 alone); the ARR trajectory (4× in a year); the revenue mix (subscriptions contribute nearly 70%, most Q1 2026 revenue from overseas); users and ecosystem (MAU past 12 million); and the cost side (video is the heaviest compute scenario in multimodal). Every row is labelled with its date and basis.)
Obviously, video models have hit the short-drama market hard. Recently the actress Qi Wei announced she was AI-izing her own likeness to appear in short dramas, which undoubtedly deepens the hit to the traditional industry.
The short-drama market generates content out of the enormous body of Chinese-language literature (web novels above all), turning the real dopamine hit that text cannot deliver into pictures, and applying AI brings costs down by more than 70%. The public figure is harsher still: DataEye's Jujucha counts AI short-drama production cost at one tenth of a live-action one, about RMB 1.5 million for live action against under RMB 200,000 for AI, a ninety percent cut. But the savings did not stay in anyone's pocket: over the same period traffic-buying costs rose more than 100% year on year, with the whole industry spending RMB 150–170 million a day, and revenue per thousand plays fell from about RMB 60 to RMB 15–30. AI flattened the barrier on the production side, so the cost simply moved wholesale to the traffic side.
And on the To-B track, video generation applied to advertising, independent media and the rest has produced substantial results. That is one reason ByteDance is chasing so hard: Seedance came up fast behind Hailuo AI, and that too is because the market response has been outstanding.
A quick aside, a few points I revised in a conversation about AI agents:
Manus's form deserves its own note. Manus was the earliest to realize that a general-purpose large model is the final model and a general-purpose agent is the final form, and the earliest to build it as a chatbot-shaped app that actually does cowork.
- agent capabilities separated by different .md files will not differ in any fundamental way;
- what matters most for an agent is attention over the context;
- an agent arranging things itself is more effective than a human arranging several agents;
- general-purpose semantic input and understanding.
All of these are the shape of future Vibe-Coding-like products, and they are why Meta liked it. On 30 December 2025, Meta announced it was acquiring Manus's parent, Butterfly Effect Pte. Ltd., for about $2 billion, Meta's third-largest acquisition ever. Manus's ARR was about $125 million at the time (only $90 million as recently as August 2025), while its April 2025 Series B valuation had been only around $500 million, four times over in a bit more than six months. The investor list included Benchmark (leading the B round), Tencent, Sequoia China and ZhenFund. Meta said at the time that the deal brought “millions of paying users”, that it would keep operating Manus independently and fold the technology into Meta AI.
The part that follows is where this story actually lands: Butterfly Effect had already moved its headquarters from Beijing and Wuhan to Singapore in mid-2025, and the deal still did not go through. MOFCOM opened a review on 8 January 2026, the NDRC summoned the two founders in March, and on 27 April the foreign-investment security review office formally rejected it and ordered the transaction unwound, citing national security and technology-export compliance. Meta began unwinding the deal in June 2026, and the founding team is negotiating a buyback at around $1 billion. Of course, given subsidies and government policy, the political factors around walling off Chinese identity are out of scope for this piece — it is just that this one deal has drawn the line clearly enough: the product can go overseas, the corporate entity can move away, and ownership of the technology cannot.
Thinking
A general-purpose large model is the final model and a general-purpose agent is the final form — Manus was the earliest to realize it, which is how it went from a $500 million valuation to a $2 billion acquisition price in a bit over six months. And that the deal was ultimately rejected by regulators and returned intact says there is a second set of books on this track beyond the technology.
Looking back at MiniMax's four product lines, the logic is actually consistent: emotional companionship is there to build the user count and the stickiness, productivity products are there to pull the ARPU up; and outside the foundation model, video and coding are the two tracks that monetize fastest. As for where the gap finally opens up, the answer is written in those points above too — attention over the context, and the agent's ability to arrange its own tasks.
10 Alibaba | One of only two vendors with the full INFRA → chips → cloud → foundation models → agent stack
There is too much in this Alibaba section; it could be a chapter of its own.
But something unpleasant has to be said up front: Alibaba has not had a good few years, exactly as Munger said when he called it a retailer rather than a technology company. At the Daily Journal annual meeting in February 2023 his words were: “I consider Alibaba one of the worst mistakes I ever made… I got charmed by the idea of their position in the Chinese internet, and I didn't stop to realize they're still a God damn retailer.” The market has borne this out lately, mainly because:
It lost the war for the entry point. The internet war ten years ago was a fight for every user's habits, and for converting those habits into MAU, DAU and money. Alibaba had only the e-commerce platform; the biggest entry points for social and for leisure were held by ByteDance and Tencent, and now e-commerce itself is threatened by PDD, JD and others, plus new tracks like Xiaohongshu. No entry point means the biggest consumption channel never gets built. Alibaba buying Yahoo, building the Laiwang social app, buying Weibo, UC Browser, Xiami Music and Youku, all of it came out of anxiety about entry-point traffic.
And now AI has arrived.
The path of development runs through how humans interact with computers: from a person and a mouse doing traditional Search on Google or Baidu, to the mobile internet era where the interaction happens between a person's fingers and a phone; and now telling an AI in natural language can compress all of the above and become the entry point itself.
The opportunity that entry point opens up is enormous: change people's behavior, funnel them to your own platform and capture their spending on food, clothing, housing and travel. That is the super entry point, and it is the crux of the transformation. That is why Doubao and Qwen fought their Spring Festival Gala war: to seize that entry point.
That battle at Chinese New Year 2026 was fully public: the Qwen App put in RMB 3 billion (~$420M), took exclusive title sponsorship of the Spring Festival galas on the Henan, Dragon, Zhejiang and Jiangsu satellite channels and co-created content with them; ByteDance's Volcano Engine took “exclusive AI cloud partner” on the CCTV gala, supporting programme production, online interaction and streaming technology, and Doubao ran three rounds of prize draws across the gala. Over the same window Tencent's RMB 1 billion and Baidu's RMB 500 million red-packet campaigns were all pressed onto the same entry point. (By the same logic, the Doubao Phone blowing up and getting pushed back on by everyone else follows the same reasoning; the details will be in Part II, on AI combined with hardware.)
To seize that entry point and rewrite the rules of the old internet war, accumulated technology is essential, and Alibaba's three strongest weapons (Alibaba Cloud, the Tongyi Lab and T-Head) give it the best possible conditions for transforming and fighting:
Alibaba Cloud
Start with cloud among the three. Understand Alibaba Cloud's products and you understand cloud-computing architecture better.
Having used products on the Alibaba Cloud platform, the strongest impression is how complex and how complete it is. There is so much of it that you are not quite sure which one to pick.
This part goes in this order: Alibaba's architecture and IDC first, then cloud, large models and the agentic community, and finally Alibaba's products.
The conclusion first: agents and large models are, for the moment, converted to serve the goal of selling more cloud, the same purpose as selling IaaS — all of it is to hold on to the SaaS and PaaS layers, where the best profit is.
Those two steps went by fast, so let me fill them in. Profit in the hardware industry is visible and fixed: what a machine sells for, how many points of gross margin, laid out flat, that is what it is. Software is different. Software blurs that part, and it compounds — the deeper the same capability is embedded and the more customers it is copied to, the lower the marginal cost and the higher the stickiness, and what you can extract from it snowballs.
So selling agents and selling large models is, in the short run, selling models; in the medium run it is pulling customers onto the cloud. And once they are on the cloud, what can actually be charged for over the long run, and charged more and more for, is precisely those two layers, SaaS and PaaS, the “higher up, higher profit” end of the table below.
The opening of the MaaS battlefield has also moved everyone's attention onto grabbing customers on that track, and the upstream and downstream effects of it should not be underestimated. Alibaba happens to be China's strongest cloud vendor, and one of only two companies with all five layers: INFRA, chips, cloud, large models and MaaS, agentic. The other is Google.
(Fig. 35 Artificial Analysis's map of players across the AI value chain: four layers down the side: applications, in-house foundation models, in-house cloud inference and accelerator hardware, with darker colour meaning a stronger position. In the whole table, only two columns are dark on all four layers and boxed with a dashed line: Google and Alibaba. Source: Artificial Analysis.)
Only with a full-stack architecture do you have the right to choose. That is one reason Alibaba can make its large models this complete too: from the smallest model to the general-purpose one, open weights and closed, video, audio, text and image, all of it.
Now Alibaba's IDC situation. Founded on Wang Jian's work, Alibaba Cloud counts as the developer and pioneer of cloud computing in China, and it also has the largest installed base. Alibaba Cloud's architecture and its IDC construction matter a great deal to the whole industry.
Alibaba's IDC build speed and scale keep pace with its global expansion: 29 regions and 92 availability zones worldwide (Alibaba Cloud's published figures), of which 60 machine rooms sit across 15 provinces in China.
The IDC architecture has reached the CUBE DC 5.0 generation, with PUE at or under 1.1 across the board (the published figures are air-cooled PUE ≤1.15 and liquid-cooled PUE ≤1.10): a cooling design with a common source for air and liquid, with the electrical and mechanical turned into a modular-cabin productized delivery. The overall modularization rate across the five systems (power, cooling, security, intelligence and fire protection) went from 30% in the early years to 90%, the delivery cycle was compressed to 100 days and total cost came down more than another 10%. By capacity tier, 5.0A is the 31/34.4 MW class and 5.0P is the 105 MW class.
The core of cloud computing is also shifting CAPEX into OPEX, plus elasticity and not having to build it yourself. You can read it straight off the definition: highly elastic, bought on demand, highly reliable, visible and quantifiable returns, all of it solving direct needs for the owner.
There are many reasons to go to the cloud, and just as many services to match.
IaaS / PaaS / SaaS: which layer you want to manage decides where the profit is
The traditional IaaS, PaaS and SaaS cloud business is about which layer the owner wants to manage. The higher up you go, the higher the profit.
| Layer | What is sold | Matching Alibaba Cloud product |
|---|---|---|
| IaaS (Infrastructure as a Service) | Sells resources | ECS compute, OSS/NAS storage, VPC networking, DDoS protection and so on |
| PaaS (Platform as a Service) | Sells product capability | ACK containers, the Yaochi database, big-data compute, MSE middleware, and things like Oracle and ServiceNow |
| SaaS (Software as a Service) | Sells the result directly | The PAI machine-learning inference platform, the Bailian platform, Qwen, plus Lingyang and DingTalk |
Alibaba Cloud is very strong in the traditional cloud market today, and it is also first in the To-B business on the AI-cloud market as a whole. But note that being first on the overall market and being first on call volume are not the same company — first on call volume is still ByteDance.
📊 First in traditional cloud, first on the AI cloud market as a whole, but first on call volume is ByteDance
- Traditional IaaS: per Gartner's Market Share: IaaS, Worldwide, 2025 (published April 2026), Alibaba Cloud is first in mainland China with 32.8%, revenue up 34.4% year on year; it is the largest IaaS in Asia-Pacific with 22.5% share; and fourth globally with 7.7%.
- The AI cloud market as a whole: Alibaba is first in the To-B business, with 38.1% share per Omdia, more than the second through fourth combined. But note what that basis covers: the “China AI cloud market”, meaning AI IaaS plus MaaS together (a total pool of RMB 56.7 billion in 2025, 69% of it IaaS and 31% MaaS), not MaaS revenue share on its own. Volcano, breaking through with Doubao and Seedance, has done better on the consumer market, and Trae as a coding platform got people using it earlier. Switch to the model-call-volume basis and it is an entirely different story: per IDC, model calls on China's public cloud in 2025 came to 1,944 trillion tokens, with Volcano Engine first at 49.5%, Alibaba Cloud second at 28% and Baidu AI Cloud third at 10%.
- The size of the pool: the MaaS market is exploding. On IDC's basis, China's public-cloud MaaS market already reached RMB 30.7 billion in 2025, with about RMB 186 billion forecast for 2026 (IDC itself flags that this is a high-growth scenario, contingent on multimodal maturing, agents landing at scale, and compute supply and compliance staying stable). For comparison, IDC's forecast in the October 2024 edition only dared put RMB 250 million in the first half of 2024 and RMB 3.8 billion in 2028 at a 64.8% CAGR, overturned by reality in eighteen months. That contrast says more than any growth rate.
- The financials: FY2026 Q4 Cloud Intelligence Group revenue was about RMB 41.626 billion, up 38% (external commercial revenue up 40% within that; do not mix the two bases). AI-related revenue was about RMB 8.971 billion for the quarter, the eleventh consecutive quarter of triple-digit year-on-year growth, and passed 30% of external commercial revenue for the first time.
The reason I went and dug through job postings is that how a cloud vendor cuts the market shows up more directly on the careers page than on any slide. The three screenshots below are all publicly searchable results on the official careers site, searched in July 2026: Alibaba Cloud's experienced-hire listings have an MTE track for AI sales, with very specific verticals: government, finance, gaming, MNC, SNB, overseas expansion, embodied intelligence, automotive and more.
(Fig. 36 Searching the careers site for “AI sales”: the result bar shows 17 open positions. Source: [7])
(Fig. 37 The same site searched for “MTE”: 10 open positions. Source: [7])
(Fig. 38 The same site searched for “business technology engineer”: 73 open positions. Same set of verticals, one different keyword, seven times the count. Source: [7])
Line the three searches up: 73 business technology engineer roles, 17 AI sales roles, 10 MTE roles. The vertical coverage matches the list above; in numbers, the pre-sales technical track far outweighs the sales track, and MTE as a standalone keyword is actually the smallest. The careers page loads dynamically, so the numbers are as of the moment of the screenshot. For example, the “Alibaba Cloud Intelligence – Business Technology Engineer (Gaming) – Shanghai” posting[7].
Harness and Agent: what the MaaS platforms are fighting over
As of the day of writing, Tencent's WorkBuddy has emerged from nowhere on the MaaS platforms, pulling customers in on free call allowances and a very strong ecosystem.
Since I do not use the three big vendors' products much, my impressions here are taken from other people's. Generally the experience of Vibe Coding depends on differences in harness engineering, things like PLAN mode and global memory.
The harness itself is software engineering, and what engineering tests is the design framework of the system. Citing “Understanding Harness Engineering: the six-layer architecture, context management, and front-line team practice”[8]:
Agent = Harness + Model, where the harness's six layers are: Agent.md — MCP tools — task orchestration — memory management — evaluation and verification — global constraints.
(Fig. 39 The diagram that splits an Agent in two: the orange Model in the middle is the “CPU”, the source of capability; the ring around it (system prompt, tool calls, file system, sandbox environment, memory system, context management, orchestration logic, hook middleware, feedback loop and constraint mechanisms) is collectively the Harness, i.e. the “OS”. Source: [8])
(Fig. 40 Three nested layers: Prompt Engineering ⊂ Context Engineering ⊂ Harness Engineering, solving expression, information and execution respectively; the line underneath, “the model sets the ceiling, the harness sets the floor”, can serve as the footnote to this whole section. Source: [8])
(Fig. 41 Analysys, “Visits to desktop AI-native office agent platforms in China”, June 2026: WorkBuddy leads by a distance at 20.97 million monthly visits, followed by the domestic TRAE IDE at 12.79 million, QoderWork 7.88 million, CodeBuddy 6.56 million and QClaw 3.54 million, with a cliff below that. “Emerged from nowhere” has numbers behind it. Source: Analysys.)
Before getting into the Tongyi Lab's models, one thing first. Faced with a dizzying array of AI, there is a way of positioning what AI can do for you that works like choosing a cloud service, and it comes down to what you need. AI has to be defined: find which step of the problem needs solving first, then decide what to hand the AI:
- Question + context: traditional RAG handles it, and this is the layer most people use, Doubao / ChatGPT / DeepSeek;
- Material + the boundary of the material: the traditional software from the first layer can answer this too when the material is there, plus Perplexity;
- Task + clear steps and a path: Vibe Coding platforms, where you formally start calling agents, Manus / WorkBuddy / Trae and the like;
- Project + development environment: pushing a project forward continuously, organized around the material, Cursor, Claude Code, Codex;
- System + complex environments and task handover: multiple systems, multiple projects, with OpenClaw and Hermess taking over the process autonomously and implanting memory.
Those five tiers ask, from the user's side, “how far do you want AI to take you”. Switch to the enterprise side and the same question has to be asked twice. The first time is about how you move what you already have — going to the cloud itself comes in four paradigms: stay put = Retain / lift-and-shift = Rehost / optimize = Replatform / rebuild = Refactor. The further right you go, the more you move, the bigger the gain and the bigger the cost.
The second time asks how far AI gets rolled out once the move is done. The To-B MaaS paradigm generally runs three tiers, L1–L2–L3:
- L1: POC (a pilot project) and MVP (minimum viable product), evaluating short-term ROI and standing up a framework to verify feasibility;
- L2: business expansion, rolling AI out to 30% of core business processes, replicating L1 laterally, optimizing it deeply and adapting it to different scenarios;
- L3: building the organization and technical architecture around AI, integrating it into strategic planning, innovating products on top of AI and validating in stages: market validation, product iteration, commercialization, scaled conversion; and at the same time building an API and MCP marketplace, enriching the usage patterns, running quantifiable and monitored reviews of the AI itself, and reallocating flexibly.
The three classifications are three cross-sections of the same question: the individual asks “what do I hand to AI”, the enterprise asks first “how far do I move” and then “how far do I roll out”.
Of course it is still a land grab right now: different large models and MaaS platforms are prying MNC (broadly, enterprise) customers away with deeper discounts to expand the market, which itself proves the MaaS market is early. And traditional BTE sales, following what MaaS products require, have started carrying MTE targets alongside their own.
As for how AI applications themselves get classified, I have seen an AI-CAF four-quadrant chart: the horizontal axis is “efficiency gain ↔ product innovation”, the vertical is “internal application (facing employees and processes) ↔ external business application (facing customers and the market)”, cutting out four types: service optimization, product innovation, internal efficiency, capability evolution. It maps onto concrete scenarios neatly too: intelligent customer service and marketing assistants top left; AI SaaS, AI hardware and embodied intelligence top right; process automation and intelligent legal bottom left; intelligent business analytics and product demand forecasting bottom right.
(Fig. 42 The AI-CAF four quadrants: efficiency gain ↔ product innovation on the horizontal, internal application ↔ external business application on the vertical, giving service optimization, product innovation, internal efficiency and capability evolution, each with seven or eight typical scenarios listed underneath. I am reproducing this chart second-hand; the original source is not marked.)
The Tongyi Lab
The second weapon is the Tongyi Lab, which in product terms means Qwen.
Now the foundation-model tier. There are as many as 51 Qwen models, and the strongest of them, Qwen3.8 Max, is right up front.
Alibaba's models also cover every domain: video models, audio models, coding models, small flash models, general-purpose large models, open weights and closed, the widest coverage of anyone, with benchmark scores near the top too. The three screenshots below are all from Artificial Analysis, taken on the day of writing (August 2026):
(Fig. 43 Artificial Analysis's “intelligence index vs. price” scatter filtered to Alibaba alone: all 51 models on one chart, with Qwen3.8 Max standing on its own in the top right at about 57 points and roughly $1 per million tokens, while the green “optimal quadrant” and the dashed Pareto frontier string Qwen3.7 Plus, 3.6 Plus and 3.6 Max Preview into a line. One vendor covering an entire price-performance curve by itself.)
(Fig. 44 The same thing across all vendors: the top half is the same intelligence-index vs. unit-price scatter, with Qwen3.8 Max in the $0.90 per million tokens band at an index around 53; the top nine on the board below are Claude Opus 5, Claude Fable 5, GPT-5.6 Sol, Kimi K3, GPT-5.6 Terra, GPT-5.5, Grok 4.5 and Claude Opus 4.7.)
(Fig. 45 Artificial Analysis's full independent-evaluation table: the top of the screen has bar charts for GDPval-AA v2, τ²-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR; the lower half (AA-Briefcase, AutomationBench-AA, Harvey LAB-AA, IFBench and so on) only caught the titles, the charts had not loaded, hence the large blank area. Reading the individual rankings is far more useful than reading one composite score.)
Qwen3.8 Max (released 3 August 2026, still the newest as of the day of writing) has 2.4 trillion total parameters (2.4T), 95B active per pass, a 1-million-token context window, and native visual understanding.
I learned about the architectural changes from the Zartbot public account[9]; the technical conclusion is this: overall, Qwen3.8's ability to execute autonomously with no additional prompting is on par with Opus 4.8. What keeps it from reaching Fable 5's level is that it has not yet learned to use OR-Tools.
While we are here, what OR-Tools is: Google OR-Tools is an open-source operations-research optimization toolbox, built for exactly the kind of problem where you “pick the best solution out of a vast combinatorial space under a pile of hard constraints” — rostering, bin packing, route planning, shop-floor scheduling. Its main engine, CP-SAT, is a constraint-programming solver you can call straight from Python. Why is “can it use this” a watershed? Because combinatorial optimization is precisely what large models are worst at: they write the next token by probability, they do not backtrack and they do not prove optimality, so once the problem gets big all they can do is guess a plausible answer, and the more confident they sound the further off they are. A strong model does not brute-force it; it recognizes “this is a combinatorial optimization problem”, translates the variables, constraints and objective into a model the solver can read, hands it to CP-SAT, and confines itself to modelling and checking the result.
On the foundation-model fight, we are still in the land-grab phase, and it has a very strong long tail. SemiAnalysis's TokenBudgeting gives the distribution as: a median of just $136 per person per year, $7,300 at the 90th percentile, $90,000 at the 99th — $1,430 lands in the 75th–90th percentile.
(Fig. 46 A comparison table flattening out “what AI actually costs”: Qoder CN enterprise edition at about ¥99–199 per seat per month, GitHub Copilot Business/Enterprise at $19/$39, M365 Copilot and Cursor Business at $30/$40, heavy enterprise use of Claude Code actually paying $150–250 a month, and Qwen's token prices on Bailian. The lower half has three big enterprise-side numbers: median AI spend of $136 per person per year; agents cutting the cost of repetitive functions 40–70% with ROI payback under 3 months; and Gartner's figure of $2.52 trillion in global AI spend in 2026, up 44%. Every row is sourced.)
And the profit flows to the foundation-model layer, not the hardware layer. Take company A: Anthropic's annualized revenue went from $9B to $44B+ (the $44B ARR figure is as of May 2026), with gross margin rising from 38% to 70% (that is SemiAnalysis's third-party estimate of inference gross margin; Anthropic has never confirmed it). The incremental profit flows disproportionately to the model labs, not to the hardware layer.
Blackwell's token output per card is 30× Hopper's a year earlier, and VR NVL72 is up to 32× H100 throughput on FP4. Per the calculations in the SemiAnalysis report, VR NVL72's cost-based pricing floor is about $4.92 per GPU-hour and its value-based pricing ceiling $9.63–12.25 an hour. NVIDIA clearly has plenty of room to raise prices and has not used it up, which is a separation between pricing power and the act of pricing.
The margin it gives up buys speed of ecosystem expansion and a safety margin against antitrust. In effect it banks “invisible profit” into future market share.
But the foundation-model layer is the most fragile one for gross margin. It is built on holding frontier capability exclusively, with no process node, no capacity and no physical moat, and the moment open-source models close in, the profit falls back faster than it climbed.
T-Head: Lingjun M890, plus some gossip
The third weapon is T-Head. Before the chips, a piece of gossip.
🍿 On 7.23, overseas model subscriptions on the second-hand platform went to zero overnight
Here is a piece of gossip. On 7.23 I did what I always do, opening the second-hand platform hoping to buy a topped-up Claude5X subscription. To my considerable surprise and shock, after asking more than thirty sellers and hearing they were all out of stock, my thinking changed. Word was that a certain big internet company had come through sweeping up everything — one line, we'll take however much you have. The sellers guessed it was to build a resource pool and reverse-proxy it as a pool. All I can say is, some spending.
I have kept thinking about this. When supply tightens, the first tier to go dry is always us retail buyers, earlier than any earnings report or research note.
It was bad enough that the US-China fight buried me on stocks. Why does the large-model war have to hit those of us scraping by, going to all this trouble just to get our hands on Claude. On a day without Claude I do not want to go to work at all…
The two screenshots below are the chat list from that day. Seller nicknames, message content and quotes have all been redacted; they are kept only for the impression of a screen full of Claude and every one of them out of stock.
(Fig. 47 The message list that afternoon: an entire column of Claude-related product thumbnails down the right, and an unbroken run of “out of stock for now” on the left. Seller nicknames redacted.)
(Fig. 48 Scrolling further down the same list, the composition is unchanged — “sorry, none right now”, “apologies”, “cleared out yesterday”. This one screen is every reply I got that day. Seller nicknames redacted.)
This showing had the most famous T-Head, the most practical AI recording card from DingTalk (cheaper than on PDD, even), the Bailian coding platform, Yaochi storage from the traditional cloud side, the now open-sourced 3.8 Max (post-update), QoderWork and more. There was genuinely too much to take in.
💡 Why a future POD node is most likely 64–128 cards (Lingjun M890 + scale-up/out + the DeepSeek V4 bandwidth calculation)
The Lingjun Zhenwu M890 supernode, capable of serving inference on a 10T MoE model:
| Metric | Figure |
|---|---|
| FP16 compute | 0.36 PFLOPS |
| FP8 compute | 0.722 PFLOPS (inferred, slightly under the 950) |
| HBM capacity | 144 GB |
| Interconnect bandwidth | 800 GB/s, 400 GB/s unidirectional = 3.2 Tb/s |
| Power | 400 W |
| Benchmarked against H200 | 80%–100% |
Four items in that table (FP16 compute, FP8 compute, power and “benchmarked against H200”) come from conversations on the floor; there are no official figures for them. The FP16 number I recorded is 0.36 PFLOPS, while the only third party to have given a specific figure writes 0.6 PFLOPS; the two are on different bases. On the benchmark, what I heard on the floor was H200, whereas the official line and both Chinese and foreign media say H20 (the previous-generation Zhenwu 810E was described as “comparable in overall performance to the H20”, and the M890 claims 3× the 810E). The officially published specs are: 144 GB HBM per card (up 50% from the 810E's 96 GB), 800 GB/s die-to-die interconnect (up from 700 GB/s), native support for full precision from FP32 down to FP4, 3× the performance of the 810E, and 128 chips integrated per rack; on the supernode side, 64 cards out of the box with 9 TB of memory, officially said to support MoE inference at the ten-trillion-parameter level, and already running the 2.4-trillion-parameter Qwen3.8. Converting: 400 GB/s × 8 = 3.2 Tb/s.
Now GPU scale-up (vertical) and scale-out (horizontal):
- Scale-up: interconnecting GPUs inside a single compute POD, i.e. making sure the GPU chips work together efficiently, on technologies like NVLink and UCIe;
- Scale-out: interconnecting multiple PODs into a compute cluster, on topology and protocol optimization, based on RoCE, RDMA and the like.
NVIDIA's DGX SuperPOD[10], for example, uses this approach to achieve coordinated computation across a multi-node GPU cluster. Mainstream AI infrastructure today commonly uses a hybrid architecture of “scale-up within the node, scale-out between nodes”, which both guarantees efficient computation inside a single node and supports very large-scale distributed training. (Citing “A complete breakdown of GPU scale-up/scale-out: architectural principles, scenario selection and industrial-grade deployment”[11])
(Fig. 49 An item-by-item comparison of scale-up and scale-out: use cases, architectural characteristics, hardware selection (NVLink/UCIe vs. InfiniBand/RoCE), software frameworks (NCCL vs. PyTorch Distributed/Megatron-LM), key performance indicators, cost and deployment threshold, and typical cases. The dividing line between 8/16/32 cards in a single node and hundreds or thousands in a cluster is laid out most clearly in this table. Source: [11])
Per the conclusion cited by the Zartbot blogger: over the next two to three years, 128-card scale-up is already enough. What the DeepSeek V4 paper says is the other side of the same coin:
(Fig. 50 The compute-to-bandwidth ratio ceiling given in the DeepSeek V4 paper: C/B ≤ 2d = 6,144 FLOPs/Byte. Source: the DeepSeek V4 paper.)
Which is to say, every GB/s of bandwidth can support at most 6.1 TFLOPs of compute before bandwidth becomes the constraint. Run it backwards: around 800 GB/s of scale-up bandwidth pairs with compute approaching 5 PFLOPS, which is already near the theoretical limit. So a future POD node is most likely 64–128 cards.
An aside: DingTalk Talk and AI hardware
DingTalk Talk has been popular for a while. Weighing it up on weight, number of microphones, transcription accuracy and how well the templates fit, I hesitated a long time and ended up buying a PLAUD, then worked with Claude to refine the workflow for extracting the content, with a plan to feed the results into Obsidian as notes. You can see the AI chain coming; more in chapter two, “Thinking in the age of AI”.
One gripe: we are at the early stage of product application, and fitting each usage pattern still takes deep fiddling on your own part. Very geeky, and also far too inconvenient for the user.
(Fig. 51 My own prompt-template library for recording transcription, split into cards along a learning/cognition line: L1 rapid industry-orientation cards, L2 AI/DC sparring and fact-checking, L3 distilling the thrust of a talk or a lecture. The body text on the cards is redacted; only the structural layer of the templates is left: how many tiers there are and what each tier solves.)
This kind of product is actually the best attempt yet at converting large-model applications into profit at the hardware end, with absurdly high margins (that is a feel, with no public teardown cost or vendor gross margin to check it against).
Underneath, it plugs into various ASR (speech-to-text) models and turns the hardware's recording function into text. It could in fact hand you the result directly at the next step, but for the sake of productization it inserts one more step into each vendor's own recording platform, say Feishu's minutes or DingTalk's AI transcription. The purpose is the same next step in both cases: charging for AI large-model software and acquiring users. For a product like this it kills several birds at once: it keeps the margin, it adds a customer-acquisition track, and it charges a second time on the large-model fee.
In the age of AI, how you combine with hardware to change people's lives and habits is the final purpose of an AI product.
At the same time, AI minutes let us review content more logically. Attention is still the most precious resource there is, and for each of us AI minutes will replace the part of meeting note-taking that needs no brain, saving the brain cells. But it also means you need dedicated time to run post-training (AFTER TRAIN) and fine-tuning (SFT) on your own ability to summarize and express logically.
Thinking
Losing the war for the entry point is where everything in this section starts: with no entry point for social or leisure, the biggest consumption channel never gets built. AI reshuffled the entry point, and the three weapons are what buy Alibaba a seat at the table.
Only with a full-stack architecture do you have the right to choose — that is the thread running through this whole Alibaba section. Selling agents and selling large models are, in the short run, both about selling more cloud, about holding on to the profit in those two layers, SaaS and PaaS. Software is worth defending more than hardware precisely because hardware profit is visible and fixed, while software blurs it and compounds it.
But the same thread also explains where it is fragile: the profit really is moving toward the foundation-model layer, and foundation-model gross margin is the most fragile layer there is, with no process node, no capacity and no physical moat. Once open source closes in, the profit will fall back faster than it moved up.
At the individual end the conclusion is the same: in the age of AI, how you combine with hardware to change people's lives and habits is the final purpose of an AI product. AI minutes can save you the brain cells you did not need to spend, but the time you save has to be spent back on yourself — one round of post-training on your own ability to summarize and express logically.
11 Moonshot AI / Kimi (月之暗面) | An open-source 3T super-size model: top specs, top price
Kimi put on a show the day before the exhibition opened, carrying on the shock of GLM-5.2. The release of K3 blew apart that distillation fiction from the White House science adviser and Anthropic in one stroke: Anthropic's Fable only became publicly available on 1 July, and K3 shipped two weeks later on 16 July, with several third-party experts saying flatly that pure distillation is simply not doable on that timeline. An open-source 3T super-size model shoved straight in their mouths — is open source supposed to be distilling closed source now? That is also the funniest joke I have heard lately.
But the cost to the user is just as considerable. K3 may be the most expensive general-purpose large model in China. I tried it on the OpenCode GO plan and burned 5 hours' worth of allowance in 3 minutes. Being broke is a me problem.
(Fig. 52 Same plan, same day: on 2 August, kimi-k3 $3.35 and deepseek-v4-pro $0.33. Same work, a full order of magnitude apart.)
(Fig. 53 The usage bars after the run: rolling usage at 100% with 1 hour 16 minutes until reset, while weekly usage is only 49% and monthly 45%. What throttles you is the rolling window, not the monthly allowance.)
On raw capability, K3 beats Opus 4.8 across the board on the six agentic boards Moonshot put out itself (Terminal-Bench 2.1 88.3 vs 84.6, FrontierSWE 81.2 vs 66.7, DeepSWE 67.5 vs 59.0, GDPval-AA Elo 1668 vs 1600…), and its Artificial Analysis intelligence index of 57 puts it third in the world, in the same band as Opus 4.8 and GPT-5.5. But Opus still carries SWE-bench Verified 88.6%, SWE-bench Pro 69.2% and USAMO 2026 96.7%, which K3 has not matched, and the two harnesses are different. So “won the arm-wrestle” is more accurate than “crushed it”.
KDA + AttnRes compressed the KV Cache 12.5×: of K3's 93 layers, 69 are KDA (linear attention with a fixed-size recurrent state that does not grow with sequence length) and 24 are Gated MLA (keeping a growable compressed KV cache). The reason the KV Cache compression ratio directly determines token cost is that every single call involves reading that data.
Kimi K3 optimized three things this time: 1. parameter count; 2. context; 3. model depth. What follows is compiled from the Bilibili creator “Algorithm Magician”'s video “Dissecting Kimi K3's three core architectures: understanding the next generation of large models”.
💡 K3's three architectures: bigger, longer, deeper, and what each was traded for
How a large model keeps getting bigger: on a MoE expert architecture, apply Latent MoE, adding a compressed space to deal with load imbalance and poor communication. Compared with NVIDIA Nemotron 3 Super, the idea is more experts and a lower activation ratio.
How context keeps getting longer: KDA, compressing historical information into a dynamic note that is updated and computed in real time, at the cost of possibly losing some information; in practice MLA and KDA are interleaved.
How the network gets deeper: Kimi Attention Residuals, replacing the Transformer's residual connection, a fixed unit-weight sum, with a learnable softmax depth attention. Compared with the old residual connection (the previous layer's output is the next layer's input), that old form ignores the contribution of shallow layers; and most of a Transformer can in fact be pruned with very little loss.
The result: cross-stage caching + two-stage inference; training overhead under 4%, inference latency under 2%.
The three together complete the RNN-to-Transformer shift in the sequence dimension. My own reading: the RNN side compresses history into one context, while the Transformer still chooses to access every historical position selectively.
2.8T parameters (2.78T total and 104.2B active, to be exact), 1M of context. Kimi has always been known for long context, and 1M is enough to feed in an entire novel. The importance of context in the AI era goes without saying, and it is the key thing that separates Kimi from the other large models: it held the 1M line on context while the parameter count leapt from the B era into the T era.
The deployment side is worth studying too: I saw someone deploy K3 straight onto 64 B200s (Threads @3cpj, July 2026).
💡 Three tiers of the bill for running K3 on 64 B200s: gets it running / production-ready / actually pleasant
| Tier | Hardware cost | Notes |
|---|---|---|
| Should get it running | RMB 8 million (~$1.1M) | Barely gets it up |
| Production grade | RMB 15 million (~$2.1M) | Can serve external traffic |
| Actually pleasant | RMB 20 million and up (~$2.8M+) | Electricity not counted yet |
And that is already the level before you count the electricity. That order of magnitude matches public reporting: the 64-card tier alone is over $2.4 million to buy (RMB 17 million-plus), with total system draw around 45 kW, which is already machine-room-scale power (unwire.hk, 21 July 2026).
If the tuning is done well, 80 RTX 5090s can take it: 80 × 32 GB gives 2,560 GB of VRAM, hardware cost $200,000–250,000 (RMB 1.5–1.8 million), electricity a little over $10,000 a month, and 20 tokens/s on a single stream on day one before tuning (Kocpc, July 2026). The best standard would be 22 × H100 — which is also the minimum configuration for native MXFP4 weights; at $25,000–30,000 per card, a 22-card cluster is about $550,000–700,000.
(Fig. 54 Another bill of materials for a “minimum usable K3 inference cluster” (FP4 precision, short conversations): 20 × H100 80GB at $600,000, two DGX H100 systems at $400,000, InfiniBand NDR networking at $50,000, NVMe storage at $8,000, rack-level liquid cooling and power at $30,000, $1.088 million in total. Note that it counts servers, networking and racks, so it is not on the same basis as the three-tier bill above, which only gives a total.)
(Fig. 55 Cost per million tokens, self-hosted versus calling the API: Kimi's official API $0.50, self-hosted on 22 × H100 about $0.12 and on 8 × H100 (TP+EP) about $0.15, against Claude Opus 4.8 API at $15 and GPT-5.6 API at $10; the right-hand column converts these into a monthly bill at 1 billion tokens. The original table notes it is an estimate based on public-cloud GPU pricing.)
Now the renting side. K3's official API: input $3 per million tokens (cache hit $0.3), output $15 per million, 1M context (Moonshot official pricing page, checked August 2026). Artificial Analysis lists 10 third-party providers, with up to about 1.8× spread between them and peak throughput around 174 tokens/s. The reference price on the self-hosting side: B200 on-demand hourly rates have a median of about $6.1 per card-hour (spot as low as about $3.2; AWS/GCP want $14–16), which puts monthly rental of 64 B200s roughly in the $150,000–280,000 range (converting from the spot floor to the on-demand median at either end; figures taken August 2026).
(Fig. 56 Artificial Analysis's on-demand hourly GPU rates by cloud: in the B200 band, AWS $14.2, CoreWeave $9.6, Lambda $6.7, Nebius $6.5, RunPod $5.9; in the H100 band anywhere from $3.0 to $12.3. The same card is three or four times apart across clouds; the “on-demand median about $6.1” in the text comes from this table. Source: Artificial Analysis.)
The full hardware requirements for local deployment are in OpenModelMap's Kimi K3 Hardware Requirements — GPU, VRAM & Server Config[12].
Kimi's booth was extremely simple. A large-model company being this plain: just a few computers set out, chatting with whoever came by, with the prizes co-provided by Zhihu and others. Most people there were not from a model-technology background either, so the conversation could only stay general.
Thinking
A large-model company's booth can be plain enough to be a few computers and some general conversation with whoever turns up — because what it actually sells was never on the booth.
K3 pushed all three of parameter count, context and model depth up a notch at once, and long context is still what separates Kimi from the other large models: 1M of context swallows an entire novel in one go, while the parameter count has crossed into the T era. The cost is written on the bill — it is the most expensive general-purpose large model in China, and the OpenCode GO plan burned 5 hours' worth of allowance in 3 minutes. Being broke is a me problem.
12 Fudan MOSS | An open-source matrix you can run locally at 16 billion parameters; the most interesting part is real-time video understanding
MOSS counts as a large-model startup I ran into by coincidence, out of Fudan's NLP group (repo description: OpenMOSS/MOSS: An open-source tool-augmented conversational language model from Fudan University).
It is very small in scale, which suits local deployment well. Take the moss-moon series open-sourced in April 2023: the model has 16 billion parameters, runs on a single A100/A800 or two 3090s at FP16, and on a single 3090 at INT4/8.
What I found most interesting on the floor was real-time video recognition; it had been a long time since I came across this kind of application.
Seeing that real-time video understanding line (the name I recorded from the booth panel was MOSS-Video-and-Audio; in the official repo it should correspond to MOSS-Video-Preview / MOSS-VL-Realtime, released April 2026, with MOSS-Audio and MOVA in the same string alongside it), the first thing I thought of was Ant's Robbyant, and the VLA work happening on the embodied-intelligence side. The purpose is the same: recognize images and audio and produce understanding. The VLA route is LLM → VLM (Vision-Language Model) → VLA (Vision-Language-Action Model).
(Fig. 57 The MOSS model wall at the booth: the speech line runs from MOSS-TTSD and MOSS-TTS-v1.5 down to MOSS-TTS-Nano (0.1B, runs on CPU alone), then the MOSS-Transcribe and Diarize series (long-audio transcription and speaker separation, claimed first among open source on the OpenASR Leaderboard), then VoiceGenerator, MOVA, MOSS-SoundEffect, MOSS-Audio and MOSS-Music, with MOSS-VL-Realtime only in the last box. One wall holding the entire open-source matrix.)
The video-plus-audio line should not be the same thing as Vision-Action; it counts as a weakened VA. But at the application end I already find it interesting: the first product that came to mind was digital-human reactions, delivered the VA way.
Trying it in person felt remarkable. VA can now recognize in real time what the person in front of the camera is doing and give feedback; at the same time a small-parameter model can be implanted into multiple terminals and still performs well across different video recognition tasks. Excellent.
(Fig. 58 The demo unit for real-time video understanding: on the left, what the camera is seeing; on the right, the model's text understanding in sync, with latency low enough that it picks up the moment you raise a hand. The bystander on the right is redacted.)
Above all, the name MOSS immediately makes you think of the robot in The Wandering Earth — expecting humans to stay rational forever is asking too much.
Thinking
Seeing an open-source matrix of models pointing in this many directions, I think building MOSS looks easy enough now; it is only a matter of setting the boundaries and the functional direction. The hard part was never building it, it is where you draw the line.
Breaking through the barriers of traditional physics and mathematics with large models is within reach, MAYBE.
H1 wrap-up · Looking back at the model layer from INFRA
Having walked the whole of H1, the strongest thing I felt is this: on one value chain, every layer is anxious about something completely different.
At the infrastructure layer, the anxiety is delivery. Technology stopped being the barrier long ago: the depth of liquid-cooling R&D cannot hold a barrier up, and anyone can buy their way in. What actually blocks you is the certification stamps, the product form factor and engineering delivery; it is having to negotiate for that plot of land overseas a year ahead; it is that piece of paper from CE/UL; it is labor abroad costing three times as much at half the efficiency. All of it is supply-chain business, and not one bit of it happens in a lab. So Vertiv turns the machine room into a container and compresses an overseas year of delivery into 5 months.
Cloud and engines are anxious about position. That DaoCloud product person said there is still a long road to breaking through CUDA. It sounds modest; it is actually precise. CUDA sits above the operating system and below the model, stacked layer by layer out of more than a decade of software ecosystem. You cannot go around it and you cannot break it. What decides the outcome is how much code other people have accumulated on your platform; how much you write yourself matters less.
The model layer's anxiety is staying fresh. That set of SemiAnalysis numbers says it plainly: incremental profit skips over the hardware layer and flows disproportionately into the model labs. But it is also the most fragile gross margin there is, built on holding frontier capability exclusively, with no process node, no capacity and no physical constraint it can defend. The moment open source closes in, the profit falls back faster than it climbed. The day K3 shoved 3T open-source parameters straight in everyone's mouth, that logic was on the table.
The application layer's anxiety is the plainest of all: the entry point. Everyone is fighting over the same thing, getting users into a habit. Tencent grabs the knowledge base with IMA, Kingsoft black-boxes tokens with per-seat pricing, Sangfor lowers the barrier with visual workflows. This layer competes on retention.
String the four layers together and one line runs through them: the higher you go, the thinner the barrier and the closer the money. Infrastructure is the heaviest and the hardest to replace, and also the hardest to get a premium on; the model layer is the lightest and the easiest to catch up with, and it takes the most of the incremental profit. Nobody is right or wrong here; the industry is simply at this stage. In an arms race, what is scarce is capability. Capacity is not short at all.
But stages change. The day model capability levels out, profit will sink back toward whatever has a physical constraint on it.
Coming in Part II: apart from infrastructure companies like Vertiv, Hall H1 was mostly the software end: cloud, foundation data models and so on. The companies to visit in H2 own the chip layer, which is upstream of cloud (foundation models sit downstream of it), so Part II focuses on the chip layer and the application of AI-grade hardware. The logic and the frame of the whole report sit right here: the vantage point is looking at AI from INFRA, and structurally it runs from H1 through H2 and H3. In the H2 part I will also bring in my own study of LPUs and NPUs, which the main text will get into as well.
Appendices
Appendix A · Route walked and schedule
The route map:

(Fig. 59 H1 floor plan, with the route walked)

(Fig. 60 H2 floor plan, with the route walked)

(Fig. 61 H3 floor plan, with the route walked)
Visit log and route: H1-H3-H2-H4. This was my third WAIC. In past years I went through at a canter, to witness the future and the new order and to watch the shape of the industry and its effects. My habit is to attend on the afternoon of day one and study the products while it is quiet, then watch the livestreams on the morning of day two to learn which way the government is leaning and hear the big names talk.
On which way the government leaned on AI this time, see a video by Xiao Wang Albert, the AI part from July.
On day two I went round Hall H2 with a hangover from the previous night's entertaining, so I did not study it thoroughly, and it set back the writing of the log too. On top of that I know very little about the chip industry and am unclear on the key technology iterations in semiconductors, so writing H2 means learning as I write. I am also travelling for work at the moment; I will get through it as fast as I can.
Further reading on CDUs: a long feature on CDUs on WeChat. This is material I went through while writing the CDU passages in section 01, attached here as further reading.
Appendix B · Glossary
| Abbrev. | Full name | In one line |
|---|---|---|
| ⭐ CDU | Coolant Distribution Unit | The “heart” of liquid cooling: it distributes coolant to every rack by pressure and flow |
| ⭐ HVDC / 800V | High Voltage Direct Current | Power stops being converted back and forth and feeds the server directly; one less conversion is one less loss |
| ⭐ DC / IDC | (Internet) Data Center | The building the servers sit in; managing power and managing cooling are its only two lifelines |
| ⭐ Colo | Colocation | Machine-room leasing: the landlord builds the electrical and mechanical, you move your servers in and pay rent by the kilowatt |
| ⭐ Attention | Attention mechanism | Lets the model decide for itself which words in the input to focus on. The bedrock of large models |
| ⭐ Transformer | — | The model skeleton proposed in that 2017 paper; every large model today grows on it |
| Chinchilla law | Chinchilla Scaling Law | Parameters and data have to be balanced; piling on parameters without enough data is like buying the card and never switching it on |
| ⭐ GPU / NPU / XPU / LPU | Graphics / Neural / eXtended / Language Processing Unit | Parallel compute chips. GPU is the most general; the other three are each vendor's AI-specific variant |
| ⭐ Token | — | The smallest billable unit of text a model handles, roughly half to one Chinese character |
| ⭐ TDP | Thermal Design Power | A chip's rated heat output. Cooling is designed to it; it is the first domino in every thermal design |
| Sparsity | Sparsity | Only part of the parameters or the attention works each time. It saves compute, and you give a little accuracy back |
| ⭐ UPS | Uninterruptible Power Supply | The “power bank” for the few dozen seconds between the grid dropping and the genset starting. If it fails, everything stops |
| Busbar / busway | Busway | The machine room's “power rail”: one copper bar runs down the row and racks tap power close by |
| End-of-row cabinet | End-of-Row PDU | The shared distribution and metering cabinet for one row of racks; when that row has a problem, check it first |
| ⭐ 2N / N+1 / DR / RR / 4N3 | Power redundancy architecture tiers | From “two of everything” to “shared backup across paths”; the more expensive, the less you fear an outage |
| Genset | Diesel generator set | The last line of defence once the grid is fully down; those few dozen seconds the UPS holds are spent waiting for it to start |
| Dry cooler / cooling tower / chiller | Dry Cooler / Cooling Tower / Chiller | Three ways of dumping heat outdoors: blow it away, evaporate it, or fire up a compressor and force the temperature down |
| In-row refrigerant-pump unit | In-row refrigerant-pump air conditioner | An air conditioner set between racks; the ceiling of the air-cooling route, about 50–60 kW per rack |
| ⭐ Liquid cooling / cold plate | Liquid Cooling / Cold Plate | Liquid pressed against the chip carries the heat away; power density that air cannot move has nowhere else to go |
| ⭐ PUE | Power Usage Effectiveness | Total power ÷ IT power. The closer to 1 the better; liquid cooling can push it to 1.1 |
| SST | Solid State Transformer | Solid-state transformer: power electronics replacing the iron lump. Small and adjustable, but a long way from deployment |
| OCP | Open Compute Project | The open server and rack standard led by Meta, making racks as interchangeable as Lego |
| PSU | Power Supply Unit | The power module inside the server; as rack power climbs, it is the first bottleneck |
| ⭐ PoD | Point of Delivery | The smallest deliverable unit of compute: a full set of GPUs + networking + power packaged into “one lump” |
| CE / UL certification | CE Marking / Underwriters Laboratories | The mandatory safety tickets into the EU and North America; without them the product cannot go overseas |
| DDP / DAP | Delivered Duty Paid / Delivered At Place | Trade delivery terms: DDP means the seller pays duty and delivers to the door, DAP means the buyer carries the duty |
| T+6 | — | The domestic delivery convention: electrical and mechanical installed and handed over 6 months after the shell tops out. Overseas it is often 12 months or more |
| ⭐ IRR | Internal Rate of Return | The annualized rate of return on how fast a project pays back; the first number the capital side grills a machine room on |
| ⭐ RAG | Retrieval-Augmented Generation | Look things up in a knowledge base before answering. The first layer of enterprise AI, and the most homogeneous |
| ⭐ Agent | Intelligent agent | An AI that breaks the task up itself, calls tools, runs it to the end and hands the work in. Not just chat |
| OCR | Optical Character Recognition | Turns the words in an image into editable text; the gateway technology for scans and invoices |
| ⭐ MCP | Model Context Protocol | The “USB port” for a model calling external tools: wire it once and it works everywhere |
| To B / To C / To G | Facing enterprises / individuals / government | Same model, three sets of pricing, three sets of delivery, three sales playbooks |
| ⭐ IaaS / PaaS / SaaS / MaaS | Infrastructure / Platform / Software / Model as a Service | The four layers of cloud: sell resources, sell a platform, sell the result, sell model calls. The higher up, the higher the margin |
| ⭐ CUDA | Compute Unified Device Architecture | NVIDIA's GPU programming base; more than a decade of ecosystem is its real barrier |
| cuDNN / cuBLAS | CUDA Deep Neural Network / Basic Linear Algebra Subprograms | NVIDIA's hand-tuned operator libraries; frameworks run fast mainly because of these two |
| ⭐ K8s | Kubernetes | The container orchestration system, the machine room's “shift roster”, deciding which job goes on which card |
| ⭐ PFLOPS / TFLOPS | Peta / Tera FLOPS | Floating-point operations per second, a chip's “horsepower”; 1 P = 1,000 T |
| FP16 / FP8 / FP4 | Floating-point precision | How many bits a number is stored in. Halve the bits and you double the speed, halve the memory and take a small discount on accuracy |
| ⭐ MoE | Mixture of Experts | Split the model into a pile of experts and wake only a few each time, which is how total parameters can get so enormous |
| ⭐ KV Cache | Key-Value Cache | Caches the context already computed so it is not recomputed; the longer the context, the more memory it eats |
| Multimodal | Multimodal | One model that can take in text, images, audio and video at once |
| Total / active parameters | Total / Active Parameters | In the MoE era you have to read them together: the first decides memory, the second decides what each inference costs |
| ⭐ Context length | Context Length | How much it can hold in mind at once; going from 256K to 1M is going from a stack of documents to a whole novel |
| benchmark | Benchmark score | An exam result on a standard question set. Useful for shortlisting, not the same as being good to use |
| ⭐ ARR | Annual Recurring Revenue | Annual recurring revenue, a subscription company's “revenue pulse”; valuation follows it |
| PS | Price-to-Sales | Valuation ÷ annual revenue. AI companies have no profit yet, so this is all anyone has to argue over price with |
| B / T | Billion / Trillion (parameter scale) | 32B fits on one card; 2.8T takes a whole rack |
| ⭐ Per-seat pricing | Per-Seat Pricing | Charging by head rather than by call volume. The signal that AI has moved from selling tokens to selling productivity |
| MAU / DAU | Monthly / Daily Active Users | Monthly and daily actives; the higher DAU÷MAU, the closer the product is to “needed every day” |
| ⭐ Harness | — | The engineering layer outside the model: tools, orchestration, memory, context. Agent = Harness + Model |
| ⭐ Vibe Coding | — | State the requirement clearly in natural language and let the AI write the code; the human only signs it off |
| Foundation model | Foundation Model | A general-purpose base trained from scratch, on top of which vertical applications grow |
| MSA | Mixture of Sparse Attention | MiniMax's sparse attention: coarsely screen for the passages worth reading, then go back and read them closely |
| ⭐ KDA | Kimi Delta Attention | Kimi's approach: compress history into a dynamic note updated as it goes. Saves memory, loses detail |
| GQA | Grouped Query Attention | Lets multiple query heads share one set of KV; the most common memory-saving modification |
| ARPU | Average Revenue Per User | Average revenue per user; $5 on emotional companionship, $6,000+ on enterprise customers. The whole gap is here |
| ⭐ Distillation | Distillation | Have a small model imitate a large one's output, approaching its performance at a fraction of the cost |
| CAPEX / OPEX | Capital / Operating Expenditure | Building your own machine room is CAPEX, paying a monthly cloud bill is OPEX; going to the cloud is essentially that swap |
| CAGR | Compound Annual Growth Rate | Compound annual growth rate, folding several years of growth into “how much per year on average” |
| POC / MVP | Proof of Concept / Minimum Viable Product | Verify it works first, then build the smallest sellable version. The first step on an AI project |
| ⭐ Retain / Rehost / Replatform / Refactor | The four postures for going to the cloud | Leave it, lift it, change the base, rewrite it; effort and payoff rise in that order |
| ⭐ OR-Tools | Google Operations Research Tools | An open-source solver for “find the optimal” problems like rostering, bin packing and routing (briefly explained in section 10, Alibaba) |
| ⭐ Scale-up / Scale-out | Vertical / horizontal scaling | Tying the cards tightly together inside one PoD (NVLink) vs. tying PoDs into a cluster (RoCE) |
| ⭐ NVLink | — | NVIDIA's high-speed card-to-card link, an order of magnitude faster than going over the network. The workhorse of scale-up |
| ⭐ RoCE | RDMA over Converged Ethernet | RDMA over ordinary Ethernet. Cheaper than InfiniBand and the mainstream for large clusters in China |
| ⭐ RDMA | Remote Direct Memory Access | Servers reading and writing each other's memory directly, bypassing the CPU. Far lower latency |
| ⭐ HBM | High Bandwidth Memory | The high-speed memory stuck next to the GPU; in inference it often runs out before the compute does |
| Supernode | Super Node / Rack-scale System | A rack-level product that turns dozens or hundreds of cards into “one big machine” over a high-speed bus |
| TTS | Text-to-Speech | Text to speech; turning a recording into text is the reverse, ASR/STT, which is what a voice recorder uses |
| ⭐ Post-training / SFT | Post-training / Supervised Fine-Tuning | The catch-up class after pre-training: use labelled data to tune a general model into one that can do the job |
| Latent MoE | — | Adds a compressed space to expert routing, easing load imbalance and communication overhead |
| MLA | Multi-head Latent Attention | DeepSeek's approach: compress KV into low dimensions before storing it, with a harder compression ratio than GQA |
| Residual connection | Residual Connection | Puts a “lift” into a deep network so shallow information is not diluted by the dozens of layers after it |
| Pruning | Pruning | Cut the connections that barely do anything in exchange for speed; a Transformer can lose quite a few without losing score |
| RNN | Recurrent Neural Network | The route before Transformers: read straight through and compress history into one state |
| ⭐ Quantization / INT4 / INT8 | Quantization | Drop weights from floating point to integers; memory falls by half to three quarters, accuracy slightly |
| ⭐ LLM / VLM / VLA | Large Language / Vision-Language / Vision-Language-Action Model | Can speak → can see → can act. The three steps of embodied intelligence |
Appendix C · How many tokens one kilowatt-hour produces, and the commercial logic of exporting tokens
First, pin down the order of magnitude. In published calculations, generating a million tokens takes about 15–20 kWh on average (China Power Enterprise Management, “From selling electricity to selling tokens”, chinapower.org.cn/detail/456250.html), which works back to roughly 50,000–67,000 tokens per kWh. Academic measurements put Llama 3 405B on H100 at about 163–236 joules per token, which converts to about 15,000–22,000 tokens per kWh, while small models can reach several hundred thousand (grokipedia.com/page/Energy_per_Token). Squeeze it from both ends and one kilowatt-hour produces “tens of thousands to a few hundred thousand tokens”.
The industry has no authoritative absolute benchmark yet, and “tokens per watt” is becoming the new metric: NVIDIA and CoreWeave publish only relative figures, e.g. Vera Rubin NVL72 running DeepSeek-R1 does 10× the tokens/s/MW of GB200 NVL72 (coreweave.com blog, July 2026). The metric is moving from “compute per watt” to “tokens per watt”; electricity has become the denominator of output.
The commercial logic of going overseas can therefore be written as a chain: China's electricity prices, build speed and supply chain make the production cost per token lower; rack rents and compute prices overseas are higher; so as long as “the electricity-price gap × the delivery-speed gap” exceeds “the certification, logistics and labor premium of going overseas”, selling the compute (or selling the tokens directly) holds up. What containerized delivery compresses is exactly the cost on the right-hand side of that inequality.
Appendix D · A short note on Vertiv's product lines (public material only)
800V/HVDC: a DC power combination for the next generation of high-density racks, aligned with the industry's 800 VDC roadmap (Vertiv press release, 13 October 2025: the combined products release progressively from the second half of 2026, aligned with NVIDIA's 2027 Rubin Ultra platform). The direction is to merge the multiple conversion stages in the AC chain and shorten the path from grid to chip.
CoolChip CDU: the published product line for coolant distribution units spans 70 / 121 / 600 / 1,350 / 2,300 kW (in-rack through centralized; vertiv.com product catalogue). The CDU is the “heart” of a liquid-cooling system, responsible for distributing cooling capacity to each rack on demand.
SmartRun: prefabricated containerized delivery. The electrical and mechanical systems are integrated and tested in the factory before shipping to site, compressing an overseas build cycle that routinely runs a year down to around five months, while sidestepping expensive, low-efficiency on-site labor.
One Core: a whole-machine-room integration concept in which power, cooling and monitoring are pre-integrated on one architecture, with the goal of turning an “engineering project” into a “product delivery”.
(Note: all of the above is on public material only; nothing involving specific customers, quotes or product-selection correspondences is written here.)
Appendix E · What a UPS is, and what a modular UPS is
A UPS (uninterruptible power supply) solves the few dozen seconds between the grid dropping and the genset taking over: while the grid is healthy it rectifies and inverts to the output and charges the batteries; at the instant of an outage the batteries carry the load; once the diesel generator is up, it takes over. In a data center the genset is the busbar-level backup and the battery (lead-acid mostly, lithium more often overseas) is the UPS's backup, working alongside redundancy architectures like 2N and N+1.
A modular UPS breaks the whole unit into a number of hot-swappable power modules: one fails, you swap one, capacity expands on demand, and the granularity is far finer than a tower unit. And precisely because “the whole unit is broken apart and sold as components”, unit price and gross margin get pulled down; layer large customers' category-level centralized procurement on top and pricing power across the whole UPS industry has been thinned. When the main text says the UPS has gone cabbage-cheap, this is the product-side background.
Appendix F · Revision notes on cooling systems
The four components of the refrigeration cycle: compressor (compresses low-pressure gaseous refrigerant into high-pressure, high-temperature gas) → condenser (rejects heat outward, turning it liquid) → throttling device (usually an expansion valve; drops pressure and temperature) → evaporator (absorbs heat and vaporizes, carrying the machine room's heat away) → back to the compressor. Remember the order and you have the skeleton of every cooling product there is.
By heat-transfer medium there are two big systems: water systems (chillers/cooling towers/dry coolers as the cold source, water doing the transport) and refrigerant systems (refrigerant straight to the terminal unit, condenser outside). The liquid-cooling era added a third role: the CDU, which isolates and distributes between the primary side (building chilled water) and the secondary side (the coolant going into the rack), and can be folded into a water system or a mixed water-refrigerant system.
Every one of those dizzying cooling product names is really a variant of the four components in different positions and combinations, which is why after every new-product pitch the customer asks “how is this different from chilled water or a refrigerant system”. See the diagram in the main text (the AI-generated four-component cycle figure).
Appendix G · What to look at when siting a data center
Power: the price level and the peak-valley spread (electricity is the bulk of operating cost), grid capacity and the queue for expansion, green-power quotas.
Climate: mean annual temperature and humidity decide how many hours of free cooling you get, which feeds straight into PUE and the cooling investment.
Policy: energy assessment and the PUE red line, local subsidies and window guidance, speed of approval.
Network: latency to backbone nodes, carrier resources; latency-sensitive workloads (inference services) are pickier about this one.
People and land: supply of operations talent, land price and ground conditions (bearing capacity, water source, flooding).
Plots that meet all of these at once are scarce, which is why data centers naturally cluster: everybody likes the same batch of land.
References
12 in total, numbered in order of first appearance in the text. Accessed 11 August 2026.
[1] About Us | AI Infrastructure Map. AI Infra Map (aiinframap.com). https://aiinframap.com/about
[2] AIinfra Map · CH01 · Overview(全景图). aiinframap.vercel.app. https://aiinframap.vercel.app/ch01-overview.html
[3] Foresight 2026: Panorama of China's IDC (Internet Data Center) Industry in 2026 (market size, competitive landscape and trends)(预见2026:《2026年中国IDC行业全景图谱》). Qianzhan, qianzhan.com. https://qianzhan.com/analyst/detail/220/260429-16bd92f7.html
[4] Create a cluster with Minikube | Kubernetes(使用 Minikube 创建集群). kubernetes.io. https://kubernetes.io/zh-cn/docs/tutorials/kubernetes-basics/create-cluster/cluster-intro
[5] [Large-model basics] A complete breakdown of OCR: an in-depth guide from principles to practice(【大模型基础】OCR技术全解析). CSDN, blog.csdn.net. https://blog.csdn.net/weixin_42969320/article/details/155875670
[6] Multi-dimensional growth in the games industry in 2025; the industry expects sentiment to improve again in 2026 | 2025 year-end review(2025游戏行业多维增长). Cailianshe. https://m.cls.cn/detail/2243678
[7] Alibaba Cloud Intelligence – Business Technology Engineer (Gaming) – Shanghai(阿里云智能-商业技术工程师(游戏方向)-上海). https://careers.aliyun.com/off-campus/position-list
[8] Understanding Harness Engineering: the six-layer architecture, context management and front-line team practice(一文搞懂 Harness Engineering)| JavaGuide. javaguide.cn. https://javaguide.cn/ai/agent/harness-engineering.html
[9] On WAIC and Qwen 3.8(谈谈 WAIC 和 Qwen 3.8). Zartbot, WeChat public account. https://mp.weixin.qq.com/s/jZc-3-h2pIYnhFPhLSlTQQ
[10] DGX SuperPOD: AI Infrastructure for Enterprise Deployments | NVIDIA. https://www.nvidia.com/en-us/data-center/dgx-superpod/
[11] A complete breakdown of GPU scale-up/scale-out: architectural principles, scenario selection and industrial-grade deployment(GPU Scale-up/Scale-out全解析). Zhihu. https://zhuanlan.zhihu.com/p/1980587409814077874
[12] Kimi K3 Hardware Requirements — GPU, VRAM & Server Config | OpenModelMap. openmodelmap.com. https://openmodelmap.com/kimi-k3/hardware?lang=zh