header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

DeepSeek, Huawei, 'Must Succeed': 160,000 Fireflies Illuminate the Wasteland

Read this article in 29 Minutes
Must succeed.
The original title: "DeepSeek, Huawei, 'Must Succeed': 160,000 Fireflies Illuminate the Wasteland"
The original author: Dongcha Beating


In 2019, the most expensive asset in China's quantitative trading circles was not in Lujiazui, but in the server room of an office building in Hangzhou.


1,100 GPUs — High-Flyer paid nearly 200 million yuan in real money. The racks were lined up, occupying an area close to a basketball court. There was no day or night in that server room; the only background noise was the piercing howl of high-speed fans.


The person managing the machines gave this cluster a codename: "Firefly."


The name was light and airy, but the calculation behind it was brutally realistic. In a year when large models had not yet become a prominent discipline, 1,100 cards ran day and night without rest, their sole task being to compute the next basis point of alpha for their owner from massive tick data before the next day's opening auction.


A pure money-printing machine.



In the same year, over a thousand kilometers away in Shenzhen, Huawei was added to the Entity List. The world's most advanced process nodes and semiconductor IP were completely cut off on that day.


Both groups were paying for unknown variables.


High-Flyer believed in algorithms. A few young Zhejiang University graduates just wanted to turn massive electricity bills into excess returns on the books before the market caught on; the self-developed chips in Huawei's hands were a costly Plan B, whose best destiny was originally to never be deployed in its lifetime.


At that time, they had no intersection whatsoever.


No one could have predicted that the cluster of computing power lit up for the secondary market in the deep of Hangzhou nights would, years later, travel all the way south and ultimately land in the silicon wafers of that old warehouse in Shenzhen.


Firefly


For a long time, Liang Wenfeng had almost no face in China's tech world.


He was born in Zhanjiang, Guangdong, scored first in the entire city on the college entrance exam, and then went north to Zhejiang University to study machine vision. During the most frenzied years of the mobile internet, most of his brilliant peers rushed to big tech companies to do recommendation algorithms, or squeezed into the CV track to work on facial recognition.


Liang Wenfeng chose something that seemed completely unsexy at the time: teaching machines to trade stocks.


The business logic was actually extremely dry — in the thousandth of a second when a matched trade is completed, turning chaotic high-frequency data into excess returns on the books. He and a few Zhejiang University classmates grew this quantitative institution called High-Flyer to over 100 billion yuan in assets under management.


Liang Wenfeng fundamentally distrusted human judgment. Traders compete on reflexes, analysts on connections—but High-Flyer flipped the script entirely. They believed in machines, and only machines.


In 2019, they built "Firefly No. 1" with 1,100 GPUs; by 2021, the bet had quintupled. High-Flyer shelled out 1 billion yuan, sweeping up tens of thousands of Nvidia A100s in one go, with a server room spanning ten basketball courts. This was "Firefly No. 2."



At the time, many thought he was insane. A quantitative fund hoarding a pile of power-hungry metal—why sink billions into infrastructure with no apparent rationale?


Until October 2022, when the U.S. Department of Commerce dropped a ban that welded shut the most advanced computing channels. Liang Wenfeng had quietly bought up all the chips he needed before the iron curtain fully closed.


He is the kind of person who walks too far ahead of his time.


Ren Zhengfei took a completely different path. He built things first, tossed them into the shadows, and then waited quietly.


That wait lasted a full fifteen years.


Ren Zhengfei is a full forty years older than Liang Wenfeng. In 1987, this 43-year-old man from a small county in Guizhou, with 21,000 yuan scraped together from everywhere, founded Huawei in a cramped residential room in Nanyou, Shenzhen.


The rest is history—starting as a distributor of Hong Kong switches, moving to self-developed communications equipment, then sweeping the globe with 5G base stations and smartphones. The business empire sprawled enormously, but there was a hidden thread Ren Zhengfei buried deep, rarely dissected under the spotlight.


In 2004, Huawei established a wholly-owned subsidiary called HiSilicon.


HiSilicon had only one purpose: to make chips. The ultimate metric Ren Zhengfei set for this team was simple—if external supply ever got cut off, Huawei needed a fallback. In an era when global division of labor was held as gospel, pouring money into this bottomless heavy industry seemed utterly counterintuitive to most.


On December 1, 2018, Canadian police detained his daughter Meng Wanzhou at Vancouver airport.


From that moment on, HiSilicon—a subsidiary that had stayed underwater for fourteen years—was forced to surface in an extremely brutal way. It was no longer a seemingly superfluous "idle move," but the sole lifeboat for the entire giant ship.


On May 16, 2019, the U.S. Department of Commerce entity list took effect.


In the early hours of the next day, HiSilicon President He Tingbo wrote in a company-wide letter that all the backup plans that had lain dormant for years were officially activated overnight.


Three months later, Huawei unveiled the "Ascend 910."


The contrast between the two scenes was stark. Firefly was locked away in a temperature-controlled server room in Hangzhou, with the outside world knowing nothing beyond the numbers on the books; Ascend, meanwhile, was pushed into the center of the spotlight, subjected to the industry's scrutinizing and critical gaze.


People on both ends were spending enormous cash flows in advance for something that had not yet happened.


It was just that the gate would close faster than anyone had anticipated.


The Blunt Knife


The first thing to be cut off was the terminal business.


In September 2020, TSMC halted wafer foundry services, and the 5-nanometer Kirin 9000 became a swan song. Huawei held first-tier chip design capabilities, yet could not find a single foundry anywhere in the world willing to take its orders.


The real shadow war shifted to the server rooms.


Ascend was pushed to the front line. But in the face of Nvidia's mature CUDA ecosystem, almost no commercial customer was willing to pay for an unproven domestic system. Since single-chip computing power could not catch up to Nvidia at the physical limit, Huawei simply switched to a solution defined by extreme engineering brute force.


If one chip wasn't enough, they would forcibly link thousands of slightly inferior chips into one cluster.


The cost of this approach was soaring power consumption and spinning electricity meters, but Huawei accepted the bill. China has no shortage of cheap green electricity, nor of engineers in batches who can chew through hard problems. This was an extremely clumsy, resource-devouring path—but also one that only they could afford to take.


On another track, Liang Wenfeng faced his own major test.


In October 2022, export controls took effect, with Nvidia's A100 and H100 completely banned from sale to China. After that, the only chips flowing into the country were the precisely neutered special-edition A800 and H800.



Domestic tech giants and startup teams fell into unprecedented FOMO. Some scoured the market for second-hand cards; others looked for underground computing pools in Southeast Asia. Liang Wenfeng did not join the buying frenzy. He had already prepared his cards well in advance.


But stockpiling ahead of time could only solve the immediate fuel problem. A colder reality lay ahead: from now on, even if you paid several times the premium, you could never again buy the fastest blade of the same generation globally.


When the tool itself falls short, what do you use to build something that rivals your competitors?


Forcing a craftsman to use a blunt knife to carve intricate patterns indistinguishable from those made with a sharp blade.


Everything DeepSeek did afterward was push this engineering capability of carving flowers with a blunt knife to its absolute limit.


By the end of 2024, they had trained DeepSeek-V3, a model whose overall performance rivaled that of top-tier labs, using only 2,048 Nvidia H800 chips with their interconnect bandwidth slashed, at a cost of approximately $5.576 million.


This cost ledger stunned Silicon Valley outright. To reach the same level, leading overseas labs typically need to deploy tens of thousands of top-tier GPUs and burn through tens of millions to hundreds of millions of dollars.


DeepSeek won on a set of almost brutally obsessive engineering discipline, squeezing every megabyte of compute and memory bandwidth dry.


Geopolitical blockades did not strangle them; instead, they forced out the company's most moat-worthy core asset. But everyone knows that running on castrated Nvidia chips means the knife handle is still firmly in Californians' hands.


The narrow gate ahead leaves only one option.


Later, people gave this path a word steeped in compromise: domestic substitute. But this was never some shrewd cost-performance choice—it was clearly a path clawed open through sheer grit after being choked by the throat.


Hook Punch


If Huawei's counterattack was a grinding trench war, what Liang Wenfeng threw was an unreasonable hook punch.


In 2024, DeepSeek first used extremely aggressive token pricing to blow through the entire large model commercial API price floor, forcing peers to revise their pricing sheets overnight.


On January 20, 2025, DeepSeek open-sourced DeepSeek-R1 without any warning, a model with deep reasoning capabilities, with a publicly disclosed training cost of only a few million dollars—a mere fraction of Silicon Valley's top budgets.


Capital markets completed their pricing with lightning speed. Seven days later, when U.S. stocks opened on Monday, Nvidia's market cap evaporated by nearly $590 billion in a single day, setting the largest single-day market cap drop for a single listed company in U.S. stock market history.



What's more interesting is DeepSeek's capital structure. While the large model track was frantically grabbing Middle Eastern hot money and big tech strategic investments, DeepSeek never took a single cent from external institutions. The cash flow that High-Flyer earned penny by penny through high-frequency algorithms in the secondary market became Liang Wenfeng's war chest for this bold gamble.


The aftershocks of R1 toppling Nasdaq have yet to subside, and Huawei has simultaneously launched the full deployment images for R1 and V3 on the Ascend community.


It runs.


The conclusion that "domestic chips simply can't run cutting-edge models" was pierced through on this day.


However, according to our understanding, the deep binding between DeepSeek and Huawei actually began much earlier than the outside world has seen. It wasn't until early 2025 that Huawei suddenly discovered that DeepSeek had long been secretly conducting extremely deep adaptation and stress testing on Ascend's underlying environment behind everyone's backs.


The people doing hardware weren't even the first batch of insiders in this mysterious lab.


The generation gap remains enormous. The single-card floating-point computing power of Ascend 910C is only about one-third that of Nvidia's flagship B200. Unable to compete on a per-card basis, Huawei chose to continue brute-forcing it at the engineering architecture level, using thousands of high-spec optical fibers to forcibly connect 384 910C chips into a massive "supernode," barely pulling total computing power up to the same level through network topology.



The cost is obvious: terrifying power consumption at four times that of Nvidia's architecture.


But China's energy endowment happens to accommodate this kind of consumption. On the inland Gobi Desert, wind and solar power are cheap enough.


The posture isn't exactly graceful, but at least it can run. The same impenetrable Iron Curtain finally pushed two parties who originally had no intersection at all to face each other.


Liang Wenfeng's models run on Nvidia chips that could be cut off at any moment; while Huawei has built the largest-scale AI chip in China and urgently needs a world-class model to endorse the actual combat capability of this infrastructure.


Two puzzle pieces, fitting perfectly together at this moment.


The most solid alliances in the business world have never been because of shared ideals.


They're all forced into existence.


Must Succeed


In 2026, DeepSeek rarely sought external funding.


Before this, Liang Wenfeng had never bowed to the capital markets. Even in the most sluggish years, High-Flyer's cash flow was more than enough to support an AI lab of a few dozen people.


When a geek who never lacks money suddenly speaks up asking for money, there's only one reason. The boulder he wants to push has grown so massive that no individual pocket could possibly contain it.


What he aims to do is team up with Huawei to completely migrate the next-generation flagship model onto a purely domestic computing power base.


In a funding pool totaling $7.4 billion, the largest single check came from Liang Wenfeng himself, at 20 billion yuan, accounting for nearly two-fifths of the total.


But money is often the easiest variable to solve in hardcore industries.


The real tough nut to crack is the software stack. Uprooting the model from Nvidia's CUDA ecosystem, which has dominated the industry for nearly two decades, and migrating it to Huawei's CANN architecture is tantamount to tearing down an entire skyscraper and rebuilding it.


This is not a conventional move. You have to use a different, less familiar set of tools to rebuild the same skyscraper from the foundation up on flat ground. The underlying operators must be manually rewritten and reconstructed one by one, numerical precision must be realigned across tens of millions of inferences, and even the stack traces thrown during compilation errors are completely unfamiliar.


Workstations in Hangzhou and Shenzhen stay lit late into the night. A large number of engineers, facing development documents full of unknowns, work through the night manually tuning operators until the glaring red lines on the terminal screen are eliminated one by one in the early morning.


Reactions soon came from across the ocean.


On April 15, 2026, Jensen Huang said on Dwarkesh Patel's podcast that if DeepSeek is the first to successfully run and release its next-generation flagship on Huawei's Ascend platform, "that would be catastrophic for the United States."


Huang, who has fought hard battles for over thirty years and rarely shows emotion in public discourse, unprecedentedly uttered the word "catastrophic."


The loss of orders for tens of thousands of chips is not fatal to him. What truly unsettles him is the rules.


For twenty years, the world's top model teams have had an unspoken iron law.


The best algorithms must run on Nvidia's CUDA ecosystem. That is the deepest moat he has built.


And now, for the first time, a first-tier outlier has publicly refused to pay for this set of rules.


When it comes to moats, once someone breaches a gap in a blind spot, the water can never be held back again.


Liang Wenfeng has done the math: to train a model that reaches top overseas benchmarks, about 50,000 of Nvidia's latest flagship GPUs would be needed; switching to Huawei's Ascend 950 series, that number would need to balloon to a full 200,000.


A four-to-one hardware attrition rate, plus at least a two-year generational lag.


What's even more constricting is the current capacity supply. The single-batch quota Huawei can currently free up for DeepSeek is only about 16,000 chips.


With such meager chips in hand, there is simply no possibility of going head-to-head on parameter scale.


But he still bet the entire company's training infrastructure wholly on Huawei chips. This is the biggest bet since DeepSeek was founded.


According to The Information, citing people familiar with the matter, Liang Wenfeng said only one thing at the time:


"It must succeed."


Grassland


The end of this road lies on a stretch of grassland in Inner Mongolia.


In September 2026, DeepSeek was reported to be planning a super-large data center in Inner Mongolia, which will directly pack in a full 160,000 Huawei Ascend chips.


The energy consumption of the entire campus is planned at the GW level. The core model of this batch of chips is the Ascend 950DT, each of which packs a massive 144GB of memory, relying on an extreme intra-chip bus to distribute tokens without delay at the very moment the model completes inference.


This specific model was, in fact, projected onto the big screen by Rotating Chairman Xu Zhijun as early as September 2025 at Huawei Connect. From the 950PR and 950DT to the later 960 and 970, Huawei has set the pace of computing power evolution at one generation per year.



When Xu Zhijun outlined the product blueprint line by line on stage, global partners sat in the audience; but Liang Wenfeng, who a year later would stake his entire fortune on this batch of cards, was not in the venue at the Expo Center at that time.


These chips will ultimately light up a stretch of grassland.


There is almost nothing on the grassland except howling winds and unobstructed blazing sun. But in this brutal computing power equation, cheap green electricity has become the most critical variable.


For an economy whose lifeblood is choked off in advanced processes, the only chips it can currently put on the table are wind, sunlight, and a boundless, nearly zero-cost wasteland.


Liang Wenfeng knows all too well the weight of these 160,000 chips. When he said "It must succeed" in front of investors, he personally burned all his retreat routes.



He has been proud his whole life. He doesn't mingle in circles or chase trends. When he first set out to build large models, his original intention was extremely pure—he had had enough of the Chinese tech community only picking up crumbs behind Silicon Valley.


For such an extremely arrogant technical founder to be willing to make these four words explicit, he has already laid his cards on the table.


Proud, and with no other choice.


From orders landing, production line scheduling, to the data center being powered on and lit up, it takes at least a long cycle of a year and a half. How much can ultimately be delivered depends on the fragile and narrow manufacturing yield upstream. No one can guarantee it.


Deep in the grasslands of Inner Mongolia, the long wind howls across from the far horizon. The air is dry and transparent, and looking up late at night, one can see the entire complete Milky Way.


160,000 chips will eventually light up one by one in the wind and sand of the northwest. Before the vast grasslands, those faint indicator lights are still as tiny as they were in a Hangzhou data center late at night years ago.


But in the long cycle, the lights still have to be turned on.


As for whether these 160,000 points of faint light will ultimately connect into a bright expanse, or be scattered by the strong wind across the vast Gobi.


This question can only be left to time itself to answer.


-END-


Original link


Welcome to join the official BlockBeats community:

Telegram Subscription Group: https://t.me/theblockbeats

Telegram Discussion Group: https://t.me/BlockBeats_App

Official Twitter Account: https://twitter.com/BlockBeatsAsia

举报 Correction/Report
Choose Library
Add Library
Cancel
Finish
Add Library
Visible to myself only
Public
Save
Correction/Report
Submit