header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

China's large model kill line has been killed.

Read this article in 25 Minutes
Price wars are the final siege tactic of scale monopolists.
Original title: "The Kill Line of Chinese Large Models Has Been Killed"
Original author: Sleepy


On October 8, Anthropic released Haiku 5.5. For short requests of no more than 100,000 tokens, it charges $0.1 per million input tokens and $0.5 per million output tokens, just one-tenth of the previous generation Haiku 4.5.


That morning, Dragonfly managing partner Haseeb Qureshi posted two Artificial Analysis scatter plots on X, with the caption:


"If you want the cheapest LLMs, you should now buy American."


(If you want the cheapest large models, you should now buy American.)



The horizontal axis of the scatter plots is the cost of completing one benchmark task, and the vertical axis is the score. The dashed line connecting all optimal solutions is called the Pareto frontier in microeconomics. Every point on the line means that at the same cost, you cannot find a smarter model.


In the June chart, the cheapest end of the dashed line was Xiaomi MiMo, DeepSeek V4 Pro, MiniMax-M3, and Zhipu GLM-5.2, almost all names of Chinese large models.


By October 8, the entire line had been taken over by Anthropic and OpenAI. Among Chinese models, only MiMo barely clung to the edge, while Zhipu GLM-5.3-Flash and DeepSeek V4.1 Flash were pushed to the lower right of Haiku 5.5, meaning that by comparison they cost more and scored lower.


When the Hong Kong stock market opened that day, large-model concept stocks fell in response. At 10:26 a.m., MiniMax's decline widened to 10%, and Zhipu fell 5.6%. By the midday session, the Hang Seng Tech Index was down 1.93%, with the two large-model newcomers underperforming the broader market by nearly five times.


Over the past two years, Chinese open-source and high-cost-performance models have welded a kill line into the global large-model market. As long as prices were driven into the mud, overseas volume-tier closed-source models became difficult to sustain. Star Silicon Valley companies lined up to replace their underlying pipelines with Chinese open-source large models.


Now, the table has been flipped.


The kill line has been killed.


The Formation of the Kill Line


When GPT-4 first launched in March 2023, the official pricing for the 8K context version was $30 per million input tokens and $60 per million output tokens.


Intelligence was extremely scarce, supply was monopolized by a handful of Silicon Valley oligarchs, and pricing was entirely a seller's market. It was a honeymoon period that believed "intelligence alone can command a premium."


What shattered this myth was a technical team incubated by a Chinese quantitative hedge fund. On May 6, 2024, DeepSeek launched V2. With 236 billion total parameters and only 21 billion activated per token; relying on the MLA architecture and extreme engineering trimming, the team cut KV cache by 93.3% and boosted throughput by 5.76x. The saved compute was directly converted into commercial pricing: 1 yuan per million input tokens and 2 yuan per million output tokens.



Domestic tech giants were practically dragged kicking and screaming into the price war.


On May 15, ByteDance's Volcano Engine launched Doubao, offering an enterprise price of 0.0008 yuan per thousand tokens for its flagship model, claiming it was 99.3% cheaper than the industry average. On May 21, Alibaba Cloud announced a cliff-edge price cut for the Tongyi Qianwen series, with Qwen-Long charging only 0.5 yuan per million tokens on the input side; hours later, Baidu declared two of its flagship ERNIE models free of charge.


By August, DeepSeek further introduced hard disk caching technology, charging only 0.1 yuan per million input tokens on cache hits.


Across the ocean at the time, the giants did not fully follow suit. In July 2024, OpenAI released GPT-4o mini, with input and output prices dropping to $0.15 and $0.6; but four months later, Anthropic launched Claude 3.5 Haiku, priced at four times that of the previous generation. Dario Amodei's logic at the time was that the model had gotten smarter, and pricing must reflect the advancement of intelligence.


What truly sent a bone-chilling shiver through Silicon Valley was January 2025.


DeepSeek-R1 burst onto the scene, with reasoning performance approaching OpenAI o1, while output cost only $2.19 per million tokens, whereas o1 was priced at $60 at the time.


On January 27, Nvidia lost nearly $589 billion in market value in a single day, setting a record in U.S. stock market history. Marc Andreessen called it the Sputnik moment for AI, and Microsoft CEO Satya Nadella remarked on social platforms that the Jevons paradox had struck again.


The kill line was thus drawn.


Open weights make this blade even sharper. With closed-source models, customers can only do arithmetic among a few monopoly price lists; once weights are open, developers can both deploy privately on their own clusters and run the same model on any third-party inference cloud. The same model has multiple suppliers, and customers gain the option of self-deployment, shrinking the room for model providers to maintain high premiums.


A group of U.S. companies crossed the river first.


In October 2025, Airbnb CEO Brian Chesky publicly stated that the company's intelligent customer service Agent relied heavily on Alibaba's Qwen, with the entire system orchestrating 13 models. Although it also integrated OpenAI's latest flagship, it was rarely invoked in the production pipeline because there were more cost-effective alternatives.


That same month, Cognition launched SWE-1.5, claiming it was built on "an industry-leading open-source base." Zhipu later confirmed that the base was GLM-4.6.


Shortly after, a16z partner Martin Casado told The Economist that among U.S. AI startups pitching with open-source model architectures, about 80% were using Chinese bases.


In March of this year, code editor Cursor released Composer 2, priced at $0.5 per million input tokens and $2.5 per million output tokens. The team eventually admitted that the base was fine-tuned from Moonshot AI's Kimi K2.5.


Developers are voting with code and budgets. A report released by Mozilla shows that the token share of Chinese open-weight models on OpenRouter rose from less than 2% at the end of 2024 to more than 45% by April 2026.



In the ecological niches of volume-driven usage and structured invocation, the haughty closed-source models were once powerless to fight back.


The Cost


Slaying opponents does not necessarily mean living decently oneself.


In the same Mozilla report, there is also a set of comparative data. From May to September 2025, open models accounted for about 20% of model usage on the OpenRouter platform, yet received only about 4% of model-layer revenue. This is a sample of one platform during a specific period and cannot represent the global market, but the gap between usage share and revenue share is already very clear.



Creating value has never been equal to capturing value.


In January of this year, Zhipu and MiniMax rang the opening bell in succession. The interim results subsequently disclosed allowed the outside world to see clearly the costs and losses behind revenue growth.


Zhipu's first-half revenue was RMB 954 million, of which open platform and API business revenue was RMB 825 million, accounting for 86.5%. The gross margin of this business has turned positive from negative, rising to 24.6%; but the company as a whole still recorded a net loss of about RMB 2.072 billion, narrowing by 12.1% year on year. Selling more tokens has begun to generate gross profit, but there is still a distance before it can cover all of the company's expenses.


MiniMax's first-half revenue was about $117 million, up 283.1% year on year; net loss was about $358 million, narrowing by 11% year on year. After excluding share-based payments, changes in the fair value of financial liabilities, and listing expenses, the adjusted net loss was about $293 million, expanding by 111.2% year on year.


After listing, the growth brought by low prices needs to withstand another test. How much gross profit can each call leave behind, and how much of that gross profit can cover R&D and operating expenses. Revenue is growing quickly, but the company is still far from overall profitability.


In late spring of this year, the Agent wave pushed computing power consumption to a critical point.


The open-source project OpenClaw swept through the global developer community, and by early March its cumulative token consumption on OpenRouter exceeded 8.52 trillion, ranking first. Within one week in mid-March, Chinese models ran 7.36 trillion tokens in a single week, surging 56.9% month on month.


GLM-5, launched on February 12, once topped the usage rankings, and surging demand instantly overwhelmed infrastructure quotas. Overseas and domestic developers began intensively complaining that GLM Coding Plan suffered second-level latency and high-frequency rate limiting. This was the first time a leading Chinese AI team publicly issued an urgent call for computing power.


On February 23, Zhipu's stock price plunged nearly 23% in a single day, wiping out more than HK$70 billion in market value in one day.


What followed was a round of defensive price hikes.


Zhipu raised its rates when GLM-5 was launched, and the price of GLM-5-Turbo rose another 20% in March, making it an average of 83% more expensive than the previous generation of specifications.


DeepSeek's V4-Pro, released on April 24, has 1.6 trillion total parameters and offered a 75% peak discount upon launch. This fire first burned its peers, with MiniMax falling about 9% and 10% over two consecutive trading days.


But by mid-August, DeepSeek also shifted to time-of-use pricing, with V4-Pro's peak output climbing to $3.96 per million tokens, and cache-hit costs rising more than 12-fold.


In July, Moonshot AI released Kimi K3, with API list prices pushed up to $3 per million input tokens and $15 per million output tokens, more than three times as expensive as K2.6.


Yet, unexpectedly, the price increases did not trigger severe customer churn.


Zhipu CEO Zhang Peng publicly revealed that after an 83% price adjustment in the first quarter, API call volume instead surged 400%. By the end of August, Zhipu's MaaS platform's annualized revenue reached $1.6 billion, and the API business's gross margin miraculously turned positive to 24.6%.


During the months-long window when computing power was in short supply, Chinese model teams touched precious pricing power for the first time.


However, the kill line that once kept overseas giants awake at night was also, imperceptibly, pushed higher by themselves.


Blitzkrieg


The other side of the ocean completed a strategic pivot in the summer.


At the end of June, Anthropic released Sonnet 5, listing a launch promotional price of $2 per million input tokens and $10 per million output tokens, having originally explicitly stated that prices would return to $3 and $15 in September.


Then in July, OpenAI launched the GPT-5.6 series, and three weeks later suddenly launched a blitzkrieg, cutting Luna's price by 80%, with the input side falling to $0.2 and the output side to $1.2.



OpenAI attributed this price cut to pure underlying engineering reconstruction: rewritten GPU kernels reduced end-to-end serving costs by about 20%, and a retrained speculative decoding draft model increased generation throughput by more than 15%.


On August 10, Anthropic withdrew its price increase notice and announced that Sonnet 5 would be permanently locked at a low price.


On September 22, Anthropic rolled out Opus 5.5, cutting comprehensive usage costs by 40%; on the same day, OpenAI launched GPT-6 Sol and Luna, with Luna locking its entry-level price at $0.1 for input and $0.5 for output. What followed next was Haiku 5.5.


Two years ago, Anthropic firmly believed that "intelligence must command a separate premium." Now it has torn off its own label with its own hands.


The rhetoric around large model releases has changed as well. No longer is it a pure listing of marginal wins on MMLU; they have begun repeatedly emphasizing output per dollar, latency, and the engineering efficiency of every step of inference.


Moreover, at the same time Haiku 5.5 made its debut, Anthropic announced that it would fully allocate API credits to subscribers, granting $100 to $200 per month at the Max tier, with enterprise Team users eligible for up to $500 in offsets.


This move strikes straight at the heart of the battlefield. Discounts on a price list only affect single-call decisions, but once a customer's code engineering and agent orchestration are fully adapted, the barrier to switching will be as high as a city wall.


That American giants dare to launch this price war rests on unfathomably deep capital pools and supply chain privileges.


Anthropic's annualized operating revenue soared from about $9 billion at the end of 2025 to more than $65 billion by the end of July this year. The $65 billion financing completed in May directly pushed its valuation to $965 billion.


On the compute front, Anthropic announced in April that it would expand cooperation with Google and Broadcom, obtaining multi-gigawatt-scale next-generation TPU compute capacity, expected to come online progressively starting in 2027.


By contrast, Zhipu struggled within the year to raise more than HK$70 billion through placements and convertible bonds, which, converted into U.S. dollars, is less than one-seventh of Anthropic's single round of ammunition.


Those who can afford to pay are still only the super giants.


Liquidation in the secondary market has never been merciful.


Zhipu listed on the Hong Kong stock market on January 8 with an issue price of HK$116.2. The spring compute shortage failed to stop the bulls' frenzy. On June 22, its share price surged to HK$2,980, and its market value briefly touched the ceiling of HK$1.33 trillion.


Then came the long slide.


In July, lock-up shares were released from restriction, and two rounds of placements hammered the private placement price from HK$1,588 all the way down to HK$714. On July 16, the day after Kimi K3 debuted, Zhipu plunged 28.48% for the full day. On September 22, Opus 5.5 and GPT-6 launched in tandem, and Zhipu fell more than 10% again.


On September 25, Zhipu touched an intraday low of HK$610.5, its total market value shrinking to around HK$300 billion, a drawdown of nearly 80% from its all-time high.


MiniMax's trajectory was equally brutal. Falling from a peak of HK$1,330, it plunged 17.98% on the day its lock-up expired and closed at HK$216 in mid-July. Nearly 60% of its revenue in the first half relied on overseas customers, and overseas was precisely the battlefield hit first and hardest by the U.S. giants' price-cut storm.


Valuation metrics also began to tighten. According to media reports, Jefferies in a September research note cut the valuation multiple for Zhipu's cloud business—based on 2026 forecast annualized revenue—from 50x to 30x, a reduction of 40%.


Chinese models once snatched food from the jaws of the giants by virtue of extreme cost-effectiveness, and now that path has been proven viable on the other side of the ocean as well. Developers come because it's cheap, and naturally leave because something is even cheaper.


The Ghost of $0.1


On August 25, 2006, on a sweltering beach in Cabo San Lucas, Mexico, Jeff Barr typed out an article introducing the Amazon EC2 beta.


At the end of the technical description, he wrote down a number destined to go down in business history: developers renting a virtual compute instance would be charged $0.1 per hour.


The rest of the story is well known. AWS proactively cut prices more than a hundred times over the following decade-plus, and in 2014 alone halved S3 storage rates by 51%.


Cloud computing never built its empire by selling the same computing power at ever-higher prices. Every cent of marginal profit squeezed out of scale, in-house custom silicon, and extreme operational efficiency was continuously returned to the end market, thereby attracting more developers. The perpetually price-cutting AWS ultimately grew into a cash cow with annual revenue of nearly $130 billion and an operating margin as high as 35.4%.



Price wars are the final encirclement tactic of the scale monopolist.


Those who today can slash model prices to a tenth and stock the shelves of all three major global public clouds the moment they launch are not just selling tokens—they are also consuming tokens. Anthropic, while slashing API prices, feeds its massive inference compute back into Claude Code, Cowork, and its own end-to-end Agent products.


China's large model teams still haven't laid down their arms. In September, DeepSeek once again cut prices for its Flash series, Xiaomi pushed MiMo-Flash to the extreme of 1 yuan per million tokens, and Zhipu also priced GLM-5.3-Flash firmly in the razor-thin margin zone. But the rules for the chips on the table have already been rewritten.


Low pricing itself is no longer an exclusive moat card.


On that scatter plot updated in real time by Artificial Analysis, not a single Chinese company's name is absent. GLM, DeepSeek, Kimi, Qwen, and MiniMax are all still there.


It's just that they can no longer draw that kill line.





Original link


Welcome to join the official BlockBeats community:

Telegram Subscription Group: https://t.me/theblockbeats

Telegram Discussion Group: https://t.me/BlockBeats_App

Official Twitter Account: https://twitter.com/BlockBeatsAsia

举报 Correction/Report
Choose Library
Add Library
Cancel
Finish
Add Library
Visible to myself only
Public
Save
Correction/Report
Submit