header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

NVIDIA's Fortress is dealt another blow: 'Bull Runner' model GLM-5.3 Flash mints 230 trillion Tokens, all running on domestic chips

Dynamic Beating AI News Flash: Following the GLM-5.3 Flash release, Intellisense has confirmed a previously undisclosed detail: during the Ox Alpha anonymous testing phase, all traffic was powered by a domestically produced Chinese AI chip for inference computation.


Ox Alpha, launched on OpenRouter, processed 23.2 trillion Tokens in just 6 full calendar days, more than double the throughput of the concurrent DeepSeek-V4-Flash. OpenCode had even suggested earlier that there was the capability to provide a daily free quota of 100 trillion Tokens.


Intellisense stated that they have optimized end-to-end inference performance on the same Chinese domestic hardware to initially triple the efficiency, and now the hardware efficiency and per-Token cost have approached levels close to mainstream NVIDIA GPUs.


SemiAnalysis specifically pointed this out. Previously, observers speculated whether the supply capacity of 100 trillion Tokens per day could only be sustained by top-tier labs and NVIDIA GPUs. However, this time, the entire process was run using domestically produced Chinese chips.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish