Beating AI News Flash: Zhipu has launched GLM-5.3-FlashX, which can be understood as a high-speed version of GLM-5.3-Flash, with a maximum output speed of 200 tokens/s. This time, the official side did not release new capability benchmark scores or architectural changes; the focus of the update is simply to continue improving inference speed.
This release comes right after yesterday's inference system optimization. Zhipu just disclosed that the Infra Agent powered by GLM-5.3 participated in optimizing Flash's inference infrastructure. In less than two weeks, end-to-end throughput increased to about 3 times the initial level, and the entire service runs on more than 100,000 domestic AI chips.

