header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Meituan Open Sources Trillion-Parameter Large Model LongCat-2.0, Simultaneously Releases Chinese Domestic Card Inference Code

According to Cognition Beating monitoring, Meituan has officially open-sourced the trillion-parameter LongCat-2.0 large model, with a total of 1.6T parameters and an average activation of about 48B, designed specifically for real Agentic Coding tasks. The architecture innovatively introduces LongCat Sparse Attention and N-gram Embedding. The former reduces fragmented memory access through flow-aware indexing and hierarchical indexing, accelerating training and inference with million-level long contexts. The latter allocates 135B parameters to the embedding layer with MoE sparsity approaching 97%, balancing parameter benefits and structural stability. Post-training adopts multi-teacher online distillation, categorizing the experts into Agent, Inference, and Interaction, seamlessly integrating them on a domestic computing cluster through the MOPD architecture. As the industry's first trillion-parameter model to complete inference on a domestic computing cluster with fifty thousand cards, LongCat-2.0 validates the domestic chip's capability to handle complex large-scale model tasks.


Addressing China's domestic chip constraints such as VRAM, bandwidth, and interconnect, Meituan has made breakthroughs in three directions: at the model level, using ScMoE to leverage the domestic chip's core control capability to achieve dense and MoE branch physical core-level parallelism, combining KV-cache partitioning to alleviate the VRAM pressure of ultra-long contexts; at the chip adaptation level, utilizing Super Kernel to reduce operator launch overhead, hiding I/O latency through Weight Prefetch, and maximizing hardware utilization under constrained conditions; at the deployment level, implementing PD separation to balance TTFT and TPOT, along with Expert-Parallel asynchronous load balancing to address load imbalance under high expert degrees of parallelism. This open-source release simultaneously provides various precision versions such as BF16, FP8, and INT8, and fully opens up the inference results optimized for domestic computing power, with the goal of enabling smooth deployment of trillion-parameter large model inference services on existing domestic cards, including older ones.


---------------------------------

Click the original text link below to join the Cognition Beating · Feishu AI News Channel, which monitors global AI hot topics and news 24/7.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish