header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Dai Jifeng's NaiveAI Debuts First Open-Source Model: AI Participates in 151 Rounds of Optimization in 6 Days

Beating AI News Flash: NaiveAI, founded by Tsinghua University Associate Professor Dai Jifeng, has released Naive-N0.5-Flash and opened the model weights. The model is adapted from Xiaomi's MiMo-V2.5 Base and is open-sourced under the MIT license.


NaiveAI retained the original MoE architecture and focused on modifying the attention mechanism. The team replaced global attention with sliding window attention and DeepSeek sparse attention, then continued training on 3.25 trillion tokens. The modified model natively supports a 1 million token context, which can reduce the computational load when processing ultra-long texts.


AI also directly participated in the R&D. It was responsible for writing code, running experiments, analyzing results, and continuing optimization, while researchers set the direction and made key decisions. Taking the inference system NaiveRT as an example, the team ran 151 rounds of optimization experiments in 6 days, of which 63 rounds were adopted.


The officially announced peak inference speed is 2,122 tok/s, but the testing conditions were quite special: using 8 GPUs, with thinking mode disabled, excluding input processing time, and taking the best one second out of 41 requests. In standard mode, the officially stated speed is about 50 tok/s per user.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish