header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Magic claims $500,000 rivals DeepSeek V4 Pro base model, cutting pre-training compute by 50x.

动察 Beating AI Flash: AI coding company Magic released a pre-training research report, claiming that a new training approach can achieve pre-training results close to DeepSeek V4 Pro Base with only about $500,000 worth of GB200 compute, requiring roughly 1/50 of the computational cost.


Base refers to a base model that has only completed pre-training and has not yet undergone post-training such as reinforcement learning. Magic tested the models using unseen code, math problems, and research papers to see which one predicts subsequent content more accurately. Therefore, this comparison focuses on the pre-training effectiveness of the base models and does not imply that chat, coding, or agent capabilities have already matched the finished version of DeepSeek V4 Pro.


Magic then scaled up the training by 10 times, equivalent to about $4 million worth of GB200 compute. In the same set of self-tests, this version surpassed the latest public Base models from DeepSeek, Kimi, and NVIDIA that it evaluated. Magic estimates that if it were to achieve the same level using DeepSeek V4 Pro's training efficiency, the compute cost would exceed $100 million.


Magic has not disclosed the full training recipe, only revealing that the efficiency gains come from dozens of changes across model architecture, optimizers, training objectives, and data processing. The model weights have also not been released yet, so the "50x efficiency" claim cannot be independently reproduced by external parties.


举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish