动察 Beating AI Flash: AI coding company Magic released a pre-training research report, claiming that a new training approach can achieve pre-training results close to DeepSeek V4 Pro Base with only about $500,000 worth of GB200 compute, requiring roughly 1/50 of the computational cost.
Base refers to a base model that has only completed pre-training and has not yet undergone post-training such as reinforcement learning. Magic tested the models using unseen code, math problems, and research papers to see which one predicts subsequent content more accurately. Therefore, this comparison focuses on the pre-training effectiveness of the base models and does not imply that chat, coding, or agent capabilities have already matched the finished version of DeepSeek V4 Pro.
Magic then scaled up the training by 10 times, equivalent to about $4 million worth of GB200 compute. In the same set of self-tests, this version surpassed the latest public Base models from DeepSeek, Kimi, and NVIDIA that it evaluated. Magic estimates that if it were to achieve the same level using DeepSeek V4 Pro's training efficiency, the compute cost would exceed $100 million.
Magic has not disclosed the full training recipe, only revealing that the efficiency gains come from dozens of changes across model architecture, optimizers, training objectives, and data processing. The model weights have also not been released yet, so the "50x efficiency" claim cannot be independently reproduced by external parties.

