header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

AMD has released the fully open-source MoE Large Model, Instella-MoE, with a parameter scale of 16B, challenging mainstream open-source models.

BlockBeats News, July 25th, AMD announced the launch of the fully open-source Mixtures of-Experts (MoE) model, Instella-MoE, which boasts 16 billion total parameters, with each Token activating 2.8 billion parameters, claiming to have leading performance among open-source language models of the same scale.


AMD stated that Instella-MoE is entirely based on its own AMD Instinct MI300X and MI325X GPUs, as well as the ROCm software stack, trained from scratch, and adopts architecture innovations such as Gated Multi-head Latent Attention (Gated MLA) and FarSkip-Collective to improve training and inference efficiency.


Performance tests show that the Instella-MoE-16B-A3B base model achieves an average score of 76.7, leading among open-source models, surpassing models like SmolLM3-3B, OLMo-3-7B, and competing with larger-scale models even with only 2.8 billion activated parameters.


Additionally, the model supports 64K Token length context processing and completes the full training process, including pre-training, mid-stage training, long-context expansion, supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement learning (RL).


AMD has simultaneously released all model weights, training configurations, data ratios, intermediate checkpoints, and inference code of Instella-MoE to promote open AI model research and reproducibility.


AMD stated that Instella-MoE showcases the capability of training large-scale MoE models based on AMD hardware and an open software ecosystem, and will continue to advance the development of open-source language models with larger scales, stronger inference capabilities, and higher efficiency in the future.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish