header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

MiniMax Preview Generation Model: A single model for both live and post-production preview based on the H3 architecture

According to Dynasty Beating monitoring, the MiniMax H3 team revealed in a Reddit AMA that they are developing a dedicated image model and plan to open-source the weights. This model will combine textual image generation and general image editing into a single framework and is currently in the post-training optimization phase.

This model shares the same architectural philosophy as H3. It will utilize H3's VAE Encoder and introduce a separately designed VAE Decoder for image generation. The envisioned workflow by MiniMax is to first use the image model to generate the initial frame, which will then be passed to H3 to continue generating the video.

The team also mentioned that H3 itself has shown potential for both image generation and editing. Previously, they only trained the model to predict the final frame based on the "initial frame + text description," without specific image editing training. However, the model still demonstrated strong zero-shot capabilities in various image editing evaluations. This serves as a key rationale for MiniMax to extend this architecture to an image model.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish