header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

MiniMax has released the full-modal generation model H3, a model that handles images, videos, and audio all in one.

According to Perceive Beating monitoring, MiniMax has released the omni-modal generation model H3. It is capable of understanding text, images, videos, and audio simultaneously, and can generate or edit videos based on a natural language command.

Users can instruct H3 to mimic camera movement from one video, use a person from another image, and refer to the sound from a third audio clip. Actions that previously required separate calls for motion transfer, person reference, sound reference, and video editing can now be consolidated into a single task.

H3 can generate videos up to 15 seconds long, supporting 2K resolution and native stereo sound. The official API prices 2K at $0.13 per second, making the cost of generating a 15-second video approximately $1.95; 768P is priced at $0.09 per second.

Each task can input a maximum of 9 images, 3 videos, and 3 audio clips, not exceeding a total of 12 files.

MiniMax plans to release the model weights in the next few days. The official evaluation of models such as Seedance and Veo has not yet been announced.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish