MiniMax H3 Launches With Multimodal Input and Native Audio
- Aug 3
- 4 min read

"Multimodal" has become one of those words that stops carrying information. Nearly every model announced in the past year claims it, and the claims describe wildly different things — anything from accepting an image alongside a prompt to genuinely reasoning across several kinds of material at once.
So it's worth being specific about what MiniMax H3 does with its inputs, and what it does on the way out, because both are more concrete than the label suggests.
Four inputs, one context
H3 accepts text, images, video, and audio, and processes them in a single shared context rather than routing each type through its own pathway and merging the conclusions afterwards.
The distinction matters when references disagree. If you supply a still that fixes a character's face and a reference clip where a different person performs the movement you want, something has to decide which source governs which property. A model holding everything simultaneously can resolve that. A pipeline that handles each input separately tends to produce an average of your references instead of an application of them.
What this enables in practice is a division of labour across sources. Stills fix identity — this face, this product, this exact garment. A reference clip fixes motion, performance rhythm, and camera language. An audio sample fixes voice and delivery. An existing edit can carry cutting rhythm and colour feel to new material. A reference set can hold up to twelve mixed files, with video and audio references capped at fifteen seconds each.
The consequence is that you stop describing and start showing. Character consistency, which used to be a prompt-engineering discipline maintained through saved paragraphs of adjectives, becomes a matter of supplying the same reference.
Native audio is the other half
The output side is where Minimax H3 makes its more unusual bet. Every generation arrives with native stereo audio produced in the same pass as the picture: dialogue, room tone, foley, music, and the timing relationships between them.
The industry default is a chain — one model for pictures, another for speech, a library for effects, a person to align them. That chain exists for defensible reasons, and it produces a characteristic failure. Sync achieved afterwards is approximate sync, and approximate is precisely where the eye catches it. Most content that reads as machine-made gives itself away through audio rather than through pictures.
The evidence that joint generation is doing real work shows up on the hardest material. Fast rap delivery, where syllable density makes a two-frame drift visible, or a line where breath placement matters as much as the words. Post-hoc alignment fails there almost by definition.
There's also a compounding effect. Because dialogue can be replaced with the mouth re-forming around new words, and because vocal timbre can be cloned or transferred from a reference, a language variant stops being a dub or a reshoot and becomes an edit. That only works if the model owns the audio; bolt-on pipelines reintroduce the seam at exactly the point you were trying to remove it.
Multimodal input plus editing
The launch framing puts input and audio in the headline, but the capability that ties them together is editing existing footage.
H3 modifies a designated element while holding the rest of the frame stable — replacing, adding, or removing objects and people; changing backgrounds, environments, and lighting; adjusting effects; modifying motion and performance; rewriting a line; transferring a voice. It currently leads the Artificial Analysis video editing leaderboard ahead of Seedance 2.0.
Precise editing is what multimodal input is for. "Make her jacket the same red as before" isn't a specification, it's a hope; an image reference is a specification. Once instructions can be comparative rather than descriptive — here's the footage, here's what should change, here's the source of truth for what it should become — editing becomes controllable rather than approximate.
The rest of the sheet
Clips run 5–15 seconds at 24 FPS, up to 1440p, across 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Prompts accept up to 7,000 characters. First/last frame mode preserves the input image's aspect ratio; all-purpose reference mode infers an appropriate output ratio.
Two of these read as limitations and aren't. 24 FPS is the frame rate of film and commercials, so output conforms into a real timeline without conversion. 1440p is the threshold at which typography, product detail, and interface elements stay legible through motion — which is what commercial work actually requires, as against a 4K figure that gets downscaled before anyone sees it. That interface stability is a genuine differentiator: game UI, app flows, and product screens survive camera movement rather than dissolving into letter-shaped noise.
What it doesn't claim
Fifteen seconds is a shot, not a film. Assembly remains a human job, and deciding order and rhythm is where a piece succeeds or fails.
Local editing is strong but bounded — the larger the change, the more of the frame it disturbs, and past a certain point you're regenerating with extra steps. There's no engine export, no alpha or scene data for compositing, and no way to photograph a real product exactly as it exists.
The pricing position is the last piece and does a fair amount of the persuading: per-second cost sits well below comparable models, Seedance 2.0 included, which matters less as a saving than as a quality mechanism, since good output is always the survivor of discarded attempts. Anyone assessing it should start with Minimax H3 Free on their own material rather than trusting a launch summary, including this one.
The short version
Two things distinguish this release. Inputs of four different kinds are reasoned about together rather than handled separately, which makes precise specification possible. And sound is generated with the picture rather than after it, which closes the seam where synthetic content usually reveals itself.
Whether those were the right bets will be answered by what teams are still using in a year — not by a benchmark, and not by the word in the headline.


