# MiniMax H3 Prompting Guide: 2K Video With Native Audio

> How to prompt MiniMax H3 for 2K video with native audio and readable on-screen text. Reference roles, timecoded shot lists, audio direction, credit costs and the mistakes that waste generations.

- Published: 2026-08-07
- Updated: 2026-08-07
- Author: Keira (Founder's Associate, VidGuy)
- Tags: MiniMax H3, AI Video Models, Prompting Guide

MiniMax H3 is an omni multimodal video model: it takes text, images, video clips and audio in the same request, and it generates picture and sound together in one pass. In VidGuy it produces 2K video with native audio, from 4 to 15 seconds.

That "omni" part is the whole story. Most video models take a prompt and maybe one image. H3 takes a small pile of reference material and expects you to tell it what each piece is *for*. Get that right and it is remarkably obedient. Skip it and you get an expensive slideshow.

Here is how to actually prompt it.

## The one habit that matters: give every reference a job

The single biggest difference between a good H3 prompt and a bad one is whether you told the model what each attached file controls.

Do not do this:

> Make a video of this woman with this bag in this location.

Do this:

> Use Image 1 for the overall mood, location and film texture. Use Image 2 for the talent — preserve her face, the half-up black hair and the indigo ribbon exactly. Use Image 3 for the bag; the hardware and stitching must match.

Same three images. The second version tells the model that Image 1 is a *lighting and grade* reference, Image 2 is an *identity lock*, and Image 3 is a *product* it must not redesign. H3 will respect that division. Without it, the model averages your references into a mush and you cannot tell which one caused the problem.

The same logic applies to every input type:

- **Images** lock identity and detail. Faces, products, sets, typography.
- **Video clips** transfer motion, camera language and pacing. "Match the camera move in Video 1" is a real instruction.
- **Audio clips** transfer voice. "Match the voice in Audio 1" clones delivery, not just timbre.

In reference-to-video mode you can attach up to twelve files — roughly nine images, three video clips and three audio clips. That is a lot of control, and all of it is wasted if the prompt does not say which is which.

## Type the words you want on screen

Legible on-screen text is the thing H3 is obviously better at than its competitors, and it is the easiest win available to a marketer.

The rule is blunt: **if a word has to be readable, type it.** Describing it gets you text-shaped pixels. Typing it gets you your text.

> The words **SUMMER DROP** appear centred, in wide-tracked cinematic typography — not pure white, with restrained material texture, subtle illumination and a faint edge glow.

Then direct how it behaves, because the default is usually too much:

> The type resolves with the focus: begin slightly blurred at low opacity, then fade into clarity over 0.3 to 0.5 seconds. No spins, bounces or large fly-ins.

And say what you do not want. Negative direction is unusually effective here:

> No garbled characters, no misspellings, no non-Latin glyphs.

That last line is worth including every single time you put text on screen. It is the difference between a usable clip and a re-run.

## Timecode anything longer than one beat

H3 will happily drift into a slideshow if you give it a list of things that happen with no sense of when. A timecoded shot list fixes it:

> [0 to 2s] High-angle overhead shot of the desk, slow push in.
> [2 to 4s] Cut to her hands opening the box, shallow depth of field.
> [4 to 6s] Rack focus to the product, camera settles, title resolves.

You are not obliged to use every second. You are telling the model where the beats fall, which is what stops it from spending nine seconds on the first idea.

Prompts can run to about 7,000 characters, so a complete shot list with sound design fits comfortably in one request. Use the room.

## Direct the camera like a DP, not like a menu

H3 responds to cinematography vocabulary far better than to effect names. Describe the physical behaviour of a camera and lens:

> Subtle handheld shake, then push in quickly and rack focus. Wide-angle lens with strong perspective distortion, backlit exposure breathing, slightly coarse noise in the shadows.

Transitions work the same way. Instead of "whip pan transition", describe what the camera actually does:

> Fast binocular-scan transition with whip movement, motion blur, optical smearing and a brief exposure flicker.

Naming an effect gets you the model's average idea of that effect. Describing the physics gets you something that looks shot.

## Art-direct the audio — it is being generated either way

This is the most commonly wasted capability. H3 generates sound natively, in the same pass. If you do not direct it, you still get audio; you just get audio nobody chose.

Treat it as its own department:

> Sound: a deep sub-bass pulse, distant metallic resonance, and one restrained hit as the title locks into focus.

For music, specify instrumentation and structure rather than genre:

> Low drone, tense string pizzicato, cool synth pulses, a low kick, sparse brushwork, fragments of walking bass.

"Cinematic music" is a wish. The line above is a brief.

For dialogue, you can replace a line while keeping the performance:

> Replace the woman's line with the line from Audio 1, matched to her mouth movement.

## A prompt template that works

Assemble in this order. It maps to how the model reads the request.

1. **Reference roles** — what each attached file controls
2. **Shot list** — timecoded beats
3. **Camera and lens** — physical behaviour
4. **Identity locks** — the specific details that must survive
5. **On-screen text** — typed verbatim, with animation behaviour
6. **Sound** — music, effects, dialogue
7. **Negative direction** — what must not happen

Worked example:

> Image 1 sets mood, location and film texture. Image 2 is the talent — preserve the face, freckles, and the rust-coloured jacket exactly. Image 3 is the product; do not redesign the label.
>
> [0 to 3s] Wide shot of the rooftop at golden hour, slow dolly left.
> [3 to 7s] Cut in to her hands turning the bottle to camera, shallow focus.
> [7 to 10s] She looks up and smiles; camera settles; title resolves.
>
> Anamorphic 40mm feel, gentle handheld float, warm backlight with visible haze, fine grain in the shadows.
>
> The words **FIELD NOTES** appear lower-third, wide-tracked, warm off-white with a faint edge glow, fading from blur to clarity over 0.4 seconds.
>
> Sound: soft rooftop ambience, distant traffic, a single warm synth swell landing with the title.
>
> No text distortion, no garbled characters, no fly-in animation, no lens flares.

## What it costs in VidGuy

In VidGuy, H3 runs at **3 credits per second**, in a 4 to 15 second range — so a 4-second clip is 12 credits and a full 15-second clip is 45. It sits on the Pro plan and above.

Because H3 generates audio in the same pass, that price includes the sound design. Compare it against the real alternative, which is generating a silent clip and then paying for music, effects and a voiceover separately.

## Common mistakes

**Passing references without roles.** The most expensive mistake, and the easiest to fix.

**Assuming identity is preserved.** It is not, unless you name the features. "Preserve the half-up black hair, the silver crown, the indigo ribbon" is what keeps a character stable across a shot.

**Leaving audio to chance.** It generates regardless. Direct it.

**No negative direction.** One line of "no soft dissolves, no morphs, no garbled text" prevents most re-runs.

**Describing text instead of typing it.** Say the word, in quotes, exactly as it should appear.

**No timecodes on multi-beat shots.** Without them, pacing drifts and you get one idea stretched thin.

## Frequently asked questions

**How long can a MiniMax H3 video be?**
In VidGuy, 4 to 15 seconds in a single generation, at 24 frames per second.

**What resolution does MiniMax H3 output?**
2K — roughly 1440 pixels on the short edge, across aspect ratios from 21:9 through to 9:16.

**Does MiniMax H3 generate sound?**
Yes, natively and in the same pass as the picture, including music, effects and dialogue with lip-sync. This is a genuine difference from most video models, which are silent by default.

**Can MiniMax H3 render readable text?**
Yes, and it is one of its strongest features. Type the exact words rather than describing them, and add negative direction against garbled characters.

**How many reference files can I attach?**
Up to twelve in reference-to-video mode — around nine images, three video clips and three audio clips. Label each one's job in the prompt.

**Is MiniMax H3 better than Seedance?**
They solve different problems. H3 is the choice for short, high-resolution work where sound and on-screen text matter. Seedance 2.5 is the choice when you need one continuous shot far longer than fifteen seconds. Most teams end up using both.

## Try it

MiniMax H3 is available now in VidGuy Studio — open the model picker under the prompt box and choose **MiniMax H3 (Omni)**. Bring your references, give each one a job, and type the words you want people to read.


---

- Canonical page: https://www.vidguy.ai/blog/minimax-h3-prompting-guide
- About VidGuy (machine-readable overview): https://www.vidguy.ai/llms.txt
- Pricing (machine-readable): https://www.vidguy.ai/pricing.md
- Get started: https://www.vidguy.ai/auth/signup
