Multimodal video

Limited-Time Offer: 50 Credits

Fast creative iteration with synchronized native audio — now 50 credits (regular 88 credits) for 480P 5-second video.

Try MiniMax H3 Max

MiniMax H3 Max on Voor AI — move from text or a first frame to a 5–15 second video with synchronized native audio, built for fast creative iteration.

Example Video
MiniMax H3 Max — Rain-lit trail shoe thumbnailMiniMax H3 Max — Last table after closing thumbnailMiniMax H3 Max — A paper city wakes thumbnailMiniMax H3 Max — Animate an approved first frame thumbnail

Speed-tuned H3 · native audio

MiniMax H3 Max turns fast inference into a tighter review loop

MiniMax H3 Max is a speed-tuned MiniMax H3 variant. Its useful idea is simple: a video model that returns direction checks quickly enough to compare several shots while the brief is still fresh. You can begin with text or an optional first frame, choose 5–15 seconds and 480P or 768P, and ask for synchronized dialogue, effects, ambience, or music in the same pass as the picture.

The product shot beside this copy is not a launch reel or borrowed showcase footage. It is one of four Voor test renders made on August 28, 2026 using the same MiniMax H3 Max model available above. We kept every test at 480P and five seconds so the first question was creative rather than expensive: did the model follow the action, camera, material, and sound brief closely enough to deserve a higher-resolution pass?

Choose the H3 route that fits the job.H3 Max prioritizes speed and prompt adherence. The original MiniMax H3 remains a separate route with different resolution and reference options.
Voor test render · Text to video5.2s · 480P · AAC stereo
5–15sWhole-second duration control
480P / 768PDraft and review tiers
6 ratiosFrom 21:9 through 9:16 in text mode
Native audioDialogue, effects, ambience, and music

Four live test renders

Inspect the output, prompt, and timing together

Every clip below is 832×480, 5.184 seconds, H.264 video with AAC stereo audio. Our test recorded 0.738–0.789 seconds of model inference across the set. That timing excludes queueing, processing, download, and your network, so it is an observed detail—not a promise that every job completes in under a second. Play with sound and review the named failure point instead of judging only the poster frame.

Text to video

Rain-lit trail shoe

0.738s inference
Premium live-action product film. A matte graphite trail shoe rests on a rain-dark basalt plinth at blue hour. One slow quarter-orbit reveals the sole and heel; beads of water roll naturally across the fabric while a cool rim light travels over the silhouette. Preserve the exact shoe geometry and lace pattern. Close rain ambience, one soft footstep on gravel, no music, no text, no logo. End on a stable three-quarter packshot with negative space on the left.

Review: Watch the laces, heel, and tread through the orbit, then listen for the single gravel step under the rain bed.

Text to video

Last table after closing

0.751s inference
Live-action cinematic medium shot in a quiet restaurant kitchen after closing. A tired chef plates one final bowl, looks toward the camera and says exactly: "Last table. Make it count." One restrained push-in, realistic hand movement, warm practical lights, stainless reflections that stay consistent. Plate contact, ventilation hum and a distant rainstorm; no music, no cut, no text. Hold the finished bowl for the last second.

Review: Check the short spoken line, lip timing, hand-to-bowl contact, and whether the warm reflections stay attached to the kitchen.

Text to video

A paper city wakes

0.776s inference
A handcrafted paper city wakes at sunrise in one continuous macro shot. Windows unfold from cream cardstock, tiny paper bicycles roll between buildings, and a folded tram rounds the corner as the camera tracks beside it at tabletop height. Every object remains visibly made from cut and folded paper with crisp edges and believable contact. Paper rustle, tiny wheel clicks and soft room tone; no text, no logo, no music.

Review: Look for material continuity: every facade, bicycle, and moving tram should continue to read as cut paper.

Image to video

Animate an approved first frame

0.789s inference
Preserve the baker, bread, counter, predawn light and camera position from the first frame. The baker slices the loaf once, steam rises, then they glance up with a small smile while the camera makes a very slow two-percent push. Natural hand contact, stable face and bread geometry. Crisp crust sound, quiet room tone and one distant street bell; no music, no cut. Hold the final frame.

Review: Compare the first frame with the moving face, apron, loaves, and doorway, then check that the bread sound has a visible cause.

A three-pass runway

Fast is useful only when each pass answers one question

Speed can reduce the cost of being wrong, but it can also produce a folder full of unlabelled variations. Give every MiniMax H3 Max pass one job. Start with a cheap direction check, correct the largest visible miss, and raise resolution only when the camera, action, and ending are already usable. Keep the prompt and settings beside the file so the reason for each change survives the session.

  1. Pass 01 · direction5 seconds · 480P · Balanced

    Test the subject, one action, one camera move, and the basic sound idea. Ignore polish. Reject the direction if the core event is not readable without your explanation.

  2. Pass 02 · correctionChange one layer, not the whole brief

    Keep the accepted subject and camera. Fix the largest error: contact, identity, timing, dialogue, or the landing frame. One controlled difference makes comparison possible.

  3. Pass 03 · review768P · chosen duration · sound on

    Move up only after direction survives. Review first, middle, and last frames; then listen once without looking. Download the pass you approve before exploring another idea.

Choose the evidence you have

Text for open direction, frames for continuity

Text to video

Use text mode when casting, composition, and art direction are all open. It exposes six aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. That makes it the better branch for comparing a cinematic wide, square feed concept, and vertical social shot without preparing source media first. The freedom is also the risk: a product or person the model never saw cannot be expected to match an approved design.

Image to video

Use image mode when a first frame already answers who, what, where, and how the scene is framed. The output follows that image's aspect ratio. Add an end image when the final pose or composition matters, but choose a destination the subject could physically reach within the duration. If you are still choosing among models, the Image to Video collection puts other still-first options beside it.

Original MiniMax H3

Choose the original H3 page when 2K output or its broader reference workflow is a hard requirement. Choose H3 Max when 480P or 768P is sufficient and fast exploration matters more. “Max” is not automatically the higher-resolution choice here; it identifies the speed-tuned route. Compare the real MiniMax H3 examples before assigning the final job. For a brief you want expanded before rendering, try MiniMax H3 Max Turbo with Balanced or Quality mode. Its dedicated workspace keeps text and first-frame generation together for 5–15 second clips with native audio at 480P or 768P.

Interactive Live Classroom

Experience MiniMax H3 Max in real time inside our Live Classroom, where the model animates consecutive 1970s cartoon lesson beats continuously on a retro CRT television stage.

A brief you can debug

Six layers for a MiniMax H3 Max prompt

Balanced prompt expansion is the practical default: it turns a compact direction into a fuller production brief without the longer delay of Quality mode. Expansion cannot rescue a contradictory idea. Separate the layers below so you can change one decision on the next pass instead of asking the model to “make it better.”

  1. 01

    Visible subject

    Name the person, object, material, wardrobe, or approved first frame that must remain recognizable. If identity matters, one uploaded image is stronger evidence than five adjectives.

  2. 02

    One action arc

    Write a readable beginning, change, and finish. Five seconds can carry a glance, a reveal, or a short product move. It cannot carry a montage disguised as one sentence.

  3. 03

    One camera move

    Choose a push, track, orbit, pan, or locked frame. State its speed and what it follows. Several competing moves make it harder to tell whether the model followed any of them.

  4. 04

    Light and material

    Describe the practical source of light and the surface response you need to preserve: wet basalt, warm steel reflections, paper grain, skin texture, or the exact geometry of a product.

  5. 05

    Sound with causes

    Put dialogue in quotation marks. Tie effects to visible actions, then give ambience and music their own clauses. A plate can clink; a vent can hum; rain can live beyond the window.

  6. 06

    A reviewable ending

    Ask for a one-second hold, a stable packshot, or a named end frame. A deliberate landing point makes the output easier to compare, cut, and approve.

Compact product briefOne matte graphite trail shoe on a rain-dark basalt plinth at blue hour. Make one slow quarter-orbit from side profile to a three-quarter heel view. Preserve the shoe geometry, lace pattern, and fabric while water beads roll naturally. Cool rim light, close rain ambience, one gravel step, no music, no text, no logo. Hold a stable packshot with negative space on the left.

Review picture and sound

Do not let a fast result skip the approval pass

Watch once at normal speed, once while scrubbing, and once with your eyes away from the screen. On the picture pass, check identity, joints, contact, reflections, object geometry, camera continuity, and the last frame. On the audio pass, check that speech is exact, effects have visible causes, room tone belongs to the space, and the mix does not hide the most important cue. The four examples above all contain real AAC stereo tracks, including the quiet bakery room tone; a muted autoplay preview cannot prove that part of the result.

Use H3 Max to find the shot, not to pretend review is finished. For final delivery, verify the exported dimensions, listen on the device where the work will appear, and label generated scenes honestly. Compare the same prompt in Gemini Omni 1.1 Flash vs H3 Maxwhen 4K or mixed references matter, or inspect H3 Max vs Wan 3.0when the accepted brief outgrows this route's 768P or fifteen-second ceiling.

A six-point acceptance note

  1. 01 The main action reads without extra explanation.
  2. 02 Identity and product geometry survive every frame.
  3. 03 Camera motion is continuous and motivated.
  4. 04 Dialogue is exact enough for the intended use.
  5. 05 Effects and ambience match visible causes and space.
  6. 06 The final frame is stable enough to cut or loop.

To steer upcoming scenes while watching a continuous stream, use H3 Max Director and send new directions during the session.

MiniMax H3 Max FAQ

What is MiniMax H3 Max?

MiniMax H3 Max is a speed-focused MiniMax H3 variant for rapid creative iteration. Its two modes create video from text or an optional first frame, with native synchronized audio, 5–15 second duration, and 480P or 768P output.

Is MiniMax H3 Max the same as MiniMax H3?

No. MiniMax H3 is the original model route and is the better choice when you need its 2K or broader reference workflows. H3 Max is tuned for faster iteration at 480P or 768P. Treat them as related tools with different production priorities, not as interchangeable version names.

Can MiniMax H3 Max generate audio?

Yes. H3 Max preserves MiniMax H3's native synchronized audio capability. Put exact dialogue in quotation marks, then describe ambience, effects, and music as separate clauses. Review the finished MP4 with sound on; good motion does not guarantee that every voice or effect lands correctly.

How long can a MiniMax H3 Max video be?

Both generation modes accept whole-second durations from 5 through 15 seconds. A short five-second pass is useful for testing direction, while 10–15 seconds gives an action or dialogue beat more room. For a longer sequence, generate separate shots and assemble only approved clips in an editor.

Which MiniMax H3 Max resolution should I choose?

Choose 480P for cheap, rapid direction checks and 768P when the shot is ready for a more detailed review or delivery. H3 Max does not expose 1080p or 2K output. Use the original MiniMax H3 route when 2K is a hard requirement.

Which aspect ratios does MiniMax H3 Max support?

Text to video supports 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Image to video follows the first image's aspect ratio. If no first image is supplied, it behaves like text to video and defaults to 16:9.

Can MiniMax H3 Max use a first and last frame?

Yes. Image to video accepts an optional first image and an optional end image. A first frame anchors subject, composition, and light; an end frame adds a target landing point. Keep both frames visually compatible so the requested transition can happen naturally within 5–15 seconds.

What does prompt expansion do in MiniMax H3 Max?

Disabled sends the prompt without expansion, Balanced performs a quick rewrite and is the default, and Quality can spend up to roughly 30 seconds building a richer production brief. Use Balanced for fast exploration and Quality only after the creative direction is already worth the extra wait.

How fast is MiniMax H3 Max?

H3 Max is built for faster H3 iteration. In four Voor 480P tests on August 28, 2026, measured model inference was 0.738–0.789 seconds for each 5-second clip. Queue time, prompt expansion, transfer, and current demand are separate, so that observation is evidence from this test set rather than a guaranteed response time.

Can MiniMax H3 Max be used for continuous video or education?

Yes. Voor's Live Classroom uses MiniMax H3 Max to animate sequential 5-second educational cartoon scenes on a live 1970s classroom TV stage. Anyone can queue custom lesson topics to watch the model illustrate concepts beat-by-beat.

Make the next MiniMax H3 Max pass

Start at 480P, judge one shot decision at a time, then move the approved direction to 768P.

Create with MiniMax H3 Max