MiniMax H3 Max turns fast inference into a tighter review loop
MiniMax H3 Max is a speed-tuned MiniMax H3 variant. Its useful idea is simple: a video model that returns direction checks quickly enough to compare several shots while the brief is still fresh. You can begin with text or an optional first frame, choose 5–15 seconds and 480P or 768P, and ask for synchronized dialogue, effects, ambience, or music in the same pass as the picture.
The product shot beside this copy is not a launch reel or a borrowed provider demo. It is one of four Voor API smoke tests made on August 28, 2026 using the exact endpoint wired into the generator above. We kept every test at 480P and five seconds so the first question was creative rather than expensive: did the model follow the action, camera, material, and sound brief closely enough to deserve a higher-resolution pass?
Choose the H3 route that fits the job.H3 Max prioritizes speed and prompt adherence. The original MiniMax H3 remains a separate route with different resolution and reference options.
Voor API test · Text to video5.2s · 480P · AAC stereo
5–15sWhole-second duration control
480P / 768PDraft and review tiers
6 ratiosFrom 21:9 through 9:16 in text mode
Native audioDialogue, effects, ambience, and music
Four live endpoint calls
Inspect the output, prompt, and timing together
Every clip below is 832×480, 5.184 seconds, H.264 video with AAC stereo audio. The API reported 0.738–0.789 seconds of backend inference across the set. That timing excludes queueing, prompt expansion, download, and your network, so it is a measured detail from these calls—not a promise that every job completes in under a second. Play with sound and review the named failure point instead of judging only the poster frame.
Text to video
Rain-lit trail shoe
0.738s inference
Premium live-action product film. A matte graphite trail shoe rests on a rain-dark basalt plinth at blue hour. One slow quarter-orbit reveals the sole and heel; beads of water roll naturally across the fabric while a cool rim light travels over the silhouette. Preserve the exact shoe geometry and lace pattern. Close rain ambience, one soft footstep on gravel, no music, no text, no logo. End on a stable three-quarter packshot with negative space on the left.
Review: Watch the laces, heel, and tread through the orbit, then listen for the single gravel step under the rain bed.
Text to video
Last table after closing
0.751s inference
Live-action cinematic medium shot in a quiet restaurant kitchen after closing. A tired chef plates one final bowl, looks toward the camera and says exactly: "Last table. Make it count." One restrained push-in, realistic hand movement, warm practical lights, stainless reflections that stay consistent. Plate contact, ventilation hum and a distant rainstorm; no music, no cut, no text. Hold the finished bowl for the last second.
Review: Check the short spoken line, lip timing, hand-to-bowl contact, and whether the warm reflections stay attached to the kitchen.
Text to video
A paper city wakes
0.776s inference
A handcrafted paper city wakes at sunrise in one continuous macro shot. Windows unfold from cream cardstock, tiny paper bicycles roll between buildings, and a folded tram rounds the corner as the camera tracks beside it at tabletop height. Every object remains visibly made from cut and folded paper with crisp edges and believable contact. Paper rustle, tiny wheel clicks and soft room tone; no text, no logo, no music.
Review: Look for material continuity: every facade, bicycle, and moving tram should continue to read as cut paper.
Image to video
Animate an approved first frame
0.789s inference
Preserve the baker, bread, counter, predawn light and camera position from the first frame. The baker slices the loaf once, steam rises, then they glance up with a small smile while the camera makes a very slow two-percent push. Natural hand contact, stable face and bread geometry. Crisp crust sound, quiet room tone and one distant street bell; no music, no cut. Hold the final frame.
Review: Compare the first frame with the moving face, apron, loaves, and doorway, then check that the bread sound has a visible cause.
A three-pass runway
Fast is useful only when each pass answers one question
Speed can reduce the cost of being wrong, but it can also produce a folder full of unlabelled variations. Give every MiniMax H3 Max pass one job. Start with a cheap direction check, correct the largest visible miss, and raise resolution only when the camera, action, and ending are already usable. Keep the prompt and settings beside the file so the reason for each change survives the session.
Pass 01 · direction5 seconds · 480P · Balanced
Test the subject, one action, one camera move, and the basic sound idea. Ignore polish. Reject the direction if the core event is not readable without your explanation.
Pass 02 · correctionChange one layer, not the whole brief
Keep the accepted subject and camera. Fix the largest error: contact, identity, timing, dialogue, or the landing frame. One controlled difference makes comparison possible.
Pass 03 · review768P · chosen duration · sound on
Move up only after direction survives. Review first, middle, and last frames; then listen once without looking. Download the pass you approve before exploring another idea.
Choose the evidence you have
Text for open direction, frames for continuity
Text to video
Use text mode when casting, composition, and art direction are all open. It exposes six aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. That makes it the better branch for comparing a cinematic wide, square feed concept, and vertical social shot without preparing source media first. The freedom is also the risk: a product or person the model never saw cannot be expected to match an approved design.
Image to video
Use image mode when a first frame already answers who, what, where, and how the scene is framed. The output follows that image's aspect ratio. Add an end image when the final pose or composition matters, but choose a destination the subject could physically reach within the duration. If you are still choosing among models, the Image to Video collection puts other still-first options beside it.
Original MiniMax H3
Choose the original H3 page when 2K output or its broader reference workflow is a hard requirement. Choose H3 Max when 480P or 768P is sufficient and fast exploration matters more. “Max” is not automatically the higher-resolution choice here; it identifies the speed-tuned route. Compare the real MiniMax H3 examples before assigning the final job.
A brief you can debug
Six layers for a MiniMax H3 Max prompt
Balanced prompt expansion is the practical default: it turns a compact direction into a fuller production brief without the longer delay of Quality mode. Expansion cannot rescue a contradictory idea. Separate the layers below so you can change one decision on the next pass instead of asking the model to “make it better.”
01
Visible subject
Name the person, object, material, wardrobe, or approved first frame that must remain recognizable. If identity matters, one uploaded image is stronger evidence than five adjectives.
02
One action arc
Write a readable beginning, change, and finish. Five seconds can carry a glance, a reveal, or a short product move. It cannot carry a montage disguised as one sentence.
03
One camera move
Choose a push, track, orbit, pan, or locked frame. State its speed and what it follows. Several competing moves make it harder to tell whether the model followed any of them.
04
Light and material
Describe the practical source of light and the surface response you need to preserve: wet basalt, warm steel reflections, paper grain, skin texture, or the exact geometry of a product.
05
Sound with causes
Put dialogue in quotation marks. Tie effects to visible actions, then give ambience and music their own clauses. A plate can clink; a vent can hum; rain can live beyond the window.
06
A reviewable ending
Ask for a one-second hold, a stable packshot, or a named end frame. A deliberate landing point makes the output easier to compare, cut, and approve.
Compact product briefOne matte graphite trail shoe on a rain-dark basalt plinth at blue hour. Make one slow quarter-orbit from side profile to a three-quarter heel view. Preserve the shoe geometry, lace pattern, and fabric while water beads roll naturally. Cool rim light, close rain ambience, one gravel step, no music, no text, no logo. Hold a stable packshot with negative space on the left.
Review picture and sound
Do not let a fast result skip the approval pass
Watch once at normal speed, once while scrubbing, and once with your eyes away from the screen. On the picture pass, check identity, joints, contact, reflections, object geometry, camera continuity, and the last frame. On the audio pass, check that speech is exact, effects have visible causes, room tone belongs to the space, and the mix does not hide the most important cue. The four examples above all contain real AAC stereo tracks, including the quiet bakery room tone; a muted autoplay preview cannot prove that part of the result.
Use H3 Max to find the shot, not to pretend review is finished. For final delivery, verify the exported dimensions, listen on the device where the work will appear, and label generated scenes honestly. The AI video model collection is useful when the accepted brief outgrows this route's 768P ceiling.
A six-point acceptance note
01 The main action reads without extra explanation.
02 Identity and product geometry survive every frame.
03 Camera motion is continuous and motivated.
04 Dialogue is exact enough for the intended use.
05 Effects and ambience match visible causes and space.
06 The final frame is stable enough to cut or loop.
MiniMax H3 Max FAQ
What is MiniMax H3 Max?
MiniMax H3 Max is a speed-focused MiniMax H3 variant for rapid creative iteration. Its two API modes create video from text or an optional first frame, with native synchronized audio, 5–15 second duration, and 480P or 768P output.
Is MiniMax H3 Max the same as MiniMax H3?
No. MiniMax H3 is the original model route and is the better choice when you need its 2K or broader reference workflows. H3 Max is tuned for faster iteration at 480P or 768P. Treat them as related tools with different production priorities, not as interchangeable version names.
Can MiniMax H3 Max generate audio?
Yes. H3 Max preserves MiniMax H3's native synchronized audio capability. Put exact dialogue in quotation marks, then describe ambience, effects, and music as separate clauses. Review the finished MP4 with sound on; good motion does not guarantee that every voice or effect lands correctly.
How long can a MiniMax H3 Max video be?
Both endpoints accept whole-second durations from 5 through 15 seconds. A short five-second pass is useful for testing direction, while 10–15 seconds gives an action or dialogue beat more room. For a longer sequence, generate separate shots and assemble only approved clips in an editor.
Which MiniMax H3 Max resolution should I choose?
Choose 480P for cheap, rapid direction checks and 768P when the shot is ready for a more detailed review or delivery. H3 Max does not expose 1080p or 2K output. Use the original MiniMax H3 route when 2K is a hard requirement.
Which aspect ratios does MiniMax H3 Max support?
Text to video supports 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Image to video follows the first image's aspect ratio. If no first image is supplied to that endpoint, it behaves like text to video and defaults to 16:9.
Can MiniMax H3 Max use a first and last frame?
Yes. The image-to-video endpoint accepts an optional first image and an optional end image. A first frame anchors subject, composition, and light; an end frame adds a target landing point. Keep both frames visually compatible so the requested transition can happen naturally within 5–15 seconds.
What does prompt expansion do in MiniMax H3 Max?
Disabled sends the prompt without expansion, Balanced performs a quick rewrite and is the default, and Quality can spend up to roughly 30 seconds building a richer production brief. Use Balanced for fast exploration and Quality only after the creative direction is already worth the extra wait.
How fast is MiniMax H3 Max?
H3 Max is built for faster H3 iteration. In Voor's four 480P smoke tests on August 28, 2026, the API reported 0.738–0.789 seconds of backend inference for each 5-second clip. Queue time, prompt expansion, transfer, and current demand are separate, so that measurement is evidence from this test set rather than a guaranteed response time.
Make the next MiniMax H3 Max pass
Start at 480P, judge one shot decision at a time, then move the approved direction to 768P.