MiniMax H3 Max AI Video Generator

Create short drama shots from text, a first-frame image or reference assets with MiniMax H3 Max on Aniv. Generate 5–15 second videos in 480P, 768P or 1080P with native audio.

Text, keyframes, references and sound — one video workflow

Describe the video scene you want to generate...
Native audio included

What is MiniMax H3 Max?

H3 is MiniMax's video model. H3 Max is a post-trained variant tuned for stronger prompt adherence, audiovisual quality and aesthetics. It creates synchronized audio with the picture and follows a written shot more closely than stock H3 at the same resolution, which is why it is the default here.

H3 Max at a glance

Length
5 to 15 seconds
Resolution
480P, 768P or 1080P
Inputs
Text prompt, an image with an optional end frame, or up to 12 reference images, videos and audio clips
Aspect ratios
21:9, 16:9, 4:3, 1:1, 3:4, 9:16

What you can create with H3 Max

Direct the shot, keep the look consistent, and bring motion and sound together in one generation.

Camera & Motion

Direct every camera move

Describe a tracking shot, pan, push-in, or fast follow. H3 Max keeps camera and subject movement aligned with the shot you wrote.

First & Last Frame

Turn two frames into one journey

Start from an opening image and add an optional closing frame. H3 Max builds continuous motion between the two moments.

Character Consistency

Keep characters recognizable across shots

Faces, clothing, proportions, and identity can stay coherent as the setting, light, and camera angle change.

Native Audio

Generate sound with the picture

Create ambience, effects, dialogue, and music in the same pass, timed to what happens on screen.

Art Direction

Hold one visual language

Carry a chosen palette, texture, linework, or cinematic treatment across every beat of the sequence.

Prompt Adherence

Render the beats in your order

Describe the action step by step. H3 Max can follow the sequence and preserve requested visual details, including on-screen text when applicable.

A new singularity in AI video

H3 Max turns ideas into motion in seconds. Model inference for a 5-second 768P video completes in about 3 seconds, with queue and file-processing time varying.

Core controls, clearly exposed

Duration, resolution, aspect ratio, seed and end frame — the same fields the model takes.

Credits follow each render

See the exact cost before generating. Use a subscription or top up when you need more; failed renders return their credits.

Prompts in English or Chinese

H3 reads both, so you can write in either language.

H3 Max or H3?

H3 Max prioritizes prompt adherence, audiovisual quality, aesthetics and fast iteration. Standard H3 raises the output ceiling to 2K and 4K. Aniv currently offers H3 Max only.

H3 Max or H3?
Model
H3 MaxDefault here
H3
Best forPrompt fidelity and rapid iteration2K and 4K delivery
Maximum output1080P4K
InputsText or image, with an optional end frameText or image, with an optional end frame
Synchronized audioIncludedIncluded

FAQ

FAQ

Have another question? Contact us at support@aniv.ai

H3 Max is fal’s post-trained variant of MiniMax H3, focused on prompt adherence and visual aesthetics.

480P costs 8 credits per second, 768P costs 12 and 1080P costs 32. Reference mode doubles the rate. A 5-second 768P video costs 60 credits, with native audio included.

Get started

Generate your video with H3 Max now

From a text prompt or an opening frame — a 5 to 15 second cut with synchronized audio, in a single pass.