How MiniMax H3 reference-to-video works
Reference-to-video uses your photos to guide who appears in a newly generated scene. In this studio, Performer 1 is the left reference and Performer 2 is the right reference. MiniMax H3 creates an orange recording booth, a shared hanging microphone and the selected performance vibe around the pair.
Both photos are passed as reference images. They are not the first and last frames of a transition. Use one visible adult per photo, clear faces and even lighting. JPG, PNG and WebP are supported, up to 10 MB each and at least 256 × 256 pixels. The studio prepares the references for H3 without cropping away the person.
MiniMax H3 and H3 Max ref: which model does this page use?
MiniMax H3 Max ref is shorthand for H3 Max reference-to-video. H3 and H3 Max are separate variants with different supported settings. This page's generator uses H3 for every quality tier.
MiniMax H3 supports native 768P or 2K output and requests from 4 to 15 seconds. The studio keeps the familiar Preview, Standard and HD export options: Preview and Standard generate at native 768P, then export at 480p or 720p; HD generates at native 2K, then exports at 1080p. Vertical 9:16 and landscape 16:9 are available.
Make a MiniMax H3 video in three steps
1. Choose your take. Start with a 10-second Standard video at 720p for 120 credits, covered by the $5 pack before applicable taxes. You can also choose 4 or 15 seconds, Preview or HD, vertical or landscape, and a playful duet, freestyle or relaxed session. The quote updates before you generate.
2. Sign in with Google, buy enough credits and upload the two photos. Confirm that both people are adults and that you have permission for AI transformation and your intended use. This page uses the same account, credits, content checks and private library as the homepage.
3. Generate, review and download. Your task stays linked to your account while H3 runs. Return with the same account to check progress. When it completes, review both faces, movement and audio, then download the MP4. Confirmed failed tasks return reserved credits; an uncertain submission needs a support check before retrying.
What the reference inputs control
The H3 API supports up to nine reference images, three reference videos and three reference audio files. Audio references must accompany an image or video reference. First/last frame inputs and multimodal reference inputs are separate modes and cannot be mixed in one request.
This focused studio accepts exactly two photos and supplies the scene and performance prompt. Reference-video uploads, uploaded audio, arbitrary text prompts and additional reference images are not controls in this release. Your two photos guide the performers; the selected vibe guides newly generated movement.
Avoid identity drift and conflicting references
A clear face in each photo gives the model more useful information than a distant subject, heavy filter, sunglasses or a crowd. Keep distinguishing facial details and clothing visible. Review the left/right assignments before starting; likeness and exact positions can still vary in the generated result.
H3 generates fresh motion and audio. It does not supply the original Hotel Lobby recording, a licensed song, an exact dance or a cloned artist voice. If you add a soundtrack after downloading, use music you are allowed to use. Watch the whole clip before sharing it and follow your destination platform's AI disclosure rules.

