On This Page8 sections
Key Takeaways
- MiniMax H3 Max Ref is shorthand for MiniMax H3 Max Reference-to-Video, rather than a separate model officially named "Ref."
- The production endpoint is
minimax/h3-max/reference-to-video. - H3 Max is a post-trained version of MiniMax H3 developed by fal Research, with additional optimization for prompt adherence, visual quality, and inference speed.
- Reference-to-Video can use images, videos, and audio as references, allowing creators to preserve characters, products, motion, style, or voice while generating an entirely new scene.
- The current workflow supports up to 12 reference files, 5–15 second output, 480P, 768P, and 1080P options, plus start, middle, and end-frame controls.
- H3 Max Ref is especially useful for AI anime, recurring characters, UGC advertising, product commercials, visual novels, and multi-shot storytelling.
- Its biggest difference from ordinary image-to-video is that a reference image does not need to become the opening frame. It can act purely as an identity or appearance anchor.
What Is MiniMax H3 Max Ref?
MiniMax H3 Max Ref generally means the Reference-to-Video mode of MiniMax H3 Max.
Its production endpoint is:
minimax/h3-max/reference-to-video
The word Ref is simply shorthand for reference.
MiniMax developed the original H3 model. H3 Max is a post-trained H3 variant developed by fal Research, with additional optimization focused on prompt adherence, aesthetics, and high-throughput inference.
This distinction matters because users searching for minimax h3 max ref are usually not looking for another completely separate model. They are looking for the reference-conditioned generation mode capable of taking images, videos, and audio as guidance.
Why Reference-to-Video Is Different From Image-to-Video
Traditional image-to-video normally takes an uploaded image and treats it as the visual starting point of the generated clip.
Reference-to-Video solves a different problem.
A reference image can define who a character is without forcing that image to become frame one. A reference video can demonstrate motion or performance. A reference audio clip can provide voice, rhythm, or sound context.
For example, a creator could provide a studio portrait of a character as Image 1, then ask the model to generate that same person walking through a rainy Tokyo street at night.
The reference controls identity. The prompt controls the new scene.
That separation is what makes Ref-to-Video especially valuable for workflows requiring consistency across multiple independently generated shots.
What Can Be Used as a Reference?
H3 Max Ref supports three main types of reference media.
Reference Images
Reference images can help preserve:
- Character identity
- Facial appearance
- Hairstyle
- Clothing
- Product design
- Props
- Illustration style
- Color language
- Creature design
Prompts can explicitly assign roles to assets such as Image 1, Image 2, and Image 3.
For example:
Image 1 is the protagonist.
Image 2 is the product.
Preserve the protagonist's face, hairstyle, and clothing from Image 1.
Preserve the product's shape, materials, and logo placement from Image 2.Explicit mapping reduces ambiguity and makes it easier for the model to understand which properties should remain consistent.
Reference Videos
Reference video can provide temporal information such as:
- Body movement
- Dance choreography
- Character performance
- Camera motion
- Gesture timing
- Physical interactions
- Action pacing
A reference video should still be considered generative guidance rather than deterministic motion transfer. The model can interpret the movement instead of copying every frame exactly.
Reference Audio
Reference audio can help guide:
- Voice characteristics
- Dialogue delivery
- Performance rhythm
- Sound cues
- Musical direction
- Ambient sound
This is particularly useful because the H3 family can generate synchronized video and audio rather than requiring an entirely separate audio-generation pipeline.
The 12-Reference Limit
H3 Max Ref currently supports up to 12 reference files in total across images, videos, and audio.
For example, one request could contain:
- 6 images
- 2 videos
- 2 audio clips
That would use 10 references.
More references are not automatically better.
Every additional input introduces another source of information the model must reconcile. A smaller set of high-quality, non-conflicting references is usually easier to control than uploading 12 loosely related assets.
A practical production setup might use:
- 2–3 images for character identity
- 1 image for wardrobe
- 1 image for product identity
- 1 reference video for motion
- 1 audio file for voice or rhythm
Start, Middle, and End Frames
One of the most useful H3 Max Ref capabilities is the ability to combine references with explicit keyframe guidance.
The workflow can support:
- Start frame
- Middle frame
- Specified middle-frame timing
- End frame
This creates two independent layers of control.
References define what should remain recognizable.
Keyframes define where the generated shot should begin, pass through, or end.
For example, a product advertisement could use:
Image 1: the exact shoe design- Start frame: the shoe sitting on a pedestal
- Middle frame: an athlete picking up the shoe
- End frame: the athlete running while wearing it
The reference preserves the product identity, while the keyframes constrain the progression of the shot.
This makes H3 Max Ref closer to directed video synthesis than traditional image animation.
MiniMax H3 Max Ref Specifications
| Feature | Capability |
|---|---|
| Endpoint | minimax/h3-max/reference-to-video |
| Output duration | 5–15 seconds |
| Resolutions | 480P, 768P, 1080P |
| Reference types | Images, video, audio |
| Maximum references | 12 combined |
| Start frame | Supported |
| Middle frame | Supported under required conditions |
| End frame | Supported |
| Aspect ratios | Adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| Seed | Supported |
| Prompt expansion | Disabled, balanced, quality |
| Safety checker | Supported |
The current 1080P workflow uses a refinement path rather than treating 1080P as identical to native high-resolution generation. This distinction is important when comparing H3 Max Ref with competing models.
Why H3 Max Ref Is Strong for Character Consistency
Character consistency remains one of the hardest problems in generative video.
A prompt such as young woman with black hair wearing a red jacket describes a category of person, not one exact identity.
Generate ten independent shots from that description and the model may alter the face, hairstyle, body proportions, or wardrobe between clips.
Reference images provide a stronger identity anchor.
A useful character pack might include:
Image 1: front-facing portraitImage 2: three-quarter facial viewImage 3: full-body wardrobe reference
A structured prompt could be:
Images 1, 2, and 3 show the same protagonist.
Preserve her facial identity, black bob haircut, silver glasses, and dark red jacket.
She walks into a neon-lit convenience store at night.
The camera tracks backward in front of her.
She looks toward the shelves, then briefly toward camera.
Audio: natural footsteps, refrigerator hum, and distant rain.This separates four different concerns:
- Identity — controlled by references.
- Action — walking and looking.
- Camera — backward tracking.
- Audio — footsteps, refrigeration, and rain.
That structure is usually more useful than filling a prompt with vague adjectives such as cinematic, beautiful, or masterpiece.
Why H3 Max Ref Is Particularly Interesting for Anime
Reference-conditioned generation is especially useful for anime and illustrated characters because small identity changes are visually obvious.
Anime characters are commonly defined by a precise combination of:
- Hair silhouette
- Eye design
- Costume
- Accessories
- Color palette
- Facial proportions
- Illustration style
Pure text-to-video can easily drift across different scenes.
A more reliable anime workflow uses a reusable reference pack containing:
- Character sheet
- Face close-up
- Full-body image
- Costume reference
- Optional environment or style reference
Those assets can then be reused while changing the action, environment, or camera in each generated shot.
Potential use cases include:
- AI anime shorts
- Fan-animation concepts
- Visual-novel adaptations
- Character trailers
- Music-video sequences
- Episodic social content
Reference consistency does not mean deterministic animation. Hands, facial expressions, accessories, and complex motion can still drift and may require multiple generations.
Product Advertising Is Another Strong Use Case
Product advertising requires more than producing a visually attractive shot. The product itself must remain recognizable.
A generated shoe, bottle, phone, watch, or car becomes unusable if key design details change during the clip.
H3 Max Ref can keep the product as a dedicated reference while generating a completely new composition.
For example:
Image 1 defines the exact espresso machine.
Preserve its black metal body, chrome lever, proportions, logo placement, and control layout.
Create a cinematic close-up in a warm modern kitchen.
The camera slowly pushes toward the machine as espresso pours into a cup.
Steam catches the morning window light.
Audio: realistic machine sounds and quiet room ambience.This workflow is useful for:
- E-commerce advertising
- UGC campaigns
- Product launch videos
- Branded social content
- Concept commercials
- Rapid creative testing
A Better H3 Max Ref Prompt Structure
Reference-video prompts are easier to control when written in layers.
1. Map the References
Explain what each input represents.
Image 1 is the protagonist. Image 2 is the product. Video 1 is the motion reference.
2. Define What Must Stay Consistent
State the attributes that matter most.
Preserve the protagonist's face and clothing. Preserve the product's proportions, materials, and markings.
3. Describe the Scene
Specify location, time, lighting, weather, and environment.
4. Describe the Action
Explain events in chronological order.
5. Describe the Camera
Specify framing, movement, and shot style.
6. Describe the Audio
Define dialogue, ambience, music, or sound effects where relevant.
A reusable template is:
Image 1 is [subject A].
Image 2 is [subject B].
Video 1 is the motion reference.
Preserve [identity-critical characteristics].
Scene: [location, time, lighting, environment].
Action: [chronological action].
Camera: [shot size, lens feel, camera movement].
Audio: [dialogue, ambience, sound effects].
Avoid changing [critical character, product, or style details].Common Prompting Mistakes
Uploading References Without Explaining Them
The model can infer relationships between assets, but production prompts should remove unnecessary ambiguity.
Use explicit statements such as:
Image 1 and Image 2 show the same woman.
or:
Image 1 defines the character. Image 2 is only a lighting reference.
Giving References Conflicting Roles
If one image shows a realistic person and another shows an anime interpretation of the same person, the model must decide which characteristics are authoritative.
State the intended hierarchy clearly.
Using Too Many References
The maximum reference count is a limit, not a target.
A small group of clean references usually provides stronger control than a large set of inconsistent images.
Expecting Exact Motion Transfer
Reference video provides guidance rather than guaranteed frame-perfect pose transfer.
If exact choreography or pose timing is critical, a specialized motion-transfer workflow may be more appropriate.
Using Vague Aesthetic Adjectives
Words such as beautiful, cinematic, stunning, and masterpiece generally provide less control than concrete instructions describing:
- Light direction
- Shot size
- Camera movement
- Character action
- Environment
- Timing
Prompt Expansion Modes
H3 Max Ref exposes several prompt-expansion modes:
disabledbalancedquality
balanced is a practical default for general use.
quality can perform more extensive prompt expansion but adds preprocessing time.
disabled is useful when a carefully engineered prompt should reach the generation system with minimal rewriting.
For controlled model comparisons, disabling prompt expansion can also make benchmarking cleaner because less invisible rewriting occurs between the submitted prompt and generated result.
MiniMax H3 Max Ref API Example
A basic JavaScript integration can look like this:
import { fal } from "@fal-ai/client";
const result = await fal.subscribe("minimax/h3-max/reference-to-video", {
input: {
prompt: "Image 1 is the protagonist. Preserve her facial identity and clothing. She walks through a rainy Tokyo street at night while the camera tracks alongside her.",
reference_image_urls: [
"https://example.com/character-reference.jpg"
],
duration: 5,
resolution: "768P",
aspect_ratio: "16:9",
prompt_expansion_mode: "balanced",
enable_safety_checker: true
},
logs: true
});
console.log(result.data.video.url);API credentials should be stored on the server rather than exposed in client-side browser code.
For production workloads, queued jobs or webhooks are generally more robust than keeping an HTTP request open until a video finishes generating.
H3 Max Ref Pricing
Pricing can change, so current API pricing should always be checked before deploying a production workload.
At a commonly listed 768P H3 Max rate of $0.08 per generated second, approximate output cost is:
| Duration | Approximate 768P output cost |
|---|---|
| 5 seconds | $0.40 |
| 10 seconds | $0.80 |
| 15 seconds | $1.20 |
Reference processing can introduce additional cost depending on reference type and provider implementation.
Reference video is especially important to account for because processing several seconds of video can cost more than using a small collection of still images.
For high-volume applications, the better business metric is therefore:
usable outputs per dollar
rather than simply:
advertised cost per render
A slightly more expensive request can still be more economical if it produces a substantially higher percentage of usable videos.
H3 Max Ref vs Traditional Image-to-Video
Choose image-to-video when:
- The uploaded image should become the first frame.
- Only one composition needs to be animated.
- The workflow should remain simple.
- Identity across unrelated scenes is not important.
Choose H3 Max Ref when:
- A character must appear in a different composition.
- Multiple characters need recognizable identities.
- Product identity must survive scene changes.
- Motion references are useful.
- Audio references matter.
- Several shots should reuse the same character or object.
- References need to be combined with start, middle, or end-frame guidance.
The simplest distinction is:
Image-to-video animates an image. Reference-to-video uses media as instructions for creating a new video.
H3 Max Ref vs MiniMax H3
MiniMax H3 and H3 Max belong to the same underlying model family, but they should not be treated as identical products.
| Capability | MiniMax H3 | H3 Max Ref |
|---|---|---|
| Origin | MiniMax | fal Research post-training based on H3 |
| Primary role | General H3 video generation | Optimized hosted reference generation |
| Reference conditioning | Yes | Yes |
| Native audio | Yes | Yes |
| Main optimization | Broad multimodal generation | Prompt adherence, aesthetics, speed |
| Best fit | Flexible H3 workflows | Fast reference-driven production |
H3 Max is therefore better understood as an optimized production variant than a simple rename of H3.
Where H3 Max Ref Still Has Limitations
Fine-Detail Drift
Hands, jewelry, small logos, text, and tiny facial details can still change during motion.
Multi-Character Confusion
When multiple visually similar characters appear together, attributes can occasionally be swapped or blended.
Explicit reference mapping helps, but cannot guarantee perfect separation.
Motion Is Not Deterministic
Reference video does not guarantee frame-perfect reproduction.
Fast motion, complicated interactions, and exact choreography may require several generations.
Conflicting References
Style references and identity references can pull the output in different directions.
Every input should have a clearly defined role.
Short Output Duration
The workflow is built around short clips, so longer stories still require:
- Shot planning
- Multiple generations
- Continuity management
- External editing
1080P Uses a Refinement Workflow
The current 1080P option should not automatically be interpreted as identical to native 1080P generation.
This matters when comparing technical specifications across AI video models.
Best H3 Max Ref Workflows
AI Anime and Character Shorts
Create a reusable identity pack and use the same core references across every shot.
AI UGC Advertising
Combine a creator reference with a product reference, then generate different hooks, backgrounds, actions, and camera setups.
Product Commercials
Use the product as a persistent identity anchor while independently changing environment, cinematography, and lighting.
Visual Novels
Turn recurring character illustrations into dialogue shots, reaction shots, and animated environments.
Social Video Iteration
Generate multiple short candidates quickly, then select the strongest outputs for editing.
A Practical Multi-Shot Consistency Strategy
For a sequence containing several shots, avoid changing every variable at once.
A stronger workflow is:
- Create a canonical reference pack. Reuse the same identity assets across every shot.
- Lock wardrobe and defining features. Repeat critical attributes in each prompt.
- Keep style references stable. Avoid contradictory visual references.
- Generate short shots. Five-second shots are easier to control than one long generation containing several story beats.
- Use seeds selectively. Seeds can help with nearby variations but do not guarantee identical characters across fundamentally different scenes.
- Use keyframes where composition matters. Start, middle, and end frames can be more effective than increasingly verbose text instructions.
- Edit externally. H3 Max Ref is a shot-generation system, not a replacement for continuity editing.
Is H3 Max Ref Better Than Text-to-Video?
They solve different problems.
Text-to-video is ideal during early ideation when no fixed identity needs to be preserved.
Reference-to-Video becomes much more useful after the creative direction is established and specific characters, products, styles, or performances must remain recognizable.
A practical workflow is:
Text-to-video → choose concept → build reference assets → generate controlled Ref-to-Video shots → edit final sequence
This keeps early experimentation fast while adding stronger control during production.
Why the Search Term "MiniMax H3 Max Ref" Matters
The formal term is Reference-to-Video, but creators increasingly shorten it to H3 Max Ref.
Users commonly search for the exact wording they encounter in:
- Social posts
- AI video showcases
- Creator workflows
- Tool interfaces
- Community discussions
rather than searching for a formal API endpoint.
That means the search intent behind minimax h3 max ref includes much more than a simple definition.
Typical questions include:
- What is H3 Max Ref?
- Is H3 Max Ref different from H3 Max?
- How many reference images can it use?
- Can it use reference video?
- Can it preserve anime characters?
- Does it support audio?
- How much does it cost?
- What prompts work best?
- Is there an API?
This makes the topic a strong candidate for a comprehensive evergreen guide rather than a short product-news article.
Frequently Asked Questions
Is MiniMax H3 Max Ref a separate model?
Not in the usual sense. The phrase generally refers to the H3 Max Reference-to-Video mode rather than a separate official model named "Ref."
What does Ref mean?
Ref means reference.
Can H3 Max Ref use multiple images?
Yes. Multiple images can be combined with video and audio references, subject to the endpoint's combined reference limit.
Can it use reference video?
Yes. Reference video can provide motion and temporal guidance.
Can it use audio references?
Yes. Audio can be included as part of the multimodal reference workflow.
Does H3 Max Ref generate audio?
H3 Max retains the H3 family's synchronized video-and-audio generation capabilities.
How long can H3 Max Ref videos be?
The current workflow is designed for short-form clips in the 5–15 second range.
Is H3 Max Ref good for anime?
Yes. Anime is one of its strongest use cases because reference images help anchor character identity, costume, palette, and visual style across different shots.
Is H3 Max Ref good for product videos?
Yes. A product can remain a dedicated reference while the model generates new environments, camera angles, and actions.
Is character consistency perfect?
No. Reference conditioning improves consistency but does not make generation deterministic. Important shots can still require several attempts.
Conclusion
MiniMax H3 Max Ref represents an important shift from generating isolated AI clips toward generating controlled shots around persistent characters, products, motion, and sound references.
Its strongest advantage is the combination of multimodal references, short-form video generation, native audio, explicit reference mapping, keyframe control, and fast iteration.
That makes it particularly useful for:
- AI anime
- Recurring AI characters
- UGC advertising
- Product commercials
- Visual novels
- Multi-shot storytelling
- Social video experimentation
The most effective approach is to build a small, high-quality reference pack, give every reference a clearly defined role, keep prompts structurally simple, generate short controlled shots, and use keyframes only when composition must land at specific moments.
H3 Max Ref does not eliminate character drift, motion errors, or continuity problems. But for creators who need substantially more control than basic text-to-video or image-to-video, it is one of the most practical reference-driven AI video workflows currently available.
Continue Reading
More articles connected to the same themes, protocols, and tools.
Referenced Tools
Browse entries that are adjacent to the topics covered in this article.








