MiniMax H3 Reference-to-Video: How to Assign Every Image, Video, and Audio File

MiniMax H3 Reference-to-Video

The most useful MiniMax H3 reference pack gives every file one written job. An image may define a face, another a coat, a video may supply movement, and an audio file may guide a voice. If two references compete to control the same detail, a longer prompt rarely resolves the conflict.

Start with a role sheet, then choose the matching H3 workflow. A first frame, identity image, motion video, and voice reference are not interchangeable inputs.

Pick the workflow before assembling the reference pack

MiniMax H3 is exposed through three broad routes in ClipDance: text-to-video, image-to-video, and reference-to-video. Selecting the route first prevents a common mistake—building a mixed reference pack and then trying to force it through a form intended for fixed frames.

Text-to-video suits a scene in which no existing asset must determine the opening composition or subject identity.

Image-to-video requires a first frame in the current browser form and can optionally use a last frame. The first establishes composition and orientation; the last gives the shot a destination. These are timeline roles, not a general character-and-style collection.

Reference-to-video allows a flexible mix of images, video, and audio without locking the output to one opening frame. The current form accepts up to nine images, three videos with a combined duration limit, and three audio files. Audio cannot be used alone; at least one image or video is required. The prompt still has to assign each asset’s responsibility.

A first frame is not an identity reference

A first frame answers, “What does the opening instant look like?” It controls composition, visible subjects, pose, lighting, and initial state. That is not the same as supplying views that define a person throughout a changing shot.

An identity reference asks which features must remain recognizable as pose, angle, or expression changes. A face view may define facial structure while a full-body image defines build and wardrobe.

The optional last frame gives image-to-video a destination. If it changes clothing, background geometry, or camera angle abruptly, the model must invent a transition between incompatible endpoints.

Use fixed frames when the endpoints matter. Use general references when identity, motion, sound, or style matters more than an exact opening composition.

Give images one visual responsibility each

Create a manifest before uploading. The labels can remain in production notes; what matters is that their order matches the live form and that the prompt describes the same roles in plain language.

File Primary role Details that must stay fixed Details allowed to change
Image 1 Character identity Face shape, hair, age cues Expression, gaze, head angle
Image 2 Wardrobe Coat cut, fabric, buttons, color Folds caused by movement
Image 3 Location Window layout, wall material, floor Camera position within the room
Image 4 Visual treatment Contrast, palette, grain character Exact objects and composition

The table exposes contradictions early. If Image 1 shows short hair and Image 2 shows long hair on the same person, choose one. If two locations disagree about the windows, do not ask both to control the set.

Avoid calling every image a “style and character reference.” That gives the model no priority when features conflict. A narrow role is easier to diagnose.

Pair an affirmative role with a boundary when the source contains distracting material. A wardrobe image may show the correct coat on the wrong person; tell H3 to preserve the garment construction without copying that person’s face. A location reference may include pedestrians who should not enter the new scene. Naming the exclusion prevents the unwanted detail from gaining authority merely because it is visible.

Boundaries should stay concrete. “Do not change anything else” is difficult to inspect when several references already disagree. “Keep the lead character from Image 1; copy only the coat from Image 2; do not introduce the model or background from Image 2” creates a decision that can be checked in the result.

Use video for time-based information

A reference video can show body mechanics, gesture timing, camera speed, interaction, and rhythm. Decide which properties matter. Borrowing movement does not require copying the performer, clothing, setting, or framing.

The first video might control walk timing while images control the character and coat. Another video might supply a camera arc but not lighting or location. This prevents motion references from redesigning the scene.

Short clips are easier to assign than a montage. If the source changes angle, performer, and tempo, H3 must infer which portion is authoritative. Trim it when only one movement matters.

Because the form limits combined reference-video duration, selection matters. One clear gesture and camera move are more controllable than a reel of unrelated motion.

Treat voice and ambience as separate audio jobs

Audio can guide dialogue, voice character, music, effects, or ambience. If a track combines speech, music, and room noise, state which layer matters and which may change.

For dialogue, record speaker, language, exact wording, pace, and emotion. If the words must remain unchanged, compare the output with the approved source and restore the master during editing when necessary. A reference is not automatic bit-for-bit preservation.

For ambience, describe continuity: a station hum remains, rain stays outside, or a machine stops on a beat. If music only guides pacing, say so; otherwise it may compete with dialogue.

H3 reference-to-video does not accept an audio-only job in the current form. Pair sound with a visual reference and explain their relationship.

Write the prompt as a set of responsibilities

After the manifest is clean, convert it into ordinary production language. Do not invent a private ClipDance command syntax or assume that labels from a local ComfyUI graph map directly to a hosted form.

A compact responsibility block can read:

Keep the woman from the first reference image as the only lead character.

Preserve the coat construction and dark green fabric from the second image.

Use the first reference video only for the pace and direction of her turn.

Use the first audio reference for the calm female voice; keep the station ambience quiet.

The camera makes one slow push in. Do not copy the people or setting from the motion video.

Each sentence identifies an asset, job, and boundary. The last prevents the motion source from taking over identity or location.

Then add a timeline: opening action, fixed details, dialogue start, and ending state. Roles define the ingredients; timing arranges them.

Do not copy instructions blindly from an open workflow

MiniMax publishes open first/last-frame and reference-to-video workflows, but a local graph and hosted form are different interfaces. Node names, file slots, placeholders, and controls may differ.

Within ClipDance, MiniMax H3 and the seedance 2.5 workspace are separate tool choices in a third-party browser catalog. Their presence there does not make ClipDance an official property of MiniMax or ByteDance. Instructions should match the controls visible in the selected ClipDance mode.

For some listings, reAPI sits upstream of the ClipDance form and handles asynchronous request and status flow across multiple model APIs. The aggregator did not train H3; the form and provider are distinct layers. A reAPI field name therefore should not be treated as a universal H3 prompting rule.

Resolve reference conflicts before the first generation

Could another editor identify the file controlling identity, find one wardrobe source, and spot motion or audio that conflicts with the prompt? If not, the pack is not ready.

Remove redundant files and replace references that hide the detail they control. If two sources share one property, set a priority: face from Image 1, hair color from Image 2, all other facial details from Image 1.

After generation, diagnose by role. A changed face points to identity; incorrect motion to the video and timeline; a changed voice to the audio brief. Do not replace the entire pack after every imperfect result.

Before adding one more file, ask what unique decision it makes. If the answer is unclear, leave it out. MiniMax H3 reference-to-video becomes easier to direct when every upload has a name, one responsibility, and a stated limit on what it is allowed to change

Leave a Comment

Your email address will not be published. Required fields are marked *