MiniMax H3 Full-Reference Prompt Guide: Tags & Template

By Madeleine Carter11 min read
minimax-h3-reference-prompt-guide.webp

When a video depends on one image, a short instruction may be enough. When it combines a character portrait, a location image, a motion clip, and a voice sample, the prompt needs a clearer system.

This MiniMax H3 full-reference prompt guide explains that workflow.

You will learn

  • how the main MiniMax H3 reference labels work

  • when to use the six-part format

  • how to build a reusable MiniMax H3 prompt template for complex reference-to-video tasks

The goal is simple: give every reference a stable role before asking H3 to combine them.

TL;DR: MiniMax H3 Full-Reference Prompt Workflow

A reliable full-reference prompt separates source files from the content you want to reuse.

  • A Subject is a reusable person, object, setting, action, or visual trait.

  • A Picture is a specific frame or composition anchor.

  • A Video represents source footage or timing structure,

  • An Audio label tracks sound that should be copied or used as guidance.

Define those roles first, state what should remain or change, and then describe the target video in playback order.

A reusable template helps prevent mismatched labels, missing references, and accidental conflicts.

Try MiniMax H3 Generator NOW

When Should You Use a MiniMax H3 Full-Reference Prompt?

Use the full-reference format when several assets have different responsibilities in the same generation.

Common examples include:

  • One image defines a character's face and clothing.

  • Another image supplies the location or product design.

  • A video demonstrates body movement, camera timing, or editing rhythm.

  • An audio file provides a voice, beat, line of dialogue, or sound texture.

Some details must remain unchanged while others move to a new subject or setting.

Existing footage must be edited or continued instead of merely used for inspiration.

If you are still deciding which input mode fits your project,

start with the broader πŸ‘‰ MiniMax H3 Prompt Guide.

For practical tests involving characters, motion, music, and multiple media files,

see the πŸ‘‰ MiniMax H3 Reference-to-Video Guide.

πŸ”Š A web generator may show simpler asset names such as @image1, @video1, and @audio1, so follow the labels displayed by the tool you are using. The structured format is most useful when a workflow exposes, accepts, or generates the full H3 reference syntax.

How MiniMax H3 Reference Labels Work

The four label types answer different questions.

Label

What it identifies

Typical role

<Subject N>

Reusable visible content taken from one or more references

Character, product, setting, clothing, style, pose, or action

<Picture N>

A particular reference image used as a concrete visual anchor

Opening frame, ending frame, keyframe, storyboard, or composition

<Video N>

A complete video asset or its timeline structure

Source edit, continuation, camera pattern, cuts, motion, or pacing

<Audio N>

A sound signal used by the target video

Voice, dialogue, music, rhythm, ambience, or effects

The number belongs to its own label category. <Subject 1> does not automatically mean β€œthe person in <Picture 1>,” and <Audio 1> does not have to come from <Video 1>. The definition establishes the relationship.

Subject vs Picture

The key MiniMax H3 Subject vs Picture distinction is the image's role.

If a portrait supplies a character for a new scene, define the character as a Subject and cite the portrait as its source:

<Subject 1> is the woman from <Picture 1>, retaining her face, short black hair, and red jacket.

If the image must anchor an actual frame, define that separately:

<Picture 1> sets the opening composition of [Shot 1], including framing, pose, and lighting.

One image can supply several Subjects.

Several images can also define one Subject.

A source image does not need a standalone definition unless it has its own frame or planning role.

Video and Audio

A person or action borrowed from footage still belongs to a Subject.

Use a Video label for the clip's role as source footage or its overall timing, cuts, and camera structure.

For a MiniMax H3 audio reference prompt, distinguish reusing the recording from following its voice, rhythm, or texture.

A video does not require an Audio label merely because its file contains sound.

Video and Audio labels are numbered independently. Their matching numbers do not establish a connection; their definitions do.

Understand the Six-Part Full-Reference Format

A complete MiniMax H3 full-reference format connects asset definitions to the final audiovisual timeline through six sections.

Section

The question it answers

subject_definitions

What does each Subject, Picture, Video, or Audio label mean?

summary

What kind of target video is being created, and how do the references contribute?

retention_analysis

Which traits are preserved, changed, transferred, copied, or loosely followed?

detailed_description

What happens visually and audibly from the beginning to the end?

overall_soundscape

Which environmental, physical, and non-verbal sounds fill the scene?

non_diegetic_music

What background score can the audience hear outside the scene itself?

Think of the structure as a chain:

Define the assets β†’ explain the task β†’ set preservation rules β†’ direct the timeline.

The same label must keep the same meaning along that chain. If <Subject 2> is a perfume bottle in the definitions, it cannot become the studio environment later. Every reference named at the beginning should have a clear purpose in the analysis or final description.

Copyable MiniMax H3 Full-Reference Prompt Template

Keep the framework reusable and the instructions specific.

This compact MiniMax H3 full-reference prompt template preserves all six sections. Replace the placeholders and include only the labels your task requires.

subject_definitions:

[Define each Subject, Picture, Video, and Audio label, its source, and its role.]

summary:

[task type + additional task if needed] Create [target video], using [labels] for [their roles].

retention_analysis:

[label] ([shots or role]): [relationship marker] - [what stays, changes, transfers, or is copied].

[Repeat for each separately defined item.]

detailed_description:

[One or two sentences describing the visual style.]

[Shot 1] [Opening composition, subjects, action, camera, dialogue, and synchronized sound. Name references where they take effect.]

[Shot 2] At [cut time], the camera cuts to [next view and action].

[Use additional shots only when needed.]

overall_soundscape:

[Ambient, physical, and non-verbal sounds.]

non_diegetic_music:

[Instruments, tempo, and dynamics; N/A when no background score is wanted.]

Before using it, check that every label keeps one meaning, its retention rule matches the goal, and the timeline includes the referenced content.

Three Steps to Build a Full-Reference Prompt

1. Make a Reference Map

Before writing prose, list every source and give it one primary responsibility.

Start with a short map:

  • Character image β†’ identity and clothing.

  • Studio image β†’ environment and lighting.

  • Motion video β†’ performance movement and camera timing.

  • Voice sample β†’ timbre and delivery for new dialogue.

This short map prevents two references from competing for the same role. It also reveals unnecessary assets before they make the prompt harder to control.

2. State What Should Remain or Change

The retention section should make creative boundaries explicit.

Useful relationships include:

  • fully_preserved: keep the defined identity or role intact.

  • partially_preserved: retain selected traits while allowing stated changes.

  • attribute_transfer: move an action, look, or other trait to a different subject.

  • weak_reference: borrow only a broad quality, such as atmosphere or composition.

  • fully_copy: use the complete source audio as the target track.

  • partially_copy: retain only part of the audio or mix it with new sound.

  • reference: generate new audio while following voice, rhythm, wording, or texture.

Choose the narrowest relationship that matches the result you want. If only the camera rhythm matters, do not ask H3 to preserve the entire reference video.

3. Write Events in Playback Order

The detailed description should read like a compact directing plan.

Establish the opening composition, then describe actions, reactions, camera movement, cuts, dialogue, and synchronized sound in the order the audience experiences them.

For dialogue, keep each speaker ID stable across every shot. Put environmental sounds in the soundscape section and reserve the music section for score that exists outside the characters' world.

Create with MiniMax H3

Full-Reference Prompt Example: Character, Motion and Voice

Preserve the performer, adapt the setting, and generate a new spoken line.

This illustrative example uses a performer image, a studio image, a movement video, and a voice sample. It demonstrates structure rather than a tested generation result.

subject_definitions:

<Subject 1> is the performer from <Picture 1>: the same face, copper hair, metallic blue outfit, and silver ear cuff.

<Subject 2> is the studio from <Picture 2>, with concrete walls and a circular light.

<Subject 3> is the half-turn and hand gesture from <Video 1>, transferred to <Subject 1>.

<Video 1> guides the slow clockwise camera orbit and timing.

<Audio 1> guides <Subject 1>'s low, measured voice (S1) for new speech.

summary:

[reference generation + audio reference] Create an eight-second fashion film with <Subject 1> in <Subject 2>, applying <Subject 3>, following <Video 1>'s camera timing, and referencing <Audio 1> for speech.

retention_analysis:

<Subject 1> ([Shot 1], [Shot 2]): fully_preserved - retain identity, outfit, and accessories.

<Subject 2> ([Shot 1], [Shot 2]): partially_preserved - keep materials and circular light; adapt the layout.

<Subject 3> ([Shot 1]): attribute_transfer - apply the turn and gesture to the performer.

<Video 1> (camera timing): partially_preserved - follow the orbit timing; add a final close-up cut.

<Audio 1> (S1): reference - follow vocal weight and pace without copying the recording.

detailed_description:

Minimal live-action fashion film with cool highlights.

[Shot 1] <Subject 1> stands in <Subject 2>, the circular light behind her. She performs <Subject 3> as the camera follows <Video 1>'s slow clockwise orbit. Using <Audio 1>'s low, measured delivery, she (S1) says: <d>[English] Design should move before it speaks.</d>

[Shot 2] At 00:05.000, the camera cuts to a closer three-quarter view. Her hand settles beside her collar, the light brightens slightly, and she holds the pose until the clip ends.

overall_soundscape:

Soft fabric movement, quiet footsteps, and a low electrical hum.

non_diegetic_music:

Slow electronic pulses with a sustained low synth tone, ending after the final pose.

The performer and studio are Subjects because their content is reused in a new shot. Movement is tracked separately from camera timing. The voice sample guides newly generated speech; it does not supply copied audio.

Make the Format Easier to Reuse

In a Reddit post, the creator of the πŸ‘‰ MiniMax H3 Prompt Builder and Media Loader described building templates, tag insertion, shot markers, and media previews to reduce manual formatting. It is a community tool, and its reported features address a practical need: keeping references organized.

A saved template works for occasional projects. An LLM can turn a reference map into structured text. A MiniMax H3 prompt builder may help with repeated local or ComfyUI workflows.

Whichever method you choose, check the labels and preservation rules yourself.

A tool can organize your decisions, but it cannot infer every creative priority.

Common MiniMax H3 Reference Prompt Mistakes

Should every reference image be a frame anchor?

No. Use a Subject for reusable content, such as a face, outfit, or setting.

Give a Picture its own entry when it anchors a frame, composition, or storyboard. Otherwise, cite the image inside the Subject definition.

How should I label a person from a reference video?

Define the person as a Subject. Use the Video label separately for source footage, camera timing, or editing structure.

This makes clear whether you want the person, the clip's structure, or both.

What if a label changes meaning between sections?

Give each label one consistent definition.

Keep a reference map and check it against the summary, retention analysis, and shot description. When a reference changes, update every affected section.

How do I reference a voice without copying the recording?

Use reference in the audio retention entry and describe the desired timbre, pace, and delivery.

Reserve fully_copy or partially_copy for tasks that actually reuse the audio signal.

What if I define a Subject but omit its retention rule?

Add a matching retention entry. State where the Subject appears, which traits must remain, and what may change. Naming a Subject alone does not explain how you intend to preserve or transfer it.

Why is a plot summary not enough?

A summary gives the story but may leave the shot unspecified.

Describe visible actions in playback order, including camera movement, relevant sound, and the ending. Replace vague phrases such as β€œa dramatic reveal” with what the viewer actually sees and hears.

Start with a small reference set, test the difficult relationship, and revise the part that failed.

Build the Reference Logic Before the Shot

A strong full-reference prompt begins before the first cinematic sentence. Map the files, assign stable labels, decide what must remain or change, and only then write the target video from beginning to end.

Once that logic is clear, the six-part format stops feeling like extra paperwork. It becomes a production map that tells MiniMax H3 which identity to protect, which movement to transfer, which sound to follow, and where each decision should appear in the final clip.

Ready to test the structure?

Start with the smallest set of references that can express your idea.

Try MiniMax H3 Generator Online

MiniMax H3 Full-Reference Prompt FAQ

Does every task require all six sections?

A complete official full-reference rewrite uses six sections. Simpler frame tasks and interfaces may use different input formats. Follow the requirements of your actual workflow.

Can one image define multiple Subjects?

Yes. A person, product, and setting from one image can be tracked separately. Multiple images can also contribute to a single Subject.

Do Video and Audio numbers need to match?

No. They have independent numbering. Explain a shared source when the relationship would otherwise be unclear.

Can ChatGPT generate this format?

It can help draft it. Supply the reference map, duration, required dialogue, fixed traits, and the official guide. Review the output for label consistency and instructions that conflict.