MiniMax H3 Max Motion Prompts: Seamless Video Transitions

Learn MiniMax H3 Max motion prompts for timestamps, motion spines, text animation, style locks, and seamless transitions, with real test results and copy-ready examples.

MiniMax H3 Max Motion Prompts: Seamless Video Transitions
JXP TeamSeptember 4, 202621 min read

Most AI video prompts describe what a shot should look like. Motion prompts need to describe what changes, what stays consistent, and how one visual element becomes the next.

That distinction matters on MiniMax H3 Max: one of its standout strengths is prompt following, especially in motion-heavy sequences. Across a batch of test generations, the model’s ability to track exact timing, keep text intact, and hold a single visual idea across multiple scene changes was consistently one of the more reliable parts of the output. In benchmark submissions, prompt understanding came back as its strongest category.

This guide covers seven techniques for writing MiniMax H3 Max motion prompts: timestamps, motion spine, text animation, seamless transitions, style locking, camera movement, and multi-scene continuity.

Try MiniMax H3 Max free on JXP →

Quick Answer: What Makes a Good MiniMax H3 Max Motion Prompt?

Motion Prompt Formula:

Visual Anchor + Motion Rule + Timeline + Transition Logic + Camera Behavior + Landing State + Style Lock

Short example:

A white circular logo becomes a tunnel opening. The camera pushes through the center. The circular edge remains visible throughout the transition. At 4 seconds, the tunnel flattens into a product frame. Hold the final frame perfectly still for 1.5 seconds.

A weak prompt only tells the model what appears. A working motion prompt also tells it what moves → what transforms → what remains → where it lands. That last part — the landing state — is the one most people skip, and it’s usually why a generation ends on an unstable or drifting frame instead of a clean hold.

1. Start With Prompt Expansion Settings

Prompt Expansion controls how much MiniMax H3 Max rewrites your prompt before generating. Get this wrong and it can quietly override the motion logic you wrote.

Prompt Type

Suggested Mode

Very short idea (“make a video with this text”)

Quality

Detailed, but you’re not sure it’s fully H3-compatible

Balanced

Fully structured prompt (timestamps, exact text, camera, transitions, final frame)

Disabled

When to Disable Prompt Expansion

If you’ve already specified timestamps, exact on-screen text, camera behavior, transition logic, and a landing frame, expansion has nothing useful left to add — and it can rewrite details you deliberately chose. Turn it off once your prompt is fully structured. This isn’t a hard rule: a detailed prompt can still run fine on Quality mode, it just adds a little generation time and gives the model more room to interpret rather than follow.

2. Use Timestamps to Control Motion

Basic Timestamp Structure

0.0–2.0s — Logo appears and settles
2.0–4.5s — Circle expands into a tunnel
4.5–7.0s — Camera moves through the opening
7.0–10.0s — Tunnel becomes a product frame
10.0–12.0s — Hold final composition

Don’t Use Timestamps Like a Shot List

The instinct is to write timestamps like a list of unrelated shots:

0–2: A. 2–4: B. 4–6: C.

That produces exactly what it sounds like — three separate clips stitched together. Instead, write each timestamp as a continuation of the one before it:

A becomes B. B creates C. C carries into D.

This single wording change is the difference between a prompt that reads like an edit list and one that reads like a continuous shot.

3. Build a “Motion Spine” Across the Video

A motion spine is one visual element or movement pattern that runs through every scene change in the video, so cuts read as transformations instead of cuts.

Elements that work well as a spine: a circle, a cable, a brushstroke, a line, a shadow, a color block, a camera trajectory, a smoke trail, negative space.

Weak prompt:

Cut from a record to a city and then to space.

Better motion prompt:

The circular record groove expands into the city tunnel. Its outer ring remains visible and becomes the orbit line of the next scene.

The second version gives the model one continuous shape to carry forward. The first gives it three unrelated instructions and leaves the transition entirely up to chance.

4. Make Every Transition Come From Something Already On Screen

Motion spine is the global thread running through the whole video. This is the local mechanic: every single cut should originate from an element that’s already visible, not appear out of nowhere.

Shape-to-Shape — circle → lens → tunnel → moon Line-to-Object — line → road → cable → horizon Mask Transition — text box → crop window → scene reveal Negative-Space Transition — camera moves through a letter O, a doorway, a product opening, a logo hole Material Transition — smoke → fog → water → ink Motion-Path Transition — a brushstroke → a sword edge → water → flame

Pick one of these six per transition rather than inventing a new mechanic for every cut — consistency in how you transition is what makes a video feel directed rather than randomly generated.

5. How to Prompt Text Animation Without Broken Letters

Treat Text as a Complete Graphic Layer

Define each phrase as one whole unit up front, not as individual letters to be assembled.

Move the Container, Not the Letters

Bad:

Letters explode and reform into the next sentence.

Better:

Keep the phrase intact. Move the full text block upward while its bottom edge becomes the next scene’s divider.

Asking the model to reform or morph individual letters is one common cause of garbled or unstable text in motion-typography generations.

Add a Readability Hold

Hold the phrase fully sharp and unchanged for 0.8–1.5 seconds.

Control Motion Blur

Moving state: directional blur. Readable state: zero blur. Specifying both explicitly stops the model from blurring text at the exact moment it needs to be legible.

6. Use One Dominant Motion Rule Per Scene

Pick one: clockwise rotation, diagonal sweep, vertical collapse, forward push, lateral tracking, radial expansion. Don’t stack zoom + orbit + spin + shake + morph + pan into a single scene.

Why Too Much Motion Fails

The model is simultaneously handling object motion, camera motion, typography, morphing, and style change. Stack too many of these at once and they start to conflict — motion paths cancel each other out or produce unstable, jittery results.

Better Prompt Strategy

One dominant transformation, paired with one camera behavior. Everything else in the scene should support that single motion, not compete with it.

If a Transition Still Feels Chaotic, Simplify in This Order

  1. Simplify the camera motion first

  2. Reduce the number of independently moving objects

  3. Extend the landing state (give the final composition more static hold time)

  4. Reduce the number of simultaneous transformations

  5. Simplify the overall motion spine

Add complexity back only once each motion has a clear, singular purpose in the shot.

7. Lock the Visual Style Before You Add More Motion

Style Lock Formula: Palette + Material + Rendering Method + Depth + Texture + Transition Style

Flat ink illustration, black and vermilion only, visible paper fibers, no photorealism, no 3D materials, every transition formed from the same diagonal brushstroke.

Test: Loose Style Prompt vs Style-Locked Prompt

Test A (loose style):

A short animated video in an artistic ink style, showing a scene transforming
from a mountain into a river into a bird in flight, smooth motion throughout.

Test B (style-locked):

Flat ink illustration, black and vermilion red only, visible paper-fiber
texture, no photorealism, no 3D shading, no gradients. A mountain ridge line
becomes a river's current line, becomes a bird's wing edge — the same
continuous ink stroke carries through all three. Every transition is formed
from that single brushstroke. Camera holds static, no zoom or pan.

Result: Test A rendered as a photorealistic-leaning monochrome ink-wash — smooth gradients, realistic feather and cloud detail, more “artistic photo” than flat illustration. Test B rendered exactly as specified: flat black-and-vermilion linework on a visible cream paper-fiber texture, zero gradients, zero 3D shading, from the first frame to the last. The difference wasn’t subtle — Test A and Test B look like two different art directors worked on them, from a single-line change in the prompt.

Takeaway: a generic style word (“artistic ink style”) gets interpreted loosely, while an explicit Style Lock formula gets followed closely.

8. Five MiniMax H3 Max Motion Prompt Tests

Testing methodology: all five tests ran on MiniMax H3 Max at 480p with Prompt Expansion disabled, so prompt wording stayed the primary variable rather than a settings difference.

Test 1 — Kinetic Typography

Use case: social hook / announcement teaser Motion spine: the shape of the word “READY”

Bold white sans-serif text "ARE YOU READY?" appears centered on a solid black
background and holds fully sharp for 1 second, zero motion blur. At 1.5s, all
words except "READY" fade out while "READY" remains intact as one complete
text layer. Between 1.5s and 3s, the outer edge of the word "READY" smoothly
reshapes into a thin rectangular frame — the letterforms are not scrambled or
rebuilt, the whole shape simply extends and squares off. From 3s to 4.5s, the
camera holds static while the interior of that rectangular frame brightens
and reveals a product photo filling the frame. Hold the final product frame
completely still, zero camera movement, from 4.5s to 6s. Directional motion
blur only during the 1.5–3s reshape; zero blur everywhere else.

Result: “ARE YOU READY?” and the isolated “READY” both rendered with correct spelling and held sharp as instructed. The reshape into a product frame worked — the video lands on a clean two-product still-life shot rather than a hard cut. One artifact worth flagging: the model carried “READY” into the product’s on-package label text, and at that small scale the label text came out slightly garbled.

Takeaway: the container-not-letters technique holds up for the primary display text, but if your landing frame includes secondary text (packaging, fine print), expect that smaller text to be less reliable than the hero phrase.

Test 2 — Product Ad Transition

Use case: luxury product ad (example: perfume) Motion spine: a continuous gold line

A single thin gold line is drawn across a black background from left to
right over 2 seconds. The gold line remains visible throughout the entire
video. Between 2s and 5s, the gold line becomes the curved edge of a glass
perfume bottle — the same line, now the bottle's silhouette. Camera performs
one slow lateral tracking movement, no zoom, no orbit. Between 5s and 8s,
light passing through the bottle's glass creates a rippling liquid
reflection on the surface below, and the gold line is now visible as the
ripple's outer edge. Between 8s and 10s, the ripple settles and flattens
into a thin gold underline beneath the logo text, which fades in above it.
Style: warm gold and black palette, realistic glass material with soft
specular highlights, no matte textures anywhere in the video.

Result: Landed on a warm gold-on-black monogram-style final frame — elegant, on-palette, glass/metal material read as intended with no matte textures showing up. The gold-and-black restriction held for the full 10 seconds.

Takeaway: this was the cleanest result of the five tests — a short, unambiguous motion spine (one line) paired with a tightly locked palette produced the most polished output of the batch.

Test 3 — Logo Animation

Use case: brand intro/outro sting Motion spine: a circle

A circular brand logo mark sits centered on screen for 1 second. Between 1s
and 3s, the circle's edge expands outward and the interior becomes a dark
tunnel opening, as if the camera is about to pass through it. Between 3s and
5s, the camera pushes forward through the tunnel; treat the tunnel's edge as
a closing lens iris that narrows as the camera moves through it. Between 5s
and 6.5s, the iris reopens and reveals the same circular logo mark, now in
its final static lockup position with the brand name beside it. From 6.5s
the camera fully stops — zero movement, zero drift, completely static hold
until the end at 8s. Style: match the exact line weight and color of the
original logo mark throughout, no added textures or gradients.

Result: Landed cleanly on a circular logo mark with the brand name set beside it, matching the requested final lockup layout. The explicit “camera fully stops, zero movement, zero drift” instruction appears to have done its job — the closing frame reads as a genuine static hold rather than a slow drift.

Takeaway: that explicit stop instruction is doing real work — don’t drop it thinking it’s redundant with “static hold.”

Test 4 — Fashion Motion Graphic

Use case: apparel/fashion promo Motion spine: a fabric ribbon

A single silk ribbon, deep red with a visible soft sheen, unfurls diagonally
across a black background between 0s and 2.5s. Between 2.5s and 5s, the
ribbon's shape becomes a wave on a dark water surface — same silk sheen,
same diagonal direction, camera tracking diagonally to follow it. Between 5s
and 7.5s, the wave's crest becomes a road or runway line stretching into the
distance, keeping the same diagonal camera track, no change in direction.
Between 7.5s and 10s, the road folds upward into a draped fold of fabric on
a garment, the silk sheen and red color remaining identical to the opening
frame. One dominant motion only: diagonal tracking. Lighting direction stays
constant (single key light from upper left) throughout all four stages.

Result: Closing frame is a deep red silk fold with the same soft sheen described in the opening — material and color both held through to the end.

Takeaway: naming the exact fabric (silk, with sheen) and locking the motion to one direction appears to be enough to keep a material-based spine intact across a full 10-second, four-stage sequence.

Test 5 — App / SaaS Brand Video

Use case: SaaS feature announcement Motion spine: a UI card’s border edge

A rounded UI card with a thin brand-blue border sits centered on a white
background for 1 second, camera fully static — all motion happens inside
the frame, no camera movement anywhere in this video. Between 1s and 3.5s,
the card's border edge extends outward and becomes a crop window that
reveals a product screenshot filling the frame, one clean text label reading
"Faster reports" appears bottom-left and holds fully sharp with zero motion
blur. Between 3.5s and 6s, the same blue border line becomes the vertical
axis of a bar chart, bars rising smoothly from zero, one text label reading
"3x throughput" appears and holds sharp. Between 6s and 8s, the border line
becomes the border of a final call-to-action button reading "Try it free",
which holds completely still for the last 1.5 seconds. Keep the same corner
radius and the same brand-blue color for the border element in every stage.

Result: This was the strongest text result of the whole batch — the final “Try it free” CTA button rendered with perfectly correct spelling, clean rounded-rectangle border, and the brand-blue color held from the opening card through to the closing button. Every shot in this prompt used exactly one text label at a time.

Takeaway: one-label-per-shot discipline looks like the single biggest lever for text accuracy — it’s the same lesson Test 1’s packaging-label issue points to from the other direction.

9. What Failed in Our Motion Prompt Tests

Each of these was run twice: an original prompt designed to surface the failure, then a revised prompt targeting the specific cause.

Failure 1 — Too Many Independent Objects

Original prompt:

A short video showing a sports car, then a mountain landscape, then a
smartphone screen with a UI, then a coffee cup on a table, smooth
transitions between each.

Problem to look for: with no shared anchor between these four objects, transitions likely resolve as hard cuts rather than continuous motion.

Revised prompt:

A sports car's headlight beam becomes a beam of light cutting across a
mountain ridge at dawn. That same light beam becomes the glow of a
smartphone screen powering on. The screen's rectangular glow becomes the
rectangular shape of steam rising off a coffee cup on a table. One
continuous light/glow element carries through all four scenes.

Result: The original opened on a red sports car against a snowy mountain and closed on a coffee cup — four visually independent objects exactly as written, with no obvious throughline between them. The revised version closes on steam rising off the coffee cup, echoing the “glow/light” thread the prompt asked for — a noticeably more connected feel than the original’s object-to-object jump.

Takeaway: this is the clearest evidence in the batch that motion spine is doing real work, not just theoretical — identical scene count and complexity, the only change was giving the scenes something to share.

Failure 2 — Style Changes Midway

Original prompt:

An artistic short video showing a forest turning into a city skyline at
night.

Problem to look for: with no palette, material, or rendering method specified, the visual style may drift between the opening and closing frame.

Revised prompt:

Flat gouache-painting style, muted green and navy palette only, visible
brush texture, no photorealism, no 3D rendering. A forest tree line becomes
a city skyline silhouette, both rendered in the identical flat gouache
style with the same brush texture and palette throughout.

Result: The actual failure mode here was more interesting than “style drifts partway through” — the original’s vague “artistic” descriptor was essentially ignored altogether, and the whole video rendered as fully photorealistic drone-style footage, forest and skyline both. There was no drift, because there was no artistic style applied in the first place. The revised prompt’s explicit Style Lock formula produced a genuinely flat gouache-painting look — muted palette, visible brush texture, no photorealism — held consistently from the tree line to the skyline silhouette.

Takeaway: a generic style word doesn’t cause inconsistency so much as it gets dropped entirely in favor of the model’s photorealistic default; only an explicit, itemized style spec actually shows up on screen.

Failure 3 — Text Turns Into Gibberish

Original prompt:

Video text that reads "INNOVATE" with the letters exploding apart and
reforming into the word "CREATE."

Problem to look for: asking the model to break apart and reassemble individual letters is a common cause of misspelled or merged text.

Revised prompt:

The complete word "INNOVATE" appears as one solid text layer and holds
fully sharp for 1 second. The whole text block then slides upward as one
piece while a new complete text layer, "CREATE," slides in from below to
replace it. No individual letters move independently.

Result: The original’s letter-explode approach produced a legible final frame (“CREATE” resolved correctly), but the mid-transition frames showed the word passing through a garbled, semi-legible state before settling — exactly the kind of transient breakage that looks bad if a viewer pauses or if the clip is used at a shorter duration than generated. The revised version — sliding the two complete word-layers past each other instead of exploding letters — stayed fully legible on both words at every point in the transition, with no unreadable middle state at all.

Takeaway: letter-morphing doesn’t always produce a wrong final answer, but it reliably produces an ugly, unpredictable path to get there — and the container-based approach removes that risk entirely rather than just reducing it.

Failure 4 — Transition Feels Like a Hard Cut

Original prompt:

Scene 1: a candle burning on a table. Scene 2: a spaceship flying through
space.

Problem to look for: scene 2 shares nothing visually with scene 1, so even a continuously-written prompt may still play back like an abrupt cut.

Revised prompt:

A candle flame flickers on a table. The flame's shape stretches and
becomes the glowing engine trail of a spaceship, which is now flying
through space — the same warm flame-colored glow is visible as the ship's
engine exhaust.

Result: The original opened on a candle and closed on a spaceship near a planet — two visually unrelated scenes with nothing carried between them, which reads exactly like an edit cut rather than a continuous shot. The revised version closes on a warm orange engine-glow that visibly echoes the candle’s flame color and shape from the opening — the connective tissue is there on screen, not just implied by the wording.

Takeaway: writing two scenes “continuously” in the prompt isn’t enough if nothing visual actually survives from one to the other.

10. How to Extend Motion Videos Beyond One Generation

5-Step Workflow:

  1. Generate — run the same prompt across several style/framing variations

  2. Select — pick the cleanest generation, or the one with the strongest opening

  3. Repair — if the last usable frame has an issue (garbled text, an extra object), enhance/clean that single frame with an image model before reusing it

  4. Extend — use the repaired frame as the first-frame reference for the next generation, continuing the same motion spine and style lock

  5. Stitch — when joining two clips, trim roughly 1 second off the start of the second clip. Motion in these generations tends to start slow and accelerate, so dropping into a fast-moving edit at full speed creates a visible speed mismatch at the cut point.

This is how you get past the single-generation length limit without the seam being visible.

11. Copy-Ready MiniMax H3 Max Motion Prompts

Based on the Section 8–9 results, single-element spines with a tightly locked palette — Test 2’s gold line and Test 3’s circle — were the most reliable performers, and one-label-per-shot was the biggest lever for text accuracy. The library below reflects that.

  1. Kinetic typography hook — Test 1 prompt above. Customize: the two headline phrases, brand color. Keep secondary/packaging text out of the landing frame if possible — small on-package text is less reliable than the hero phrase.

  2. Product commercial (spine transition) — Test 2 prompt above. Customize: product type, spine element (line/ribbon/cable), palette. This was the cleanest result of the batch — good default template.

  3. Logo reveal — Test 3 prompt above. Customize: logo shape, brand colors, final lockup layout. Keep the explicit “camera fully stops, zero drift” instruction — it’s load-bearing, not decorative.

  4. Luxury fashion motion graphic — Test 4 prompt above. Customize: fabric type, color, camera direction. Name the exact material and sheen; vague fabric words weaken material continuity.

  5. SaaS/app UI animation — Test 5 prompt above. Customize: feature names, metrics, CTA text. Keep to one on-screen text label per shot.

  6. Poster-to-video — a single static graphic (a poster/key art) whose one dominant shape (a headline’s baseline, a central icon) becomes the motion spine; camera stays locked off, only the spine element animates.

  7. Abstract motion graphic (spine only, no text) — take any of the six transition mechanics from Section 4 (try Material Transition: smoke → fog → water → ink) and run it with a Style Lock formula and zero on-screen text.

  8. Seamless loop — use a spine element whose start and end shape match (a circle, like Test 3’s logo mark) and write the final timestamp as “returns to the opening composition” rather than a new landing state, so the last frame can cut back to the first.

MiniMax H3 Max Motion Prompt Template

STYLE:
[visual style, material, palette]

MAIN MOTION SPINE:
[one continuous visual element that appears in every scene]

0.00–3.00:
[first composition]
[dominant transformation]
[camera behavior]

3.00–6.00:
[previous element becomes next element]
[keep visual anchor visible]

6.00–10.00:
[next transformation]
[maintain same material and style]

10.00–15.00:
[final transformation]
[stop camera]
[hold final frame]

TEXT RULES:
[exact text]
[keep phrases as complete layers]
[no letter scrambling]

STYLE LOCK:
[same colors/materials/rendering throughout]

FAQ

Can MiniMax H3 Max follow timestamps?

Yes. MiniMax H3 Max can respond well to timestamped motion instructions, especially when each passage continues logically from the previous one (“A becomes B”) rather than reading like a list of separate shots.

How do you make seamless transitions in MiniMax H3 Max?

Give every cut a shared visual anchor already on screen — a shape, line, or material that carries from one scene into the next — instead of describing each scene independently.

How do you keep text readable in H3 Max videos?

Define each phrase as one complete text layer, move the container rather than individual letters, and add an explicit readability hold (0.8–1.5s of zero motion blur) once the phrase is fully formed.

What is a motion spine in an AI video prompt?

A single visual element or movement pattern — a circle, a cable, a line, a material — that you deliberately carry through every scene change, so the video reads as one continuous transformation instead of a sequence of cuts.

How do you keep the same style across multiple H3 Max scenes?

Lock the palette, material, rendering method, depth, and texture explicitly at the start of the prompt, rather than relying on a single generic style word like “artistic” or “cinematic.” A vague style word doesn’t cause the style to drift — it gets dropped entirely, with the model defaulting to a photorealistic render. An explicit, itemized Style Lock is what actually shows up on screen.

Does Prompt Expansion change my motion timing?

It can. Expansion rewrites parts of your prompt before generation, which can alter or override timestamps and transition logic you specified. If your prompt is already fully structured, disable it.

How long can a single MiniMax H3 Max generation be?

Individual generations are capped at a fixed duration; to go longer, generate multiple clips using the same motion spine and style lock, then stitch them using the workflow in Section 10.

Why does text sometimes break apart in H3 Max motion videos?

One common cause is asking the model to animate individual letters — exploding, morphing, or reassembling them — instead of treating the whole phrase as one graphic layer that moves as a unit.

Final Thoughts

Better motion prompting is less about adding more description and more about defining continuity — what stays the same, what’s allowed to change, and where the camera is supposed to land.

Try MiniMax H3 Max free on JXP →