LTX 2.5 vs MiniMax H3: Which Video Model Should You Use?

Same price per second, very different models — and one licence clause that can shape the choice before you compare a single frame.

LTX 2.5 vs MiniMax H3: Which Video Model Should You Use?
JXP TeamAugust 13, 202615 min read

The LTX 2.5 vs MiniMax H3 comparison is not the resolution shootout most roundups make it look like. Both shipped open weights within a fortnight of each other, both generate synchronized audio, and at the settings most teams use, both cost the same per second. The decision turns on three things instead: what each model lets you control, how much of the system is genuinely downloadable, and — for teams in the US, EU, UK, or South Korea — a clause in H3’s licence that creates a major constraint on self-hosting. That last point is worth checking before you compare a single frame.

LTX 2.5 vs MiniMax H3: Quick Verdict

Choose LTX 2.5 for output above 2K, clips longer than 15 seconds, multi-shot sequences from one generation, or a more straightforward self-hosting path in regions excluded by H3’s default Community License.

Choose MiniMax H3 for reference-driven work — matching a face, product, motion, or voice from material you already have — for delivery beyond 16:9 and 9:16, and for editing where audio has to stay coherent.

Head-to-Head

LTX 2.5

MiniMax H3

Downloadable weights

Yes

Yes

Complete official local pipeline

More complete

Partial

Parameters

22B

33.1B

Max resolution

4K (Fast)

2K = 1440p short edge

Max duration

20 sec (Fast)

15 sec

Min duration

6 sec

4 sec

Frame rate

24/25/48/50

24

Native audio

Yes

Stereo

Aspect ratios

16:9, 9:16

21:9, 16:9, 4:3, 1:1, 3:4, 9:16 + adaptive

Multi-shot in one generation

Yes

Not documented

First / last frame control

Last frame only

First, last, or both

Reference images

Not documented

Up to 9

Reference clips (video / audio)

Up to 3 each

Video editing

Not on the current LTX 2.5 API

Yes

Excluded-territory clause

Not identified

US, EU, UK, Korea by default

Community licence revenue threshold

Under $10M

Under roughly $20M

Try LTX 2.5 in JXP

Try MiniMax H3 in JXP

What Each Model Actually Is

LTX 2.5

A 22-billion-parameter asymmetric dual-stream diffusion transformer from LTX, the company spun out of Lightricks. Two API variants: Fast, reaching 4K and 20-second clips, and Pro, capped at 1080p and 10 seconds but tuned for fidelity.

Its defining feature is native multi-shot — one generation producing several connected shots that hold character identity, environment, lighting, and voice across cuts.

One gap worth naming: retake, extend, reframe, and HDR upscale remain listed and priced in the LTX API, but against ltx-2-3-pro rather than LTX 2.5. Those editing workflows exist in the ecosystem; they have not migrated to the new model.

MiniMax H3

The model behind the Hailuo line: a 33.1-billion-parameter dense single-stream omni-modal transformer, announced July 31, 2026, weights public on Hugging Face August 3.

Its defining feature is the unified multimodal context. Rather than separate expert models per task, H3 reads text, images, video, and audio together and expresses text-to-video, editing, and reference generation as instructions over that shared context. It accepts up to 9 reference images, 3 reference video clips, and 3 reference audio clips, capped at 12 files total; audio references require at least one image or video alongside them.

Resolution, Duration, and Framing

Resolution

LTX 2.5 Fast reaches 4K, with 1440p, 1080p, and 720p below it. MiniMax H3 tops out at what it calls 2K, at 24fps.

That label deserves unpacking, because the gap is wider than “4K vs 2K” suggests. H3’s 2K puts 1440 pixels on the short edge — a 16:9 clip lands near 2560×1440, and wider formats reach roughly 3.7 megapixels, around 2976×1248 at 21:9. LTX 2.5 Fast’s 4K is 3840×2160, about 8.3 megapixels. H3’s ceiling is nearer half LTX 2.5’s pixel count than two-thirds. A 768P tier appears on H3 price sheets, but MiniMax’s documentation has described it as a later rollout, so verify availability before budgeting around it. If you need true 4K, H3 is out.

Duration

Duration favours LTX 2.5, with a caveat. Fast reaches 20 seconds only at 720p or 1080p and 24–25fps; at 48–50fps, or at 1440p and 4K, it caps at 10. Pro caps at 10 throughout. H3 generates 4 to 15 seconds with no resolution-dependent ceiling.

There is a floor difference too. H3 starts at 4 seconds against LTX 2.5’s 6, so for quick cutaways and social stingers you are not paying for two seconds you will trim anyway.

Frame rate goes to LTX 2.5: 24, 25, 48, and 50fps against H3’s documented 24.

Aspect Ratios

This is where the LTX 2.5 vs MiniMax H3 comparison flips, and most spec tables understate it.

LTX 2.5 generates 16:9 and 9:16. That is the whole list. No 1:1, no 4:5, no 21:9.

H3 supports 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, plus an adaptive mode where the model picks framing from your references. A first-and-last-frame run inherits the ratio of the image you upload.

For anyone shipping to more than two placements, that is a real workflow cost on the LTX 2.5 side. Cropping 16:9 to 1:1 throws away composition the model never planned for.

First and Last Frame Control

LTX 2.5 supports a last_frame_uri parameter on image-to-video, which fixes the closing composition but cannot be combined with automatic duration.

H3 treats this as a first-class mode: first-frame, last-frame, or both together, handled separately from its reference workflow. For product reveals, logo animations, and transformation shots, that is the cleaner path.

Prompt

Start with the product upright on the dark pedestal, exactly as in the first-frame image. The camera rotates slowly clockwise while a narrow rim light travels across the metal.

End with the product positioned as in the final-frame reference, logo facing camera.

Pricing Compared

Here the LTX 2.5 vs MiniMax H3 comparison produces a result most spec tables miss.

MiniMax lists H3 at $0.13 per generated second at 2K. A cheaper 768P tier appears on price sheets near $0.08, but availability is inconsistent across providers and MiniMax’s documentation has treated it as a later rollout — so 2K is the rate to plan against.

LTX 2.5 bills by variant and resolution: Fast at $0.09 (720p), $0.13 (1080p), $0.19 (1440p), $0.30 (4K); Pro at $0.12 (720p) and $0.17 (1080p).

So LTX 2.5 Fast at 1080p and MiniMax H3 at 2K cost the same per second — $0.13, with H3 delivering more pixels at that rate. Ten seconds runs $1.30 on either. H3’s 15-second maximum is $1.95; LTX 2.5 Fast’s 20-second maximum at 1080p is $2.60. If you need 4K, LTX 2.5 Fast charges $0.30 and H3 has no equivalent tier.

The Surcharges Most Comparisons Skip

Neither headline rate is the whole invoice.

For H3, reference audio and the first five reference images are free, with each additional image adding a small fee. Reference video is billed by its own input duration at the output rate — so a five-second reference clip feeding a six-second generation bills eleven seconds, not six. For reference-heavy work that changes the arithmetic materially.

For LTX 2.5, audio-to-video is priced separately and also billed on input audio duration. On prepaid accounts, automatic duration reserves credits against the longest clip your settings could produce.

Third-party hosts price both well above list. Treat any aggregator rate card as its own number.

Try LTX 2.5 in JXP

Open Weights: The Difference That Decides It

Both LTX 2.5 and MiniMax H3 are described as open weights. They are not equivalent, and this matters more than any benchmark.

LTX 2.5 licensing

Free for commercial and production use under $10 million in annual revenue via the LTX-2.x Community License; above that, organizations negotiate a paid agreement, and transferring fine-tunes may carry additional requirements. No comparable excluded-territory clause was identified in the published community-license terms reviewed for this comparison. The Hugging Face repository is gated — accept the licence and authenticate before download works — but acceptance is open to anyone.

MiniMax H3 licensing and the territory clause

H3’s weights ship under the MiniMax H3 Community License Agreement, dated August 2, 2026, permitting commercial use under roughly $20 million in revenue with attribution — a higher threshold than LTX’s.

But the licence defines an “Applicable Territory” that excludes the United States, the European Union, the United Kingdom, and South Korea. Reporting on the licence text indicates that, within those regions, the default terms do not extend rights to run, modify, distribute, or deploy the outputs of locally-run H3 weights. Separate authorization can be applied for, but the Community License is what applies by default to anyone downloading from the public repository.

For a team in San Francisco, Berlin, London, or Seoul, that is not a footnote. Local deployment under the default Community License requires additional licensing consideration in those territories; MiniMax’s hosted service is a separate access path governed by its own terms. Read the licence directly and take legal advice rather than relying on a summary, including this one.

What is missing from H3’s open release

Even where the licence permits local deployment, the open release is partial.

What shipped is H3-Base: two task-specific checkpoints — FL2VA for text-to-video with first- and last-frame conditioning, and Ref2VA for reference-to-video — plus the video VAE, audio VAE, and a Qwen3-VL-32B text encoder, which ships under Apache 2.0 rather than the Community License covering the rest.

Two modules did not ship. H3-Context-IR, which parses and refines the multimodal context, remains hosted. H3-Regenerate-2K, which lifts output to 2K, is not open-sourced.

The practical consequence: a fully local H3 deployment does not currently produce 2K. H3-Base generates at a native canvas of 768 pixels on the short edge, and MiniMax’s documented path for full 2K combines local H3-Base with API calls for the closed modules. If your reason for self-hosting was to avoid the API, that reason survives only partially.

LTX 2.5 has no equivalent split. The distilled checkpoints, the trainable development transformer, the text encoder, and both VAEs are downloadable, and the pipeline runs end to end locally.

Put as a set of separate questions rather than one label:

Open-model question

LTX 2.5

MiniMax H3

Downloadable weights

Yes

Yes

Fine-tuning path

Yes

Yes

Local inference

Yes

Yes, subject to licence and workflow constraints

Complete official local pipeline

More complete

Partial

Local path to the model’s top resolution

Yes

2K requires the non-open modules

Comparable excluded-territory clause

Not identified

US, EU, UK, South Korea by default

Control Paradigms: Multi-Shot vs Multimodal Reference

This is the section that outlasts the spec sheet: prices and ceilings change, the control paradigm does not. LTX 2.5 asks how to keep a sequence coherent across cuts; H3 asks how to make a generation obey material you already have.

LTX 2.5: continuity across cuts

LTX 2.5 wants you to describe a sequence — name the character and lighting once, then label the shots. The model handles identity across cuts, which is the work you would otherwise do by generating three clips and fighting drift between them. Test 1 below gives the full prompt structure.

MiniMax H3: obedience to supplied material

H3 wants you to supply material and describe its use. The prompt does less describing, more instructing: this image for the face, that one for the product, this clip for the camera move, this audio for timing. Test 2 below shows the pattern.

Neither substitutes for the other. A three-shot narrative sequence is awkward to assemble from H3’s reference workflow, and LTX 2.5 documents no equivalent to matching a client’s product from nine reference images.

Three Matched Tests

Specifications explain what each model is designed to do; output shows whether the advantage survives. These three LTX 2.5 vs MiniMax H3 tests press on different failure modes — run each prompt through both models and score against the same criteria.

Test 1: Multi-Shot Continuity

LTX 2.5

MiniMax H3

Prompt

A woman with short black hair, a dark green trench coat, and a small silver camera walks through an abandoned railway station at sunrise.

Shot one: wide establishing shot from the opposite platform as she walks beside the train. Shot two: medium tracking shot moving parallel to her as she checks the camera. Shot three: close-up as she stops and looks toward the arriving train.

Preserve her face, hairstyle, clothing, camera, station architecture, and blue dawn lighting across every shot. Natural station ambience.

Score on: face consistency, prop consistency, shot transitions, environment continuity.

Pressure point: LTX 2.5’s native multi-shot workflow. This is the test H3 is not built to win.

Test 2: Product Reference Fidelity

Feed both models the same product image as reference.

LTX 2.5

MiniMax H3

Prompt

Create a premium studio advertisement from the reference product image. Slowly orbit clockwise while a narrow rim light travels across the polished surface.

Preserve the exact shape, logo placement, typography, proportions, material, and colours of the reference. Dark luxury background, stable geometry, restrained motion.

Score on: logo preservation, geometry, typography, invented detail.

Pressure point: H3’s reference-driven control, and its first-and-last-frame mode if you specify an end state.

Test 3: Dialogue and Native Audio

LTX 2.5

MiniMax H3

Prompt

A detective sits alone in a quiet diner during a thunderstorm. Medium close-up, slow push-in. Rain runs down the window as distant thunder rolls.

He looks across the table and says quietly, “You were never supposed to find this place.” Restrained acting, realistic room tone.

Score on: lip synchronization, speech naturalness, audio-video timing.

Pressure point: both models’ native audio, generated in the same pass as the picture.

Record how many generations each model needed before one clip was usable. That number, multiplied by the per-second rate, is the cost comparison that matters most — and the one no spec table can give you.

Try LTX 2.5 in JXP

Try MiniMax H3 in JXP

Benchmarks: Who Claims What

Treat this as claims, not conclusions. Independent head-to-head testing of LTX 2.5 vs MiniMax H3 does not yet exist.

On the Artificial Analysis leaderboards, H3 has been reported first in video editing with audio and top-three for text-to-video and image-to-video. Those are third-party rankings rather than MiniMax’s own, which makes them more useful than most launch numbers — but no leaderboard tells you whether your product logo survives motion, how often a face drifts, or how many generations you need before one clip is usable.

LTX’s quality and speed figures are vendor-run: an artifact score placing LTX 2.5 Pro first among ten models, and a blind comparison above the even-split threshold, both labelled preliminary. Its speed benchmark reports a ten-second clip in 6.8 seconds self-hosted — on two NVIDIA GB200 chips, against competitors at unmatched resolutions.

A definitive quality verdict this early is extrapolation.

Local Deployment

Both landed with day-one ComfyUI support. MiniMax has stated H3 runs on GPUs including the RTX 5090 and RTX 6000 without publishing supporting benchmarks, while its own reference launch command uses four GPUs. LTX 2.5 ships distilled checkpoints at roughly 42 GB in BF16, 21.5 GB in ComfyUI INT8, and 18.7 GB in NVFP4, with FP8 quantization and CPU offloading documented. Independent VRAM testing for either is thin — validate at minimum settings on your own hardware first.

Category by Category

The spec values sit in the head-to-head table above. These are the LTX 2.5 vs MiniMax H3 judgment calls those numbers do not settle on their own.

Category

Better choice

Sequence continuity across cuts

LTX 2.5

Multimodal reference control

MiniMax H3

Resolution and duration ceiling

LTX 2.5 Fast

Delivery format flexibility

MiniMax H3

Short-clip economics

MiniMax H3

Price per second, practical tier

Tie at $0.13

Completeness of the open release

LTX 2.5

Licence territory

LTX 2.5

Overall visual quality

Needs controlled testing

The last row is the one that matters. No vendor demo settles it, and the three tests above are how you settle it for your own work.

LTX 2.5 vs MiniMax H3: The Verdict

The question worth asking is not which model is better. It is whether your harder problem is holding a sequence together across cuts, or making a generation obey assets you already have. That single distinction predicts the right answer more reliably than any spec in this article.

Two things can override it. If your delivery formats go beyond 16:9 and 9:16, H3’s framing options may matter more than LTX 2.5’s ceiling. And if you are self-hosting from a region H3’s default licence does not cover, the licence decides before the specs do.

What should not decide it is the benchmark charts. LTX’s are vendor-run and marked preliminary; H3’s leaderboard position is third-party but measures ranked preference, not fit for your brief.

Frequently Asked Questions

Is LTX 2.5 or MiniMax H3 better?

Neither across the board. LTX 2.5 leads on resolution, clip length, frame rates, and multi-shot. MiniMax H3 leads on reference control, aspect ratios, and editing, which the current LTX 2.5 API does not cover. For teams in the US, EU, UK, or South Korea, H3’s default licence adds a significant constraint on local deployment, which often settles it.

Which is cheaper, LTX 2.5 or MiniMax H3?

At the most common settings they are identical: LTX 2.5 Fast at 1080p and H3 at 2K both cost $0.13 per generated second. LTX 2.5 is cheaper at 720p and offers 4K at $0.30, which H3 has no equivalent for. Reference inputs can add materially to an H3 invoice.

Are both LTX 2.5 and MiniMax H3 open weight?

Both publish weights, but not equivalently. LTX 2.5’s release runs end to end locally. H3 published H3-Base while keeping H3-Context-IR hosted and H3-Regenerate-2K closed, so a fully local H3 deployment does not currently reach 2K.

Can I self-host MiniMax H3 in the US or EU?

Not straightforwardly. The default Community License defines an Applicable Territory that excludes the United States, the European Union, the United Kingdom, and South Korea, so local deployment there needs additional licensing consideration. Separate authorization can be applied for, and MiniMax’s hosted service is a separate access path under its own terms. Consult the licence and your own counsel rather than a summary.

Which model makes longer videos?

LTX 2.5 Fast reaches 20 seconds, but only at 720p or 1080p and 24–25fps. H3 supports 4 to 15 seconds at any setting, and starts shorter than LTX 2.5’s six-second floor.

Does MiniMax H3 support multi-shot generation like LTX 2.5?

Native multi-shot within a single generation is documented for LTX 2.5. H3’s documented strengths are the unified multimodal context, reference-driven generation, and editing rather than shot sequencing.

Do both models generate audio?

Yes, in the same pass as the picture with no separate scoring step. H3’s is documented as native stereo.

Which model supports more aspect ratios?

MiniMax H3, by a wide margin: 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 plus an adaptive mode. LTX 2.5 generates 16:9 and 9:16 only, so other formats require cropping.

Is MiniMax H3’s 2K the same as 1440p?

Effectively yes. H3’s 2K puts 1440 pixels on the short edge, so a 16:9 clip is close to 2560×1440 and 21:9 reaches roughly 3.7 megapixels. LTX 2.5 Fast’s 4K is 3840×2160, about 8.3 megapixels. LTX 2.5 Pro, despite the name, caps at 1080p.

Which model is better for first-and-last-frame workflows?

MiniMax H3, which handles first-frame, last-frame, and both as documented modes. LTX 2.5 offers a last-frame parameter that cannot be combined with automatic duration.

Which is better for product videos?

MiniMax H3, in most cases — up to nine reference images plus reference clips give tighter control over a product’s proportions, labels, and materials than a text prompt alone.