This LTX 2.5 review covers what changed, how the Fast and Pro variants differ, what LTX 2.5 costs per second, and which limitations matter before you commit a pipeline to it. The release adds native multi-shot generation, a new diffusion video decoder, a stronger text encoder, automatic duration, 4K on Fast, and downloadable weights you can fine-tune and self-host.
One detail belongs up front, because it reverses the common assumption: in LTX 2.5, “Pro” does not mean the higher ceiling. Fast reaches 4K and clips up to 20 seconds under supported settings. Pro tops out at 1080p and 10 seconds, trading range for fidelity.
LTX 2.5 Review: Quick Verdict
Native multi-shot is the headline of this LTX 2.5 review: one generation can produce several connected shots holding character identity, environment, lighting, style, and voice across cuts, where earlier versions produced a single continuous shot.
The reservations are equally concrete. The API surface is narrower than the previous generation, aspect ratio support is limited to two options, and the quality and speed figures now circulating are vendor-produced and labelled preliminary.
LTX 2.5 at a Glance
Feature | LTX 2.5 |
|---|---|
Modes | Text-, image-, audio-to-video |
Native audio generation | Yes |
Native multi-shot | Yes |
Fast max resolution | 4K |
Pro max resolution | 1080p |
Fast max duration | 20 sec |
Pro max duration | 10 sec |
Aspect ratios | 16:9, 9:16 |
Automatic duration | Yes |
Open weights, fine-tuning | Yes |
Main transformer | 22B |
Retake / extend / reframe | No |
What Is LTX 2.5?
LTX 2.5 is an open-weight video and world model from LTX, the company spun out of Lightricks, built on a 22-billion-parameter asymmetric dual-stream diffusion transformer with weights and code publicly available.
Three access paths shipped together: downloadable weights on Hugging Face, day-one ComfyUI support, and the hosted LTX 2.5 API. Licensing is free for commercial and production use under $10 million in annual revenue under the LTX-2.x Community License. Above that, organizations negotiate a paid agreement, and transferring fine-tunes may carry additional requirements.
What Is New in LTX 2.5
Native Multi-Shot Generation
This is the most important change in this LTX 2.5 review. Earlier workflows optimized around a single continuous shot, so building a sequence meant generating separate clips and fighting identity drift at every cut. LTX 2.5 moves that problem into the model. The prompt pattern that works: name the character and lighting once, then label each shot. Test 1 below shows the full structure.
Diffusion Fidelity Rendering and the New Decoder
LTX calls its new rendering approach Diffusion Fidelity Rendering, describing it as allocating compute by scene complexity rather than spending equal effort everywhere. The documented mechanism is a new diffusion video decoder replacing the previous VAE reconstruction stage, targeting faces, textures, on-screen text, and motion.
Treat “industry-leading pixel quality” as vendor language; the mechanism is the substantive claim. It matters commercially: AI video often looks fine in motion and falls apart when a viewer pauses on a hand or a label.
Better Prompt Adherence
LTX 2.5 ships with a custom Gemma 4 12B text encoder, fine-tuned for LTX and bundled with the model. Google’s stock Gemma 4 release is not a substitute — loading validates the encoder version against the checkpoint. A lightweight prompt enhancer expands short prompts into detailed instructions, so you can write shorter and let it elaborate.
Automatic Duration, and Its Three Catches
LTX 2.5 can predict clip length from the described action: send duration as null and the model chooses.
The catches are the kind of detail that costs an afternoon.
The field is still required — omitting it returns a duration is required error, so pass null explicitly.
It cannot be combined with last_frame_uri on image-to-video, since a fixed last frame requires a known length.
Prepaid accounts hold the maximum. Credits are reserved against the longest duration your resolution and frame rate allow, because final length isn’t known until generation completes. A request that would have produced six seconds is still declined if your balance can’t cover 20 seconds on Fast. The remainder releases on completion; postpaid accounts hold nothing and are invoiced for what was generated.
LTX 2.5 Fast vs Pro
The names invite the assumption that Pro does everything Fast does plus better quality. That is not how the current API works.
LTX 2.5 Fast
Optimized for speed and cost. Supports 720p through 4K, portrait and landscape, in all three modes. At 720p or 1080p and 24–25 FPS, Fast supports fixed durations from six up to 20 seconds; at 48–50 FPS those cap at 10 seconds, as do all frame rates at 1440p and 4K.
LTX 2.5 Pro
Built for fidelity rather than range. Supports 720p and 1080p only, at 24, 25, or 50 FPS, with durations of six, eight, or ten seconds. Pro does not offer 1440p or 4K through the documented API.
Resolution, FPS, and Duration Matrix
Model | Resolution | FPS | Max fixed duration |
|---|---|---|---|
Fast | 720p | 24/25 | 20 sec |
Fast | 720p | 48/50 | 10 sec |
Fast | 1080p | 24/25 | 20 sec |
Fast | 1080p | 48/50 | 10 sec |
Fast | 1440p | 24/25/48/50 | 10 sec |
Fast | 4K | 24/25/48/50 | 10 sec |
Pro | 720p | 24/25/50 | 10 sec |
Pro | 1080p | 24/25/50 | 10 sec |
Which Variant Should You Use?
Use Fast for iteration, lower cost, longer clips, and output above 1080p. Use Pro for final 720p or 1080p shots where artifacts matter more than resolution. Explore on Fast, lock the prompt and camera, then re-run the final shot on Pro.
LTX 2.5 API Pricing
The numbers matter here. LTX 2.5 bills per second of output video. Current documented rates for text-to-video and image-to-video:
Variant | Resolution | Cost per second |
|---|---|---|
Fast | 720p | $0.09 |
Fast | 1080p | $0.13 |
Fast | 1440p | $0.19 |
Fast | 4K | $0.30 |
Pro | 720p | $0.12 |
Pro | 1080p | $0.17 |
Audio-to-video is priced separately at $0.13 for Fast and $0.17 for Pro, both at 1080p — and billed on the duration of the input audio, not the output.
A ten-second LTX 2.5 clip therefore costs roughly $0.90 on Fast 720p, $1.30 on Fast 1080p, $3.00 on Fast 4K, and $1.70 on Pro 1080p.
One caution for anyone budgeting from a secondary source: LTX’s marketing page and its API documentation currently list different figures. The documentation is the more granular of the two, splitting rates by variant — verify against it before committing a large budget.
An Independent Price Check
VentureBeat compared LTX 2.5 against competitors publishing per-second pricing and found the steep cost-multiple framing doesn’t hold. By their accounting, roughly $0.90 for ten seconds at 720p runs about a quarter the cost of full Veo 3.1 and half of FLUX 3 Video, but only about 10% under Google’s budget tiers — with Veo 3.1 Lite cheaper still.
The honest read for this LTX 2.5 review: competitively priced, not category-breaking on hosted cost alone. The real economic argument is self-hosting, which removes per-generation billing. And per-second rates only tell you half the story — what decides a real budget is cost per usable clip, meaning how many takes a shot needs before one is shippable.
Generate your first LTX 2.5 clip in JXP
Native Audio and Audio-to-Video
LTX treats audio as part of generation rather than a separate post step. Both LTX 2.5 variants generate synchronized audio by default and expose an audio-to-video endpoint.
Audio-to-video inverts the usual order: you supply recorded dialogue, narration, or music, and the model builds the visual performance around it rather than chasing sync afterward.
Two notes. It is documented at 1080p only, so it sits outside the 4K path. And because it bills on input audio length, trim audio before submitting, not after. For silent output, generate_audio: false skips audio generation.
Speed and Quality Benchmarks: Read the Footnotes
The Speed Number
LTX reports a ten-second image-to-video clip generating in 6.8 seconds self-hosted and 23.7 seconds through its API, with the fastest closed alternatives in the same comparison at 52–70 seconds.
The footnotes, which LTX publishes and most coverage drops: the on-prem figure was measured on two NVIDIA GB200 chips at steady state, not a consumer GPU. Resolutions are not matched — several competitors at 720p, one at 768p, one near 1080p. The LTX API figure is internal, at 1080p. Veo 3.1 was measured on an eight-second clip.
The Artifact Benchmark
Many early LTX 2.5 comparisons cite this benchmark. LTX published a preliminary artifact score counting visible glitches per clip, lower being better:
Model | Visible glitches per clip |
|---|---|
LTX 2.5 Pro | 0.28 |
LTX 2.5 Fast | 0.39 |
LTX 2.3 Pro | 0.74 |
The test ran 98 prompts through ten models with automated rather than human scoring, and LTX labels the results preliminary. A separate blind-comparison chart reports LTX 2.5 above the even-split threshold.
These show the direction of improvement within LTX’s own evaluation, not independent confirmation against competing models. Cite them as vendor-reported, or don’t cite them.
Open Weights and Local Deployment
What You Actually Download
The repository ships several checkpoints: the distilled transformer in BF16 at roughly 42 GB, a ComfyUI INT8 build at roughly 21.5 GB, and an NVFP4 build at roughly 18.7 GB. A separate full development transformer in BF16 is the trainable base for fine-tuning and LoRA work.
Around those sit the bundled Gemma 4 12B text encoder, video and audio VAEs, and a duration head. The distilled pipeline also expects a spatial upscaler still hosted in the LTX 2.3 repository; a full distilled download runs roughly 66 GiB.
Documented efficiency levers include FP8 quantization on BF16 checkpoints and CPU offloading. Independent VRAM testing for LTX 2.5 does not exist yet, and figures from earlier releases should not be assumed to carry over — validate at minimum settings first, then scale.
The Gated Repository
This detail is easy to miss. The Hugging Face repository is gated: you must accept the licence conditions and authenticate before download works. A 401 or 403 means terms haven’t been accepted, and fine-grained tokens need the read-gated-repos scope. It is not a one-click pull.
LTX 2.5 Limitations
The feature list isn’t the whole picture. This LTX 2.5 review’s list of constraints that bite:
Pro does not reach 4K. LTX 2.5 is marketed with native 4K, but the documented API assigns 4K to Fast. Pro is 1080p and below.
Pro caps at ten seconds, where Fast reaches 20 under supported 720p and 1080p 24–25 FPS settings.
Editing endpoints did not migrate. Retake, extend, reframe, and HDR upscale remain listed and priced in the LTX API — but against ltx-2-3-pro, not LTX 2.5. They have not disappeared; they have not moved to the new model.
Only two aspect ratios. LTX 2.5 generates 16:9 and 9:16. There is no 1:1 or 4:5, so those formats require cropping after generation.
Prompt quality still matters. LTX’s model cards note that prompt following is heavily influenced by prompting style; a better encoder does not remove the need for clear instructions.
Multi-shot is not guaranteed identity locking. Angle, expression, occlusion, and camera distance all introduce variation across cuts. It is a workflow improvement, not a guarantee — demanding character work still needs shot-by-shot review.
Five Prompts to Test LTX 2.5
Probe different failure modes rather than five similar landscapes. Treat these as starting points, not verified outputs.
Test 1: Multi-Shot Character Consistency
Prompt
A young photographer with short black hair, a green field jacket, and a vintage silver camera explores an abandoned railway station at sunrise.
Shot one: wide establishing shot as she enters the empty platform. Shot two: medium tracking shot as she raises the camera beside an old train. Shot three: close-up as she takes a photograph and lowers the camera.
Preserve her face, hairstyle, clothing, camera, and warm sunrise lighting across every cut.
Tests identity retention across cuts.
Test 2: Complex Prompt Adherence
Prompt
Inside a dark jazz club, a saxophone player performs beneath a single warm spotlight while blue smoke drifts behind him. The camera begins with a slow push-in, then gently arcs to his left. A drummer stays softly out of focus in the background. Reflections shimmer on the brass.
Tests whether camera motion, lighting, and background detail survive together.
Test 3: Product Image-to-Video
Prompt
Slow cinematic orbit around the product while a narrow rim light travels across its polished surface. Keep proportions, materials, labels, logo placement, and colours of the reference image unchanged. Stable geometry, restrained motion.
Tests object preservation and invented geometry.
Test 4: Fast Motion
Prompt
A rally car races through a wet forest road at dawn. Low tracking camera beside the front wheel, water spraying from the tires through a sharp corner. The camera swings behind the vehicle as it exits the turn.
Tests motion coherence and temporal artifacts.
Test 5: Dialogue and Native Audio
Prompt
A detective sits alone in a quiet diner during a thunderstorm. Medium close-up. Rain runs down the window as distant thunder rolls. He looks across the table and quietly says, “You were never supposed to find this place.” Natural speech, soft fluorescent lighting.
Tests lip synchronization, ambience, and facial motion together.
Who Should Use LTX 2.5
Strong fit. Teams needing multi-shot continuity in one generation. Studios with on-prem GPUs that want to remove per-generation billing. Anyone fine-tuning on proprietary data, since LTX 2.5 ships a trainable development checkpoint. Finishing pipelines needing 4K output.
Weak fit. Teams choosing purely on hosted quality, where the comparison is contested and the available benchmarks are vendor-run.
Who Should Wait
If your LTX 2.3 pipeline is stable, there is no forced migration — LTX 2.3 remains supported and still holds the editing endpoints. Native multi-shot is probably the clearest reason to consider moving now. Teams depending on retake, extend, or reframe should wait until those workflows reach LTX 2.5.
Test LTX 2.5 multi-shot consistency in JXP
LTX 2.5 Review: Final Verdict
LTX 2.5 is worth testing, and the substance sits in native multi-shot and open licensing rather than in the benchmark charts. An openly licensed 22B audio-video model, free under $10M in annual revenue, with day-one ComfyUI support, is a different proposition from one that lives only behind a hosted interface.
The tradeoffs are real: Pro stops at 1080p and ten seconds, editing endpoints have not migrated, aspect ratios are limited to two, local execution stays hardware-intensive, and the most flattering numbers are LTX’s own preliminary testing. The fair summary for anyone using this LTX 2.5 review to decide — strong on openness, control, and multi-shot continuity; unproven on independent quality comparison; narrower on API surface than a version increment implies.
Frequently Asked Questions
Is LTX 2.5 open weight?
Yes. LTX 2.5 provides downloadable weights and supports self-hosting, fine-tuning, and local deployment. Commercial use is free under the LTX-2.x Community License below $10 million in annual revenue.
Does LTX 2.5 support 4K video?
Yes, with a distinction. The documented API lists 4K for Fast; LTX 2.5 Pro currently supports up to 1080p.
How long can LTX 2.5 videos be?
Fast reaches 20 seconds at 720p or 1080p under 24–25 FPS. At higher frame rates or at 1440p and 4K, it caps at 10 seconds. Pro supports six, eight, and ten seconds.
How much does LTX 2.5 cost?
Text-to-video and image-to-video run from $0.09 per second on Fast 720p to $0.30 on Fast 4K. Pro is $0.12 at 720p and $0.17 at 1080p. Audio-to-video is $0.13 on Fast and $0.17 on Pro at 1080p, billed on input audio duration.
Does LTX 2.5 support video editing or extending clips?
Not directly. Retake, extend, reframe, and HDR upscale remain in the LTX API but are documented against LTX 2.3 Pro, not LTX 2.5.
What aspect ratios does LTX 2.5 support?
Only 16:9 and 9:16. Square and 4:5 require cropping after generation.
Can LTX 2.5 be fine-tuned?
Yes. A full development transformer ships as the trainable base for domain adaptation and LoRA work. Transferring fine-tunes may carry extra licensing requirements.
How fast is LTX 2.5 in practice?
LTX reports 6.8 seconds for a ten-second clip on two NVIDIA GB200 chips, and 23.7 seconds via its API. Both are vendor measurements on hardware most teams don’t have, and the comparison uses unmatched resolutions. Expect slower results on consumer GPUs.
Is LTX 2.5 better than LTX 2.3?
LTX 2.5 adds native multi-shot, a new diffusion video decoder, a stronger text encoder, and automatic duration. Whether it is better depends on whether you need those, since LTX 2.3 retains editing endpoints LTX 2.5 does not offer.
Where do I download LTX 2.5 weights?
From the gated Hugging Face repository. Accept the licence conditions and authenticate first; fine-grained tokens require the read-gated-repos scope.
