LTX-2.5 model and workflow reference
LTX-2.5 is Sogni's default native video family: a 22B audio-video transformer that generates picture and synchronized 48 kHz stereo audio in a single pass. Fast, HQ, and Pro currently route through the release-validated official Distilled/Turbo workflows. Dev checkpoints are not publicly routed until upstream publishes an official ComfyUI Dev recipe and Sogni validates it.
Every public mode runs a fixed 8-step distilled path at guidance 1, so an LTX-2.5 render of a given shape costs and takes roughly what any other LTX-2.5 render of that shape does.
For a mode-by-mode walkthrough with side-by-side comparisons against MiniMax H3 and LTX-2.3, see H3 Directs. LTX-2.5 Brings the Camera Rig. on the Sogni blog.
#Workflow IDs
| Workflow | Creative Agent selector | Public model ID | Inputs |
|---|---|---|---|
| Text to video | ltx25 |
ltx25-22b-int8_t2v_distilled |
Prompt |
| Image to video | ltx25 |
ltx25-22b-int8_i2v_distilled |
Prompt + first frame |
| First/last-frame video | ltx25 with both frame roles |
ltx25-22b-int8_i2v_distilled |
Prompt + first and last frames |
| Audio to video | ltx25-a2v |
ltx25-22b-int8_a2v_distilled |
Prompt + audio |
| Image + audio to video | ltx25-ia2v |
ltx25-22b-int8_ia2v_distilled |
Prompt + first frame + audio |
| Video to video | ltx25-v2v |
ltx25-22b-int8_v2v_distilled |
Prompt + source video (+ mask for inpaint) |
First/last-frame generation uses a dedicated FLF workflow template while sharing the public I2V model ID. There is no public Dev selector.
#Generation limits
- Resolution 640–3840 px per side on a 64 px grid, default 1920×1088. Video-to-video uses a 128 px grid, default 1920×1024.
- Frames 25–505 on a
1 + 8ngrid, default 97 — roughly 1–21 seconds at 24 fps. Sogni's current public workflows are exercised up to 20 seconds. - Frame rate is yours to set; 24 fps is the cinematic default.
- Audio 48 kHz stereo, generated jointly with the picture rather than dubbed on afterwards.
- Steps fixed at 8 for every public distilled mode; guidance fixed at 1.
Resolution and frame rate trade against each other at the top of the range: 4K is available, and 60 fps is available at 1080p, but not both at once.
#Pricing
LTX-2.5 is priced from the shape of the job — pixels × frames — rather than a flat per-second rate, so a short 4K clip and a longer 1080p clip can land at similar cost. Measured examples from Sogni reference runs on fast-network workers:
| Mode | Resolution | Length | Spark | USD |
|---|---|---|---|---|
| Multishot text to video | 1920×1088 | 10 s | 91.83 | $0.46 |
| Image to video | 1088×1536 | 14 s | 102.72 | $0.51 |
| First/last-frame video | 1920×1088 | 8 s | 73.54 | $0.37 |
| Video to video (canny) | 1920×1024 | 8 s | 69.14 | $0.35 |
Prices are live network estimates at the standard 1 Spark = $0.005 peg. For a quote at any shape, use the job estimate calculator. Every LTX-2.5 render is also covered credit-free under fair use on Sogni Unlimited plans.
#Render time
LTX-2.5 runs on Sogni Supernet fast-network workers, whose GPUs differ, so wall-clock time varies by job. In Sogni reference runs an 8-second 1080p render ranged from about 1 m 20 s to over 3 m depending on the worker. All public modes run 8 steps, against 30 for the Dev tier.
#Video-to-video controls
The LTX-2.5 V2V family supports canny, depth, pose, detailer, inpaint, and outpaint templates. Pose requires a source video for the motion sequence and a still reference image for the subject's appearance. Inpaint also requires a mask.
#Prompt shape
Use one continuous natural-language scene paragraph with concrete subject, action, setting, lighting, camera motion, and synchronized audio cues.
- Put spoken dialogue in quotation marks, and name the language and accent when it matters.
- Break long speeches into short fragments with acting directions between them rather than one unbroken block — this is upstream LTX guidance and it materially changes delivery.
- State the camera on every shot, including when it should not move. Leaving it unstated invites drift in the final seconds.
- Keep the amount of speech realistic for the requested duration.
- For image to video, treat the still as visual truth: describe motion, performance, camera and sound, not the static contents the model can already see.
- For first/last-frame video, describe the motion and change between the supplied anchors instead of re-inventing their static contents.
#Multi-shot prompting
LTX-2.5 will cut between shots inside a single generation. To get a real hard cut rather than a dissolve, give each shot its own sentence, state a numeric timestamp for the cut, and make the shots genuinely different places. A change of room with a change of light reads as an edit; a tighter framing of the same subject usually resolves as a push-in instead. Keep multi-shot captions to roughly one sentence per shot.
#Sample gallery
More comparisons, including matched A/B pairs against LTX-2.3 and MiniMax H3, are in the LTX-2.5 launch post.
#LTX-2.3 rollback boundary
LTX-2.3 remains available through explicit rollback selectors. Voice ID-LoRA, the community transition LoRA, and 10Eros remain LTX-2.3-only; do not send those inputs or triggers to an LTX-2.5 model ID.