Sogni: Learn logo

LTX-2.5 model and workflow reference

LTX-2.5 is Sogni's default native video family: a 22B audio-video transformer that generates picture and synchronized 48 kHz stereo audio in a single pass. Fast, HQ, and Pro currently route through the release-validated official Distilled/Turbo workflows. Dev checkpoints are not publicly routed until upstream publishes an official ComfyUI Dev recipe and Sogni validates it.

Every public mode runs a fixed 8-step distilled path at guidance 1, so an LTX-2.5 render of a given shape costs and takes roughly what any other LTX-2.5 render of that shape does.

For a mode-by-mode walkthrough with side-by-side comparisons against MiniMax H3 and LTX-2.3, see H3 Directs. LTX-2.5 Brings the Camera Rig. on the Sogni blog.

#Workflow IDs

Workflow Creative Agent selector Public model ID Inputs
Text to video ltx25 ltx25-22b-int8_t2v_distilled Prompt
Image to video ltx25 ltx25-22b-int8_i2v_distilled Prompt + first frame
First/last-frame video ltx25 with both frame roles ltx25-22b-int8_i2v_distilled Prompt + first and last frames
Audio to video ltx25-a2v ltx25-22b-int8_a2v_distilled Prompt + audio
Image + audio to video ltx25-ia2v ltx25-22b-int8_ia2v_distilled Prompt + first frame + audio
Video to video ltx25-v2v ltx25-22b-int8_v2v_distilled Prompt + source video (+ mask for inpaint)

First/last-frame generation uses a dedicated FLF workflow template while sharing the public I2V model ID. There is no public Dev selector.

#Generation limits

  • Resolution 640–3840 px per side on a 64 px grid, default 1920×1088. Video-to-video uses a 128 px grid, default 1920×1024.
  • Frames 25–505 on a 1 + 8n grid, default 97 — roughly 1–21 seconds at 24 fps. Sogni's current public workflows are exercised up to 20 seconds.
  • Frame rate is yours to set; 24 fps is the cinematic default.
  • Audio 48 kHz stereo, generated jointly with the picture rather than dubbed on afterwards.
  • Steps fixed at 8 for every public distilled mode; guidance fixed at 1.

Resolution and frame rate trade against each other at the top of the range: 4K is available, and 60 fps is available at 1080p, but not both at once.

#Pricing

LTX-2.5 is priced from the shape of the job — pixels × frames — rather than a flat per-second rate, so a short 4K clip and a longer 1080p clip can land at similar cost. Measured examples from Sogni reference runs on fast-network workers:

Mode Resolution Length Spark USD
Multishot text to video 1920×1088 10 s 91.83 $0.46
Image to video 1088×1536 14 s 102.72 $0.51
First/last-frame video 1920×1088 8 s 73.54 $0.37
Video to video (canny) 1920×1024 8 s 69.14 $0.35

Prices are live network estimates at the standard 1 Spark = $0.005 peg. For a quote at any shape, use the job estimate calculator. Every LTX-2.5 render is also covered credit-free under fair use on Sogni Unlimited plans.

#Render time

LTX-2.5 runs on Sogni Supernet fast-network workers, whose GPUs differ, so wall-clock time varies by job. In Sogni reference runs an 8-second 1080p render ranged from about 1 m 20 s to over 3 m depending on the worker. All public modes run 8 steps, against 30 for the Dev tier.

#Video-to-video controls

The LTX-2.5 V2V family supports canny, depth, pose, detailer, inpaint, and outpaint templates. Pose requires a source video for the motion sequence and a still reference image for the subject's appearance. Inpaint also requires a mask.

#Prompt shape

Use one continuous natural-language scene paragraph with concrete subject, action, setting, lighting, camera motion, and synchronized audio cues.

  • Put spoken dialogue in quotation marks, and name the language and accent when it matters.
  • Break long speeches into short fragments with acting directions between them rather than one unbroken block — this is upstream LTX guidance and it materially changes delivery.
  • State the camera on every shot, including when it should not move. Leaving it unstated invites drift in the final seconds.
  • Keep the amount of speech realistic for the requested duration.
  • For image to video, treat the still as visual truth: describe motion, performance, camera and sound, not the static contents the model can already see.
  • For first/last-frame video, describe the motion and change between the supplied anchors instead of re-inventing their static contents.

#Multi-shot prompting

LTX-2.5 will cut between shots inside a single generation. To get a real hard cut rather than a dissolve, give each shot its own sentence, state a numeric timestamp for the cut, and make the shots genuinely different places. A change of room with a change of light reads as an edit; a tighter framing of the same subject usually resolves as a push-in instead. Keep multi-shot captions to roughly one sentence per shot.

Text to video, 1920×1088, 10 s. Three shots and two hard cuts from a single prompt, with synchronized audio — no editing.
Image to video, 1088×1536, 14 s. A macro portrait animated from a single still, holding skin texture and identity through a spoken line.
First/last-frame video, 1920×1088, 8 s. Two anchored stills of the same summit, sunrise to deep night, with nothing moving but the light.

More comparisons, including matched A/B pairs against LTX-2.3 and MiniMax H3, are in the LTX-2.5 launch post.

#LTX-2.3 rollback boundary

LTX-2.3 remains available through explicit rollback selectors. Voice ID-LoRA, the community transition LoRA, and 10Eros remain LTX-2.3-only; do not send those inputs or triggers to an LTX-2.5 model ID.

Last updated 2026-08-25