Creating Videos
Sogni Studio generates video from text, images, audio, or other videos. Pick a workflow based on what you've got to start from — a written idea, a still you already love, an audio clip, or a reference video whose motion you want to borrow.
#Workflows
Sogni Studio auto-picks the right workflow based on the inputs you provide:
| Workflow | When to use it | Inputs |
|---|---|---|
| Text-to-Video (T2V) | Generate motion from a written prompt alone. | Prompt |
| Image-to-Video (I2V) | Animate a still image. Optionally pin start and end frames for continuity. | Start frame (+ optional end frame) + prompt |
| Sound-to-Video (S2V) | Drive motion from an audio track — music, dialogue, ambient sound. | Reference image + audio |
| Audio-to-Video (A2V) | Generate video from audio alone, no reference image. | Audio |
| Video-to-Video (V2V) | Preserve motion or structure while restyling, enhancing, inpainting, or outpainting a source video. | Source video (+ mode-specific image or mask; LTX-2.5 pose requires a subject image) |
| Motion Transfer | Apply another video's motion to your character. | Reference image + source video |
| AnimateReplace | Keep a video's motion and scene, swap the subject. | Reference image + source video |
See Convert Image to Video for the quickest path (a single image, no reference), or Video Create Mode for the full control panel with keyframe-level overrides.
#Models
Video runs on the Sogni Supernet and uses video-specialized models. Current lineup:
- LTX-2.5 — the default Sogni-native family for text, image, first/last-frame, audio, image+audio, and controlled V2V generation with native synchronized audio. Fast, HQ, and Pro use the release-validated official Distilled/Turbo workflows.
- LTX-2.3 — explicit rollback generation plus voice ID-LoRA, transition-LoRA, and 10Eros workflows that do not carry forward to LTX-2.5.
- Wan 2.2 — first-party Sogni video model, still available.
- ByteDance Seedance 2.0 — premium hosted video in full, Fast, and Mini tiers. Use Seedance 2.0 Mini for lower-cost 720p iteration.
- Alibaba HappyHorse 1.1 — premium hosted video for text-to-video, image-to-video, and image-reference workflows with native synchronized audio.
Pick a model preset in Video Create mode; Studio matches the workflow to what each model supports. LTX-2.5 Dev checkpoints are not publicly routed until an official upstream ComfyUI Dev recipe is available and validated. See Pricing for modes, durations, resolutions, and current Spark estimates.
#After generation: stitching clips together
Generated clips can be composed into longer pieces using the Clip Mixer, a desktop-only timeline editor with FPS, aspect-ratio, and re-encode controls. See Clip Mixer.
#Privacy
Generated files are saved only on your Mac in the gallery folder you choose. Supernet rendering processes the job and purges inputs after a brief retention period — your videos don't get added to any public dataset.