Sogni: Learn logo

Creating Videos

Sogni Studio generates video from text, images, audio, or other videos. Pick a workflow based on what you've got to start from — a written idea, a still you already love, an audio clip, or a reference video whose motion you want to borrow.

#Workflows

Sogni Studio auto-picks the right workflow based on the inputs you provide:

Workflow When to use it Inputs
Text-to-Video (T2V) Generate motion from a written prompt alone. Prompt
Image-to-Video (I2V) Animate a still image. Optionally pin start and end frames for continuity. Start frame (+ optional end frame) + prompt
Sound-to-Video (S2V) Drive motion from an audio track — music, dialogue, ambient sound. Reference image + audio
Audio-to-Video (A2V) Generate video from audio alone, no reference image. Audio
Video-to-Video (V2V) Preserve motion or structure while restyling, enhancing, inpainting, or outpainting a source video. Source video (+ mode-specific image or mask; LTX-2.5 pose requires a subject image)
Motion Transfer Apply another video's motion to your character. Reference image + source video
AnimateReplace Keep a video's motion and scene, swap the subject. Reference image + source video

See Convert Image to Video for the quickest path (a single image, no reference), or Video Create Mode for the full control panel with keyframe-level overrides.

#Models

Video runs on the Sogni Supernet and uses video-specialized models. Current lineup:

  • LTX-2.5 — the default Sogni-native family for text, image, first/last-frame, audio, image+audio, and controlled V2V generation with native synchronized audio. Fast, HQ, and Pro use the release-validated official Distilled/Turbo workflows.
  • LTX-2.3 — explicit rollback generation plus voice ID-LoRA, transition-LoRA, and 10Eros workflows that do not carry forward to LTX-2.5.
  • Wan 2.2 — first-party Sogni video model, still available.
  • ByteDance Seedance 2.0 — premium hosted video in full, Fast, and Mini tiers. Use Seedance 2.0 Mini for lower-cost 720p iteration.
  • Alibaba HappyHorse 1.1 — premium hosted video for text-to-video, image-to-video, and image-reference workflows with native synchronized audio.

Pick a model preset in Video Create mode; Studio matches the workflow to what each model supports. LTX-2.5 Dev checkpoints are not publicly routed until an official upstream ComfyUI Dev recipe is available and validated. See Pricing for modes, durations, resolutions, and current Spark estimates.

#After generation: stitching clips together

Generated clips can be composed into longer pieces using the Clip Mixer, a desktop-only timeline editor with FPS, aspect-ratio, and re-encode controls. See Clip Mixer.

#Privacy

Generated files are saved only on your Mac in the gallery folder you choose. Supernet rendering processes the job and purges inputs after a brief retention period — your videos don't get added to any public dataset.

#See also

Last updated 2026-08-14