Sogni: Learn logo
Markdown

MiniMax H3 and H3 Turbo video

Sogni runs the open-weight MiniMax H3 FL2VA model through ComfyUI for text-to-video, first-frame image-to-video, and first/last-frame video, plus the separate MiniMax H3 Ref2VA checkpoint for multi-reference video. For FL2VA, FastH3 Turbo is the fastest choice. It runs the FastVideo FastH3 4-step Preview v1 VSA DataFree checkpoint and renders up to 2× faster than the LightX2V 4-step Turbo and up to 6× faster than Standard for comparable 768p, 15-second requests. FastH3 supports the same custom LoRA catalog without changing engines. Base jobs can run on 24 GB-class GPUs, including the NVIDIA GeForce RTX 4090 and RTX 3090, when at least 23 GB of usable VRAM is available; jobs with a custom LoRA route to workers with at least 32 GB. The familiar four-step LightX2V Turbo remains available as a distinct creative look, alongside eight-step Balanced and maximum-quality Standard. Video and synchronized stereo audio are generated together in one pass. Sogni has written authorization from MiniMax to offer MiniMax H3 through the Sogni platform.

New: keyframes at any moment. Pin up to eight images at exact times inside an Image to Video, First and Last Frame, Sound to Video, or Reference to Video clip, on every tier and at 2K. The video lands on each image at its time, and a keyframe with a new angle, place, or light starts a new shot, so you can storyboard a whole clip. Two keyframes are included in the price. See Intermediate keyframes.

New: Reference-to-Video in 2K. The Standard and Balanced Reference-to-Video tiers now deliver 2688×1536 (and 1920×1088) from up to 9 reference images, 3 reference videos, and 3 audio clips, through the two-stage IDs minimax-h3-ref2va-fp8_r2v_2stage and minimax-h3-ref2va-fp8_r2v_balanced_2stage — on 32 GB cards, at the tier's rate plus a per-second class surcharge. See Two-stage output and Pricing.

MiniMax H3 is a separate MiniMax model, not a Seedance derivative. Its short-form multimodal scope is comparable to Seedance 2.0, while the Sogni deployment uses released H3 weights on Sogni Supernet workers.

#Workflow IDs

Mode Model ID Inputs
Text to video minimax-h3-fl2va-fp8_t2v Prompt
Image to video minimax-h3-fl2va-fp8_i2v Prompt + first frame
First/last frame minimax-h3-fl2va-fp8_flf2v Prompt + first and last frames
Reference to video minimax-h3-ref2va-fp8_r2v Prompt + 0–9 images, up to 3 videos and 3 audio clips (12 files total; at least one image or video)
Balanced text to video minimax-h3-fl2va-fp8_t2v_balanced Prompt
Balanced image to video minimax-h3-fl2va-fp8_i2v_balanced Prompt + first frame
Balanced first/last frame minimax-h3-fl2va-fp8_flf2v_balanced Prompt + first and last frames
Balanced reference to video minimax-h3-ref2va-fp8_r2v_balanced Prompt + references as above
LightX2V Turbo text to video minimax-h3-fl2va-fp8_t2v_turbo Prompt
LightX2V Turbo image to video minimax-h3-fl2va-fp8_i2v_turbo Prompt + first frame
LightX2V Turbo first/last frame minimax-h3-fl2va-fp8_flf2v_turbo Prompt + first and last frames
LightX2V Turbo reference to video minimax-h3-ref2va-fp8_r2v_turbo Prompt + references as above
FastH3 Turbo text to video minimax-h3-fastvideo-int8_t2v_turbo Prompt
FastH3 Turbo image to video minimax-h3-fastvideo-int8_i2v_turbo Prompt + first frame
FastH3 Turbo first/last frame minimax-h3-fastvideo-int8_flf2v_turbo Prompt + first and last frames
FastH3 Two-Stage text to video minimax-h3-fastvideo-int8_t2v_turbo_2stage Prompt, sent at the half canvas
FastH3 Two-Stage image to video minimax-h3-fastvideo-int8_i2v_turbo_2stage Prompt + first frame, sent at the half canvas
FastH3 Two-Stage first/last frame minimax-h3-fastvideo-int8_flf2v_turbo_2stage Prompt + first and last frames, sent at the half canvas
Two-Stage reference to video minimax-h3-ref2va-fp8_r2v_2stage Prompt + references as above, sent at the half canvas
Balanced Two-Stage reference to video minimax-h3-ref2va-fp8_r2v_balanced_2stage Prompt + references as above, sent at the half canvas

A _2stage ID takes the same inputs as its one-stage ID, at half the finished size, and delivers the video at exactly 2× — see Two-stage output.

#Two-stage output

A _2stage workflow ID renders the video at the canvas you send, then enlarges it 2× and refines the enlargement before delivery, so the finished video is exactly twice the requested width and height with the same length, audio, and references. Send the half canvas: a 384 px short edge for 768p delivery (672×384 → 1344×768), a 544 px short edge for 1080p (960×544 → 1920×1088), and the 768p canvas for 2K (1344×768 → 2688×1536). Portrait sizes are mirrored, and no two-stage canvas edge may be under 384 px. This is Sogni's own enlargement and refinement on Supernet workers, not native 2K output.

  • FastH3 Two-Stage (minimax-h3-fastvideo-int8_t2v_turbo_2stage, _i2v_turbo_2stage, _flf2v_turbo_2stage) renders both passes on FastH3 and offers all three sizes.
  • Two-Stage Reference-to-Video, up to 2K — minimax-h3-ref2va-fp8_r2v_2stage (Standard, 20 steps) and minimax-h3-ref2va-fp8_r2v_balanced_2stage (Balanced, 8 steps) — renders the first pass on the tier's own Ref2VA model with the same reference set as the one-stage IDs (up to 9 images, 3 videos, and 3 audio clips). At 768p, the same tier's model refines the enlargement; because the reference-conditioned pass runs at a quarter of the finished pixel count, a 32 GB card such as the RTX 5090 can render Reference-to-Video with a 15-second reference video. 2K Reference-to-Video is the headline of Comfy worker 1.0.220: the 1080p and 2K classes render the first pass on the tier's model at the 768p canvas, then refine the enlargement to the finished size with FastH3 — so a 32 GB card renders every class and clip length, and the finished clip is exactly 2× the canvas you send (960×544 → 1920×1088, 1344×768 → 2688×1536). There is no Turbo two-stage Reference-to-Video ID.

Two-stage output is priced as the tier's own per-second rate plus a surcharge by canvas class — see Pricing. Worker requirements are under Worker hardware.

#Product availability

The Sogni API accepts every exact model ID in the table. Sogni Web exposes Standard, Balanced, LightX2V Turbo, and FastH3 Turbo for text-to-video, image-to-video, and first/last-frame creation; on FastH3 the Video Size picker offers 1-stage 768p or 2-stage 768p, 1080p, and 2K (Beta). Reference-to-Video exposes its dedicated Standard, Balanced, and LightX2V Turbo modes; FastH3 does not support Reference-to-Video. Every Reference-to-Video tier offers the same Video Size choices: 2-stage 768p, 1080p, and 2K (Beta) send the tier's two-stage ID with the matching canvas — 2K (2688×1536) is the top size for Reference-to-Video — and Standard and Balanced default to 2-stage 768p (as of 2026-09); on Turbo a 2-stage choice renders on the Standard ID at the Standard rate because Turbo has no two-stage build.

The Sogni Creative Agent Skill (3.39.0 and later) exposes friendly selectors for Standard, Balanced, LightX2V Turbo, FastH3 Turbo, and dedicated Ref2VA workflows. Its minimax-h3-turbo selectors continue to mean LightX2V Turbo; use minimax-h3-fasth3-turbo, or the explicit minimax-h3-fasth3-t2v-turbo, minimax-h3-fasth3-i2v-turbo, and minimax-h3-fasth3-flf2v-turbo, for FastH3. FastH3 has no r2v selector.

#Which speed should I choose?

Choose FastH3 Turbo for the quickest drafts, prompt iteration, timing tests, and other work where turnaround matters most. It renders up to twice as fast as LightX2V Turbo and up to 6× faster than standard 544/768p H3 on a 15-second clip in Sogni's qualified production configuration (shorter clips see a smaller multiplier), and it accepts the same custom LoRA catalog on workers with at least 32 GB of VRAM. Sogni Web uses FastH3 as the default Turbo engine, with a FastH3 switch under Turbo that returns to LightX2V. Choose LightX2V Turbo when you prefer its familiar four-step look. The model you select always determines the engine; attaching or removing a LoRA does not switch it. Actual wall-clock time varies with mode, input conditioning, clip length, resolution, worker hardware, and queue state.

Choose Balanced for an eight-step middle ground, or standard 20-step H3 when fine visual detail and audio polish matter more than iteration speed. At 544/768p-class output, LightX2V Turbo is 6 Spark per second and FastH3 is 4 Spark per second, against Standard H3's 16 — see Pricing.

On Sogni's primary NVIDIA RTX PRO 6000 H3 workers, a highest-quality 15-second standard H3 video typically takes about 15 minutes once rendering starts; FastH3 Turbo brings the same clip to roughly two minutes. Raw render time is the same on Unlimited and Unlimited Pro. Pro changes how much work can run and where it sits in the subscription queue, not the speed of one GPU render.

#What FastH3 is

FastH3 Turbo runs FastVideo FastH3 4-step Preview v1 VSA DataFree, the recommended FastH3 Preview v1 checkpoint from the FastVideo project at Hao AI Lab. It is a distillation of MiniMax H3 that generates synchronized video and audio in four transformer forwards instead of standard H3's 20 steps. VSA is FastVideo's video sparse attention (the checkpoint was trained with VSA-H3 at 90% sparsity), and DataFree means the distillation used data-free DMD2 rather than an external training set. Sogni runs Kijai's INT8 ConvRot conversion of the step-1300 checkpoint (about 21 GB), which is what fits the 24 GB-class worker pool. The checkpoint inherits the MiniMax H3 Community License. The Hao AI Lab FastH3 preview post covers the training details.

#Generation limits

  • Fixed 24 fps with native 32 kHz stereo audio.
  • 124–362 frames on the 124 + n×17 frame grid, approximately 5.17–15.08 seconds.
  • Size presets cover 544/768p-class output. The 480p lower tier is no longer offered as a preset or default (since 2026-09-10); it remains available only as a custom size. Dimensions use a 32 px grid, at least 480 px per side, up to 1344 px per side and no more than 1,032,192 rendered pixels.
  • Standard H3 uses fixed 20-step res_multistep / simple sampling. Balanced uses eight-step Euler/simple sampling at a video shift of 6. FastH3 Turbo uses its qualified four-step Euler/simple recipe, while LightX2V Turbo uses its qualified four-step ER-SDE/simple recipe. Guidance is fixed at 1.
  • Balanced accepts an optional shift from 4 to 12 in steps of 0.5 (default 6). Higher values spend more of the eight steps on overall layout and motion, lower values more on fine detail; if a Balanced render drifts, warps, or loses its subject, try 12 with the same seed. Values below 6 have not been reviewed. Other modes ignore shift.
  • Ref2VA reference sets: up to 9 images, 3 videos (24 fps, 2–15 s each and 15 s combined, each with an optional soundtrack), and 3 standalone audio clips — at most 12 reference files, with at least one visual reference (image or video). Audio alone is rejected.
  • Two-stage IDs: send the half canvas — a 384 px short edge for 768p delivery, 544 px for 1080p, the 768p canvas for 2K — with every edge at least 384 px, and receive exactly 2× in each dimension. Frame counts, audio, and reference limits are unchanged. Two-Stage Reference-to-Video currently admits only the 768p class (a canvas up to 672×384).

#Pricing

MiniMax H3 is priced per delivered second at the standard 1 Spark = $0.005 peg. FastH3 costs 4 Spark ($0.02) per output second at both resolution classes, so an exact 192-frame, 8-second FastH3 clip costs 32 Spark ($0.16). LightX2V Turbo remains 4 Spark ($0.02) per second at 480p and 6 Spark ($0.03) at 544/768p-class output.

Standard costs 10 Spark ($0.05) per second at 480p and 16 Spark ($0.08) at 544/768p-class output. Balanced costs 6 Spark ($0.03) per second at 480p and 10 Spark ($0.05) at 544/768p-class output.

  • Resolution tier and speed tier set the rate. Aspect ratio and mode do not change it within that tier.

  • Two-stage IDs add a per-second surcharge by canvas class on top of their tier's 544/768p-class rate. The class follows the canvas you send, and a two-stage job is never billed at the 480p rate because it always delivers 768p or larger.

    Two-stage tier 768p (canvas up to 672×384) 1080p (short edge up to 544 px) 2K
    FastH3 Two-Stage +0 +6 Spark ($0.03) per second +12 Spark ($0.06) per second
    Two-Stage Ref2VA, Standard +0 +6 Spark ($0.03) per second +12 Spark ($0.06) per second
    Two-Stage Ref2VA, Balanced +0 +6 Spark ($0.03) per second +16 Spark ($0.08) per second

    FastH3 prices Cinema's (1344×576) 768p canvas, 896×384, as 768p. On Two-Stage Ref2VA that canvas renders as a 1080p-class job and is priced as one.

  • Ref2VA video input is billed by duration and resolution. Each reference-video input second costs 10 Spark ($0.05) at 480p or 16 Spark ($0.08) at 544/768p-class output. Those full input rates apply to Standard, Balanced, and LightX2V Turbo; output-tier discounts apply only to generated output. FastH3 has no Ref2VA route. The two-stage Reference-to-Video IDs charge 26 Spark ($0.13) per reference-video second on 2K output and 16 Spark ($0.08) at the 768p and 1080p classes, and the same extra-image fee as their one-stage tier.

  • Standard and Balanced Ref2VA include five images. Each additional reference image costs 16 Spark ($0.08). Reference audio is free, and LightX2V Turbo has no excess-image surcharge.

  • Two keyframes are included. Each keyframe after the second adds output time at the clip's per-second rate, without the two-stage surcharge: 0.75 seconds on FastH3 (3 Spark, $0.015) and 0.3 seconds on every other tier. An 8-second FastH3 clip is 32 Spark with up to two keyframes, 35 with three, and 50 with eight.

  • You pay for the length that actually renders. H3 snaps to the 124 + n×17 frame grid at 24 fps, so billable durations land on grid values between 5.17 s and 15.08 s.

#480p output (custom size only)

Frames Duration Standard Balanced LightX2V Turbo FastH3
124 5.17 s 51.67 Spark / $0.2583 31 / $0.155 20.67 / $0.1033 20.67 / $0.1033
141 5.88 s 58.75 / $0.2938 35.25 / $0.1763 23.5 / $0.1175 23.5 / $0.1175
175 7.29 s 72.92 / $0.3646 43.75 / $0.2188 29.17 / $0.1458 29.17 / $0.1458
192 8.00 s 80 / $0.40 48 / $0.24 32 / $0.16 32 / $0.16
243 10.13 s 101.25 / $0.5063 60.75 / $0.3038 40.5 / $0.2025 40.5 / $0.2025
362 15.08 s 150.83 / $0.7542 90.5 / $0.4525 60.33 / $0.3017 60.33 / $0.3017

#544/768p-class output

Frames Duration Standard Balanced LightX2V Turbo FastH3
124 5.17 s 82.67 Spark / $0.4133 51.67 / $0.2583 31 / $0.155 20.67 / $0.1033
141 5.88 s 94 / $0.47 58.75 / $0.2938 35.25 / $0.1763 23.5 / $0.1175
175 7.29 s 116.67 / $0.5833 72.92 / $0.3646 43.75 / $0.2188 29.17 / $0.1458
192 8.00 s 128 / $0.64 80 / $0.40 48 / $0.24 32 / $0.16
243 10.13 s 162 / $0.81 101.25 / $0.5063 60.75 / $0.3038 40.5 / $0.2025
362 15.08 s 241.33 / $1.2067 150.83 / $0.7542 90.5 / $0.4525 60.33 / $0.3017

#Two-stage output (per output second)

Delivered size (canvas sent) FastH3 Two-Stage Two-Stage Ref2VA Standard Two-Stage Ref2VA Balanced
768p (384 px short edge, e.g. 672×384) 4 Spark / $0.02 16 / $0.08 10 / $0.05
1080p (544 px short edge, e.g. 960×544) 10 / $0.05 22 / $0.11 16 / $0.08
2K (768p canvas, e.g. 1344×768) 16 / $0.08 28 / $0.14 22 / $0.11

So an exact 8-second FastH3 2K clip costs 128 Spark ($0.64), the same as Standard at 768p, and an 8-second Two-Stage Ref2VA Standard clip delivered at 768p costs 128 Spark ($0.64) plus its reference-video seconds and any extra images — the same as one-stage Standard at 768p.

H3 runs on Sogni Supernet workers rather than an external vendor API, so it uses standard Spark access and does not require credit-card Premium Spark coverage. For a live quote at any frame count, use the job estimate calculator.

#Unlimited throughput and fair use

Both Sogni Unlimited plans cover standard H3 and H3 Turbo under credit-free fair use. The render time for one job is the same on both plans, but Unlimited Pro adds higher subscription queue priority, an additional standard H3 concurrency slot, and twice the H3 fair-use capacity of Unlimited.

Sogni intentionally does not publish a fixed number of 15-second H3 videos per 24 hours. The average paid subscriber does not encounter daily throttling, but concurrency is not the only control: paid plans also use adaptive daily fair-use throughput limits to keep the shared network responsive. Exact thresholds are unpublished to prevent gaming and can shift with GPU supply and demand. Within each plan tier, daily usage is a scheduling signal so accounts with lighter use are favored and more subscribers receive fast service with minimal queue time.

The platform is designed for an individual creator to keep a substantial queue moving throughout the day. Unlimited can queue up to 8 videos at once and Unlimited Pro up to 24; the API and Creative Agent Skill make it practical to automate a personal production queue. Fair use still applies: unattended continuous 24/7 infrastructure and multi-user production workloads require pay-as-you-go Premium Spark or an Enterprise arrangement.

#JavaScript SDK example

Use the exact FastH3 workflow ID. Sampling settings are owned by the workflow, so omit steps, guidance, sampler, and scheduler overrides:

const project = await sogni.projects.create({
  type: 'video',
  network: 'fast',
  modelId: 'minimax-h3-fastvideo-int8_t2v_turbo',
  positivePrompt: `integrated_multimodal_description: [Shot 1] A cinematic tracking shot moves through a rainy night market as vendors serve customers beneath glowing awnings.

overall_soundscape: Steady rain, footsteps through shallow puddles, and layered crowd ambience.

non_diegetic_music: A restrained electronic pulse builds beneath the scene.`,
  duration: 8,
  generateAudio: true,
  tokenType: 'spark'
});

const urls = await project.waitForCompletion();

For first-frame FastH3 animation, use minimax-h3-fastvideo-int8_i2v_turbo with referenceImage. For a FastH3 transition between two anchors, use minimax-h3-fastvideo-int8_flf2v_turbo with referenceImage and referenceImageEnd. Use the corresponding minimax-h3-fl2va-fp8_*_turbo ID when you want LightX2V Turbo. Include referenceVideoDurations in matching order for an accurate Ref2VA preflight quote; final billing uses the accepted reference-video duration. For two-stage Reference-to-Video, send minimax-h3-ref2va-fp8_r2v_2stage or minimax-h3-ref2va-fp8_r2v_balanced_2stage with the half canvas — width: 672, height: 384 delivers 1344×768 — and the same references and referenceVideoDurations as the one-stage ID; a larger canvas is refused until the 1080p and 2K classes open.

#Intermediate keyframes

Keyframes pin up to eight images at chosen moments between the first and last frame of a clip. They work on:

  • Image to video and first/last frame: the i2v and flf2v IDs of every tier — Standard, Balanced, LightX2V Turbo, and FastH3 — and the FastH3 Two-Stage i2v and flf2v IDs.
  • FastH3 Sound to Video: minimax-h3-fastvideo-int8_ia2v_turbo (image + audio), minimax-h3-fastvideo-int8_flfa2v_turbo (first/last frame + audio), minimax-h3-fastvideo-int8_a2v_turbo (audio only), and their _2stage IDs.
  • Reference to video: every Reference-to-Video ID, including the two-stage IDs.

Text-to-video IDs do not take keyframes. The video lands on each image at its time. A keyframe with a new angle, place, or light starts a new shot, so you can storyboard a clip, and on Sound to Video the speech stays in sync through the cuts. See examples of each mode in the launch post, Sogni Launches Multi-Keyframe MiniMax Generations.

In Sogni Web, open MiniMax H3 Image to Video, Sound to Video, or Reference to Video and use the Keyframes block under Frames (under Reference Media on Reference to Video; on a phone, in the Frames, Audio & frames, or References sheet). Add keyframe opens the image picker, crops the image to the video size, and places the keyframe in the middle of the widest gap. Set its time in the At field or with the slider, in 0.1-second steps, and remove it with its trash button. A keyframe that falls outside the clip, lands on the same frame as another, or has no image says so on its card and holds Imagine until you fix it; the app never moves or drops it for you. The Screenwriter sees your keyframes and writes each one into the prompt at its time, and the price estimate includes them.

In Sogni Chat, attach the images and name the moments, for example "make a MiniMax H3 video that starts on the first image and lands on the second at 3 seconds and the third at 6." The Creative Agent Skill takes --keyframe <image>@<seconds>, and the hosted API tools take a keyframes argument.

With the SDK, keyframes add to a mode's own inputs rather than replacing them: the first and last frames stay referenceImage and referenceImageEnd, Sound to Video keeps its audio, and Reference-to-Video keeps its references. Each keyframe is an image plus the frame it lands on, counted from 0 at 24 fps:

const project = await sogni.projects.create({
  type: 'video',
  network: 'fast',
  modelId: 'minimax-h3-fastvideo-int8_flf2v_turbo',
  referenceImage: firstFrame,
  referenceImageEnd: lastFrame,
  keyframes: [{ image: facingCamera, frameIndex: Math.round(3 * 24) }],
  positivePrompt: '<alignment line for Picture 1, keyframe Picture 3 and Picture 2, plus three-field prompt that names <Picture 3> at 3 seconds>',
  frames: 192,
  generateAudio: true,
  tokenType: 'spark'
});

Keyframes need SDK 5.58.0 or later, in JavaScript or Python; the Python SDK takes each keyframe as {"image": ..., "frame_index": ...}.

  • Each frameIndex must be a whole number from 1 to frames - 2 on every mode, and no two keyframes may share a frame. Frame 0 and the last frame are never keyframes, even on audio-only and Reference-to-Video jobs, which have no first or last frame image. Pass frames from the 124 + n×17 grid so the range is exact; a duration is snapped to the grid first (6 seconds becomes 141 frames).
  • Name each keyframe <Picture N> in the prompt, numbered in time order after the job's own pictures: after the first or last frame, after both, or after the Reference-to-Video reference images (audio-only jobs start at <Picture 1>). Keyframes still don't count as references.
  • Except on Reference-to-Video, open the prompt with one alignment line that lists every picture at its time: How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 2) aligns with the 2.88-second mark of the target video. On Reference-to-Video, add <Picture N> is the keyframe of [Shot M], showing ... to subject_definitions, keyframe completion to the summary tasks, and <Picture N> ([Shot M] keyframe): fully_preserved - ... to retention_analysis.
  • In the shot where a keyframe lands, write "the shot's keyframe corresponds to <Picture N>" and describe what the video shows there. H3's text encoder never sees keyframe images, so the words carry what each one looks like.
  • On Sound to Video the audio still drives the performance; keyframes pin how the picture looks at their times.
  • When a keyframe changes the framing, layout, or lighting, write a hard cut to a new shot at that keyframe's time ([Shot N] At MM:SS.mmm, the camera cuts to ..., whose keyframe corresponds to <Picture N>.) and describe the shot the way the keyframe shows it. Two keyframes that differ like that inside one continuous shot tend to cross-fade into each other.
  • If no worker that can pin keyframes is online for the model, the request is refused before any Spark is reserved. Try again shortly, or send the job without keyframes.
  • Pricing: your first two keyframes are included. Each keyframe after the second adds output time at your job's own per-second rate, without the two-stage surcharge: 0.75 seconds on FastH3 (3 Spark) and 0.3 seconds on every other tier (for example 4.8 Spark on Standard at 768p). An 8-second FastH3 clip with 8 keyframes costs 32 + 18 = 50 Spark. Quote it with keyframeCount in estimateVideoCost.

#Creative Agent Skill examples

The Creative Agent Skill exposes friendly selectors for every current workflow. The generic minimax-h3 and minimax-h3-turbo selectors infer text-to-video, image-to-video, or first/last-frame mode from the supplied frames; standard Ref2VA remains explicit.

# Standard quality-first text-to-video
sogni-agent --video -m minimax-h3 --duration 10 "<three-field H3 prompt>"

# 4-step Turbo image-to-video
sogni-agent --video -m minimax-h3-i2v-turbo --ref first.png --duration 8 "<I2V preamble plus three-field H3 prompt>"

# 4-step Turbo first-and-last-frame video
sogni-agent --video -m minimax-h3-flf2v-turbo --ref first.png --ref-end last.png --duration 8 "<FLF2V preamble plus three-field H3 prompt>"

# FastH3 Turbo (FastVideo FastH3 4-step Preview v1); --ref and --ref-end infer I2VA, L2VA, or FL2VA
sogni-agent --video -m minimax-h3-fasth3-turbo --duration 8 "<three-field H3 prompt>"
sogni-agent --video -m minimax-h3-fasth3-i2v-turbo --ref first.png --duration 8 "<I2V preamble plus three-field H3 prompt>"

# FastH3 first-and-last-frame video that also lands on two keyframes (repeat --keyframe, up to 8)
sogni-agent --video -m minimax-h3-fasth3-flf2v-turbo --ref first.png --ref-end last.png --keyframe [email protected] --keyframe wave.png@6 --duration 8 "<FLF2V preamble plus three-field H3 prompt that describes each keyframe at its time>"

# Standard Ref2VA with labelled image, video, and audio references
sogni-agent --video -m minimax-h3-r2v --ref identity.png --ref-video motion.mp4 --ref-audio voice.m4a "<six-field Ref2VA prompt>"

Use minimax-h3-turbo when you want the Skill to infer LightX2V Turbo t2v/i2v/flf2v automatically, and minimax-h3-fasth3-turbo for the same inference on FastH3. Use minimax-h3-r2v-turbo for Turbo Ref2VA; FastH3 rejects --workflow r2v.

--keyframe <image>@<seconds> pins an image, a local path or a URL, at that time strictly inside the clip, in addition to --ref and --ref-end. It works on every H3 image-to-video, first/last-frame, FastH3 Sound to Video, and Reference-to-Video selector, and never on text-to-video; a time that does not fit the clip is refused with the fix rather than moved. In the Skill's local MCP server (Claude Desktop, Codex, and other MCP clients), the generate_video tool takes the same as keyframes: [{ "image": ..., "at_seconds": ... }].

#Worker hardware

FastH3 Turbo supports custom H3 LoRAs on workers with at least 32 GB of VRAM without switching engines. Base text-to-video, image-to-video, and first/last-frame jobs can additionally run on eligible CUDA 13 workers with at least 23 GB of usable VRAM, including the NVIDIA GeForce RTX 4090 and RTX 3090.

LightX2V Turbo, Standard, Balanced, and every Reference-to-Video mode require at least 32 GB of VRAM. Reference-to-Video always uses its dedicated checkpoint, and CUDA 12 workers do not advertise FastH3 because that track lacks the required optimized INT8 kernels. Two-Stage Reference-to-Video needs Comfy Worker 1.0.218 or newer; its 768p class runs on the 32 GB class even with a 15-second reference video, whereas one-stage Standard Reference-to-Video with a reference video routes to cards with more than 40 GB of VRAM (more than 48 GB from 243 frames). FastH3 Two-Stage needs 1.0.217 or newer. Sogni Comfy Worker applies Blackwell-specific runtime tuning to the RTX 5090 and RTX PRO 6000 Blackwell, but H3 is not restricted to Blackwell. See the Fast Worker requirements.

MiniMax's hosted H3-Context-IR model stage and 2K regeneration features are not part of the current release: MiniMax has not yet published them as open weights, so current Sogni output should not be described as native 2K; the _2stage IDs deliver 2× the requested canvas through Sogni's own latent enlargement and refinement, which is upscaled delivery rather than native 2K. Sogni plans to integrate that stage if MiniMax publishes the weights. That update will be included automatically for existing subscribers. Unless a release is specifically marked otherwise, newly supported open-weight models and related updates are automatically included in both existing and new subscriptions. The separate H3 Ref2VA reference-to-video checkpoint is already included and runs on the same 32 GB worker class. The hosted model stage is distinct from the ordered Context-IR prompt document used by the released workflows.

#Prompting

Base and Turbo text-to-video, image-to-video, and first/last-frame workflows use MiniMax's ordered three-field prompt contract:

integrated_multimodal_description: [Shot 1] ...

overall_soundscape: ...

non_diegetic_music: ...

Image-to-video and first/last-frame prompts prepend their mode-specific alignment instruction, followed by a blank line before the three fields. Use [Shot N] notation, keep stable (S1), (S2), and subsequent speaker IDs across shots, and write exact dialogue as <d>[Language] words</d>. Put negative direction inside the structured prompt because H3 has no negative-prompt field.

Ref2VA uses a separate six-field contract in this order: subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, and non_diegetic_music. Its references are loose references rather than locked frames. Address them with H3's 1-based per-type labels — <Picture 1>, <Video 1>, and <Audio 1> — and give each one an explicit job such as identity, wardrobe, environment, camera movement, or voice. State which reference wins when two disagree. A reference video's own soundtrack takes the next <Audio n> number before standalone audio clips.

Official and implementation sources: MiniMax H3 model card, MiniMax H3 license, MiniMax H3 API pricing, LightX2V MiniMax H3 Turbo model card, LightX2V Turbo implementation, FastVideo FastH3 4-step Preview v1 VSA DataFree model card, Hao AI Lab FastH3 preview post, FastVideo implementation, Kijai INT8 ConvRot conversion, and ComfyUI day-zero H3 support.

Last updated 2026-09-27