API referenceGeneration
Media utilities, 3D and speech
Submit these media tools as durable workflow steps
to POST /v1/creative-agent/workflows, or use them through hosted chat. The
synchronous POST /v1/creative-agent/tools/execute endpoint runs only the
composition and planning tools and returns 400 for media tools.
#Supply source media
List source files in the request's media_references array, then select them
by index. Negative indices select request media of that type, in order: -1 is
the first image (or video, or audio file), -2 the second. Non-negative
indices select results generated earlier in the same workflow. Use dependsOn
to pass one step's output to the next. URLs must be HTTP(S); inline data:
URIs are rejected. Upload local files first with the
media URL endpoints.
#Estimate before starting
The examples send confirm_cost: false, which checks the plan and estimates it
without saving or starting anything. The API responds 400 with
details.estimatedCapacity (a per-step estimate in capacity units, not Spark)
and details.confirmationRequired: true. Resubmit with confirm_cost: true,
or without the field, to start, then read the workflow or subscribe to its
event stream for results. The exact charge is computed per job when each step
runs.
Every tool on this page also respects the request-level safe_content_filter
setting (default true); there is no per-step filter argument.
#Remove a background
{
"media_references": [
{ "kind": "image", "url": "https://your-host.example/original.png" }
],
"input": {
"steps": [
{
"id": "remove_background",
"toolName": "remove_background",
"arguments": { "sourceImageIndex": -1, "applyMask": true }
}
]
},
"token_type": "spark",
"confirm_cost": false
}
BiRefNet preserves the original foreground and returns a transparent PNG. Set
applyMask: false to return the soft foreground mask instead of the cutout. No
prompt is required. The SDK project equivalent uses
modelId: 'birefnet_image_background_removal_fp16', startingImage, and
applyMask.
A transparent cutout is not a valid first frame, last frame, or reference for Wan 3 video models, which reject PNGs with a transparency channel. Flatten it first.
#Reconstruct a textured 3D model
{
"media_references": [
{ "kind": "image", "url": "https://your-host.example/front.png" },
{ "kind": "image", "url": "https://your-host.example/left.png" },
{ "kind": "image", "url": "https://your-host.example/back.png" },
{ "kind": "image", "url": "https://your-host.example/right.png" }
],
"input": {
"steps": [
{
"id": "image_to_3d",
"toolName": "image_to_3d",
"arguments": {
"sourceImageIndex": -1,
"leftViewImageIndex": -2,
"backViewImageIndex": -3,
"rightViewImageIndex": -4,
"meshTargetFaces": 50000,
"textureSize": 2048,
"shapeResolution": 1024
}
}
]
},
"token_type": "spark",
"confirm_cost": false
}
The front image (sourceImageIndex) is required. Any subset of the optional
orbit views selects Pixal3D multi-view reconstruction; front-only uses
single-view reconstruction. Views describe the subject's own sides: left means
its left side faces camera (subject faces screen-left), right means
screen-right. Use original stills of the same subject at consistent height and
framing. The orbit views also accept URLs directly as left_view_image_url,
back_view_image_url, and right_view_image_url.
| Control | Allowed values | Default |
|---|---|---|
meshTargetFaces |
Integer 5,000–700,000 | 700,000 |
textureSize |
Integer 1,024–4,096 | 4,096 |
normalMapSize |
Integer 512–2,048 | 2,048 |
ambientOcclusionSize |
Integer 256–1,024 | 1,024 |
shapeResolution |
1,024 or 1,536 | 1,024 |
The result is a binary GLB (model/gltf-binary). Workflow artifacts report
kind: "model"; hosted chat tool results report mediaType: "model". Download
or display it as a 3D model rather than passing it to an image tool. The SDK
uses pixal3d_int8_i23d or pixal3d_multiview_int8_i23d, startingImage, and
named leftViewImage, backViewImage, and rightViewImage binary inputs.
Generic contextImages do not identify orbit views.
#Select an object
segment_image takes sourceImageIndex plus text, points, or boxes.
Point coordinates (x, y) and box coordinates (x0, y0, x1, y1) are
normalized to the original image's 0–1 range. Points carry a positive or
negative label; boxes default to positive and can carry an exclusion label
with a text prompt.
- Provide
text, at least one positive point, or a box. Negative points alone are rejected. textis 1–240 characters,pointsup to 32, andboxesup to 16, withx0 < x1andy0 < y1.- Text can combine with boxes. Points can combine with at most one positive box and cannot combine with text.
thresholdis 0–1 (default 0.5) andmaxInstancesis 1–16.multimaskdefaults totruewith points and is rejected without them.
applyMask: true returns a transparent cutout; the default returns a PNG mask.
Output matches the source's dimensions. The SDK equivalent is
sam3_image_segment_bf16 with startingImage and sam3Prompt.
#Upscale a video
upscale_video runs FlashVSR on an existing video without a prompt. It keeps
every frame, the frame rate, and the audio.
| Argument | Values |
|---|---|
sourceVideoIndex |
Negative for request videos, non-negative for earlier results. Omitted means the latest video. |
targetResolution |
Short edge 1080 or 1440. Defaults to 1440 when the source short edge is at least 720 px, otherwise 1080. |
detailPreference |
stable (default) or sharper |
processingSpeed |
stable (default) or faster |
seed |
0–4294967295 (default 0); -1 picks a random seed |
The source needs a short edge of 540–768 px (at least 720 px for 1440), a constant frame rate from 1 to 60 fps, standard dynamic range, and at most 100 MB. Output is at most 2560×1440. Options do not change the price. If no worker that supports a non-default option is connected, the step reports that it is not available.
#Speech and reference-audio cloning
generate_speech reads prompt as the literal script (up to 4,096
characters). Select model: 'voice' for a studio speaker, 'clone' for a
supplied recording, or 'design' for a voice described in voiceDescription.
- Studio voices (
voice):serena(default),vivian,uncle_fu,ryan,aiden,ono_anna,sohee,eric, anddylan.voiceDescriptioncan add direction. - Clone uses
voiceSourceIndex(orvoice_source_url) and optionalvoiceTranscript(up to 1,024 characters). Supply the original recording; a clean 3–30 second clip of one speaker works best. Cloning conditions synthesis on that recording and does not create a persistent trained voice.voiceDescriptionis ignored in clone mode. - Design uses
voiceDescription(up to 512 characters) and takes no recording.
Other controls: language (auto by default, or english, chinese,
japanese, korean, german, french, russian, portuguese, spanish, or
italian), creativity (0.1–2, default 0.9), outputFormat (wav, mp3,
flac), seed, and numberOfVariations (1–16).
#MiniMax H3 Two-Stage and audio-driven video
Two-Stage. Use videoModel: 'minimax-h3-fasth3-t2v-turbo-2stage' with
generate_video, or the corresponding i2v or flf2v selector with
animate_photo. Set targetResolution to 720, 1080, or 1440 (default).
The worker renders the base clip, then enlarges and refines it, so this is
upscaled delivery rather than native 2K. Landscape output is 1344×768,
1920×1088, or 2688×1536; portrait is mirrored. Other values are rejected. Price
the selected mode and resolution together.
For reference-to-video, use videoModel: 'minimax-h3-r2v-2stage' (Standard)
or 'minimax-h3-r2v-balanced-2stage' (Balanced) with generate_video and the
same loose image, video and audio references as minimax-h3-r2v.
targetResolution works the same way: 720 delivers 1344×768, 1080
delivers 1920×1088, and 1440 (default) delivers 2688×1536. 1080p and 2K
render at the 768p canvas, then FastH3 refines the enlargement.
Audio-driven. sound_to_video accepts these FastH3 selectors, each also in
a -2stage form:
videoModel |
Inputs |
|---|---|
minimax-h3-fasth3-ia2v-turbo |
Audio plus a first frame (sourceImageIndex) |
minimax-h3-fasth3-flfa2v-turbo |
Audio plus first and last frames (sourceImageIndex, endImageIndex) |
minimax-h3-fasth3-a2v-turbo |
Audio only |
Select audio with audioSourceIndex (-1 is the first request audio). The
uploaded audio drives the clip and is kept as its soundtrack. Clips render
between about 5.2 and 15.1 seconds, snapped to 24 fps. These modes reject
generateAudio: false, negativePrompt, and loose reference arrays.
Keyframes. Every selector in this table also takes keyframes, as do the
two-stage i2v, flf2v, and Reference-to-Video selectors above, but never
text-to-video: up to 8 images the video lands on at exact moments inside the
clip, such as
"keyframes": [{ "imageIndex": -2, "atSeconds": 4 }]. On sound_to_video,
imageIndex counts request images the same way sourceImageIndex does (-1
is the first), and duration should be set so every keyframe falls inside the
clip. See H3 Keyframes
for the rules, pricing, and an example.
Every H3 mode, including the audio-driven ones, can use loras and
loraStrengths with catalog LoRAs and compatible
personal LoRAs.