A Practical Guide to MiniMax H3 Video Generation with the Ace Data Cloud API

Video generation becomes much easier to build around when you stop treating every workflow as a separate endpoint and instead model the request as structured multi-modal input.
This guide walks through the MiniMax H3 video generation API exposed through Ace Data Cloud. The goal is practical: send text, optional reference media, and a few production parameters to create a video task that can either complete synchronously or be polled asynchronously.
What you can do
The POST /minimax/videos endpoint is designed around a unified content array. That means the same API surface can cover several common builder workflows:
- Text-to-video: provide one non-empty
textitem and a fixedratio. - First-frame animation: combine text with an
image_urlitem whoseroleisfirst_frame. - Last-frame or start/end control: use
last_frame, or combinefirst_frameandlast_frameto guide the opening and closing frames. - Reference-driven generation: use
reference_image,reference_video, orreference_audioto steer subject consistency, camera motion, visual style, action, tone, or rhythm.
For product and campaign builders, this structure is useful because you can separate intent from constraints. The text describes the shot; reference images can constrain the character, product, clothing, scene, or style; reference video can constrain motion; reference audio can constrain tone or rhythm.
How it works
The base URL is https://api.acedata.cloud. The video creation endpoint is POST /minimax/videos, and authentication is sent with an HTTP header: authorization: Bearer {token}. Requests should use accept: application/json and content-type: application/json.
The required top-level fields are model, content, resolution, and duration. The model is fixed to MiniMax-H3. resolution accepts 768P or 2K. duration is an integer from 4 to 15 seconds. ratio is conditionally required: text-to-video requires a fixed ratio and cannot use adaptive, while image-guided workflows can usually omit it or use adaptive.
By default, if you do not pass async, the endpoint waits for generation to finish and returns the complete task. If you pass async: true, or provide a callback_url, the API returns a task_id and trace_id immediately. Without a callback, the documented polling pattern is to query the task API about every 10 seconds and read the final video from task.content.url when task.status becomes succeeded.
Build a text-to-video request
A minimal text-to-video request should include one non-empty text item. A helpful prompt structure is: subject, action, scene, shot, light, and sound. Keep the technical shape strict: the API uses content, not legacy fields such as prompt, image_urls, audio_urls, messages, or first_frame_image.
curl -X POST 'https://api.acedata.cloud/minimax/videos' \
-H 'accept: application/json' \
-H 'content-type: application/json' \
-H "authorization: Bearer ${ACEDATACLOUD_API_KEY}" \
-d '{
"model": "MiniMax-H3",
"content": [
{
"type": "text",
"text": "8 seconds, 16:9 product teaser. A compact smart speaker sits on a dark desk, soft blue light moves across its fabric surface, slow push-in camera, clean reflections, quiet studio mood."
}
],
"resolution": "768P",
"duration": 8,
"ratio": "16:9",
"async": true
}'
This is the safest first integration test because it exercises the core request shape without requiring hosted media assets.
Add a first frame when visual continuity matters
When you already have a product render, poster, illustration, or character image, use an image_url item with role: first_frame. In this mode, the input image determines the visual starting point, so the documentation recommends omitting ratio or using adaptive.
{
"model": "MiniMax-H3",
"content": [
{
"type": "text",
"text": "A product hero shot comes to life with a slow orbiting camera, subtle reflections, and a clean studio background."
},
{
"type": "image_url",
"image_url": { "url": "https://example.com/product-frame.png" },
"role": "first_frame"
}
],
"resolution": "768P",
"duration": 6,
"ratio": "adaptive",
"async": true
}
For reference-media workflows, keep in mind that first/last-frame inputs and reference materials are mutually exclusive. Once you use reference_image, reference_video, or reference_audio, do not also send first_frame or last_frame in the same request.
Plan around media limits
The API accepts media through publicly accessible HTTPS URLs, mm_file://{file_id}, or base64 data URIs. For large files, public HTTPS URLs are usually the cleanest option. Base64 increases payload size by about one third, and the full request body must stay within 64 MB.
The documented single-file limits are up to 30 MB for images, 50 MB for videos, and 15 MB for audio. Image dimensions must be between 256 and 5760 pixels on both sides, with aspect ratio from 0.4 to 2.5. Reference videos and audios must be 2 to 15 seconds per segment, with total duration not exceeding 15 seconds. In multimodal reference scenarios, images, videos, and audios together are limited to 12 files.
Where this fits in a real builder workflow
I would integrate this endpoint behind a small internal job abstraction: accept user intent, normalize it into content[], validate the media-role rules, submit with async: true, store task_id and trace_id, then poll until success or failure. That keeps the frontend simple and prevents long-running video jobs from tying up a user request.
The main implementation detail is not the HTTP call itself; it is keeping the request model honest. Use content everywhere, avoid mixing legacy fields, choose ratio based on workflow, and make reference roles explicit when consistency matters.
For the full parameter tables and task-query flow, read the MiniMax H3 Video Generation API integration guide.
Comments
Post a Comment