The 15-second clip above is straight text-to-video from H3 — an Apple-ad parody for a spring onion. No reference image, no titles added in post: the Chinese type you see was rendered by the model itself.
References aren't limited to images
Most video models only take images as reference. H3's reference mode accepts all three at once:
- Images — lock a person, a product, a look
- Video — hand it motion and camera movement to follow
- Audio — hand it rhythm and mood
On the canvas, wire image, video and audio nodes into an H3 video node and switch to multi-reference — they all go in together. At least one image or video is required (audio alone won't do).
Three ways to run it
Text-to-video — wire nothing, just write the prompt.
Image-to-video — wire one image as the first frame; add a second as the last frame to move from A to B.
Multimodal reference — the one above: images, video and audio mixed.
Specs
| Resolution | 2K (2560×1440, or 1440×2560 vertical) |
| Duration | 5–15 seconds, any whole second |
| Aspect | 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16, plus "adaptive" in image and reference modes |
| Sound | Picture and audio generated together, not dubbed afterwards |
| Price | 17.5 Xins/sec — 88 for 5s, 140 for 8s, 263 for 15s |
"Adaptive" is its default: give it a vertical image and you get a vertical video, not a forced 16:9 crop.
When to reach for it
Use the reference mode when you want motion to follow a clip or the shot to sit on a beat. If you only need to keep one character's face consistent, other models on the canvas do that too — no need to detour.
The cover image above is a frame H3 generated.