MiniMax H3 Max and the Era of Agentic Video

Three MiniMax H3 Max video generations arranged as a hand-drawn storyboard

This took 1.3 seconds to generate.

It took longer than that to deliver the video file.

It took longer than that to enter the prompt!

The 1.3-second H3 Max generation

MiniMax H3 is a decent video model. By no means SotA, but good for simple stuff.

H3 Max is a further post-trained endpoint created by fal. It starts with MiniMax H3, then fal adds post-training aimed at prompt adherence and visual quality and tunes the inference stack around the model.

It's BLAZING fast!

The quality is about the same as H3, sometimes better and sometimes worse. Don't believe the weird leaderboards fal published alongside it that have it above Seedance, but that's not the point.

fal's cost-versus-quality chart
Design Arena leaderboard included in fal's announcement
Artificial Analysis leaderboard included in fal's announcement
fal's speed-versus-quality chart

Fal tested H3 Max against 12 other video models and scored overall preference, prompt understanding, and aesthetics with Bayesian Elo ratings. Fal says a five-second video takes about three seconds, giving it roughly 35 times the throughput of the official MiniMax H3 endpoint. My first result came back in 1.3 seconds.

The model keeps H3's synchronized audio and video. At launch, fal has both text-to-video and image-to-video endpoints.

The point is cost and speeeeeeeeeed.

Very low cost. It's 50% off right now, so that's 4 cents per second of output at 768p and 2.5 cents per second at 480p.

The cheapest one you can get costs 12.5 cents. This is five seconds at 480p: two guys playing basketball with a watermelon.

Five seconds at 480p: $0.125 at the launch price

Another result I got was basically just a profile picture animation. Five seconds at 768p, so that's a 20-cent generation.

Five seconds at 768p: $0.20 at the launch price

Price doubles September 1, but even then, it's still affordable.

But because of this speed, it unlocks new things.

Previously, video generation was something that you did yourself. You went back and forth with prompts to get the result you wanted.

Now you can let your agents do it, because calling it as a tool no longer takes two minutes to generate.

Your agent can enter the prompt, get the result, and review it, as long as it has video input, like Kimi K3 or GLM 5.3 Flash.

The era of agentic video is here.

Fal's H3 Max announcement