MiniMax H3 Max video model explained: fal's post-trained, speed-optimized version of MiniMax H3 with faster-than-real-time generation
ReviewsSep 3, 2026

What Is MiniMax H3 Max? Fal's Faster H3 Explained

MiniMax H3 Max is fal's post-trained, speed-optimized version of MiniMax H3. Here is how it hits 35x throughput, what it costs, and when to use it over stock H3.

MiniMax H3 Max is a post-trained version of the open-weight MiniMax H3 video model, developed by fal Research and served on fal's speed-optimized inference stack. Released on August 26, 2026, it generates a 5-second video in roughly 3 seconds of wall time, which fal reports as about 35x the throughput of the official MiniMax H3 endpoint, while ranking first in fal's human preference evaluations for overall quality, prompt understanding, and aesthetics.

That combination is the whole story: until now, fast video models and high-quality video models sat on opposite ends of a curve. H3 Max is an attempt to move the curve. This guide explains what the model is, how fal made it faster, what it costs, where it beats stock H3, and where stock H3 still wins.

TL;DR: MiniMax H3 Max at a Glance

  • What it is: fal Research's post-trained derivative of open-weight MiniMax H3, co-optimized with fal's inference engine
  • Speed: 5-second clip in about 3 seconds, roughly 35x the throughput of the official H3 endpoint
  • Resolutions: 480p and 768p native (768p default), no 2K or 4K
  • Duration and format: 5 to 15 seconds, 24 FPS, six aspect ratios on text-to-video, native synchronized audio
  • Endpoints: text-to-video, image-to-video, and reference-to-video on fal
  • Price on fal: $0.05/second at 480p, $0.08/second at 768p, plus a free daily tier
  • Best for: iteration-heavy work, interactive tools, batch runs, and beat-sequenced prompts
  • Skip it when: you need 2K/4K delivery, video editing, or voice transfer, which only stock H3 offers

What Is MiniMax H3 Max?

H3 Max starts from MiniMax H3, MiniMax's open-weights general-purpose multimodal model, but it is more accurate to call it fal's H3 derivative than a new MiniMax foundation release. fal Research took the open-weight base, post-trained it with substantial new data focused on prompt adherence and visual quality, then built a serving system specifically optimized around the result.

The cleanest one-line definition:

H3 Max = post-trained MiniMax H3 + fal's optimized inference system

Both halves matter. A post-trained model on a generic serving stack would be a better model at normal speed. A stock model on fal's stack would be a normal model at higher speed. H3 Max is the two combined, which is why fal can claim top-tier quality and faster-than-real-time generation at the same time.

Outside fal, the model has also surfaced on partner platforms: MiniMax's own Design app generates with H3 Max in seconds, and Krea lists it as an available video model.

Who Made It: MiniMax vs fal

The division of labor is unusual enough to be worth spelling out:

Role
MiniMaxBuilt the H3 foundation model, released it with open weights, and operates the official endpoint
fal ResearchPost-trained the weights into H3 Max, targeting prompt adherence and aesthetics
fal inference teamCo-designed the serving engine, squeezing faster-than-real-time throughput out of the model

fal announced H3 Max on August 26, 2026, with a press release following on September 1. The launch pricing offered 50% off for the first week.

How fal Made H3 Faster

H3 Max's speed is best understood as model-system co-optimization, not one acceleration trick. fal worked both sides at once, which is rare because frontier model research and deep kernel-level inference work seldom sit in the same team.

Post-training: better, not just faster

Traditional step-distillation aims to match the base model at lower cost. fal says it aimed higher: a model that beats the original on the qualities people actually notice, while running much faster.

The post-training data targeted prompt adherence (following what you asked for, in the order you asked for it) and visual quality. Much of the compute went into verifiable reinforcement learning tasks run through fal's in-house framework. Along the way, fal evaluated each checkpoint with head-to-head human preference studies across three dimensions scored separately: overall quality, prompt understanding, and aesthetics. Scoring them independently caught regressions that a single aggregate score would have hidden.

Crucially, the base model's core capabilities survived: unified multimodal context and natively synchronized audio and video both carry over to H3 Max.

Inference co-design: throughput with a quality gate

There are cheap ways to make any video model faster: reduce precision, cut sampling steps, approximate expensive operations. Many produce great benchmark numbers while quietly degrading output. fal's rule was that an optimization only survived if the model kept its position in internal quality evaluations. The result is an inference engine built around H3 Max, rather than a generic serving stack that happens to run H3 Max weights.

Training and serving both ran on NVIDIA GB200 NVL72 systems, which fal credits with up to twice the per-chip performance of its previous-generation accelerators.

The quality numbers

In fal's evaluations, H3 Max was benchmarked against twelve leading video models, including the official MiniMax H3 endpoint, Gemini Omni Flash, Wan 3.0, Seedance 2.5, Kling 3, and Veo 3.1. Comparisons were aggregated with Bayesian Elo ratings, and H3 Max ranked first on all three dimensions, winning the majority of head-to-head matchups against every model tested, including the original H3.

Independent benchmarks from Artificial Analysis and Design Arena also place H3 Max at number one among video models. Treat all of this as strong evidence rather than gospel: preference evaluations measure what panels prefer, and your footage may disagree. But the direction is consistent across fal's own tests and two third-party leaderboards.

What Does "35x Faster" Actually Mean?

Fal's claim: a 5-second video in about 3 seconds of wall time, roughly 35x the throughput of the official MiniMax H3 endpoint, and on average 15x faster than anything with comparable quality. A real API response on fal confirms the scale: the timings.inference field reports roughly 2.5 seconds of backend render time for a 5-second 768p clip.

The headline number still deserves a careful read, because it is easy to over-claim. What it is not:

  • Not a local speedup. The comparison covers the post-trained model plus fal's hosted stack. No official H3 Max checkpoint has been released, so there is nothing to download, and a consumer GPU running open-weight H3 will not magically gain 35x.
  • Not a latency guarantee. Throughput measures how much work the system processes over time. Independent reports have also captured slower outliers on the API (a 4-second 768p clip taking around 96 seconds, likely queueing or load related). Faster-than-real-time is a demonstrated capability, not a universal promise.

The practical takeaway for creators: on fal's hosted endpoint, H3 Max's speed advantage is real and repeatable, and fast enough to change how you work.

MiniMax H3 Max Specs

MiniMax H3 Max generates 5 to 15 second clips at 24 FPS in 480p or 768p, across six aspect ratios on text-to-video, with native synchronized audio, lip-sync, and end-frame control.

SpecMiniMax H3 Max
Resolutions480p, 768p (both native, 768p default)
Duration5 to 15 seconds
Frame rate24 FPS
Aspect ratios (T2V)6 options in the schema enum
Native audioYes, composed with the picture
Lip-syncYes
End-frame controlYes, via end_image_url
Reference inputsUp to 12 files (images, video, audio)
Prompt expansion modesbalanced, quality
EndpointsText-to-video, image-to-video, reference-to-video
Response timingstimings.inference reports backend render time
WeightsNot released (hosted only)

Since launch, fal has also added H3 Max Turbo, a further-tuned variant with stronger prompt adherence, available as minimax/h3-max-turbo/text-to-video.

MiniMax H3 vs H3 Max: Full Comparison

fal hosts both models side by side, and the differences that matter in an actual API call are resolution ceiling, feature coverage, and price.

MiniMax H3MiniMax H3 Max
Built byMiniMaxfal Research (post-trained from H3)
Resolutions480p, 768p, 2K, 4K480p, 768p
Default resolution2K768p
Native vs upscaled480p/768p native, 2K/4K upscale a 768p base480p/768p both native
Prompt ceilingUp to 7,000 charactersLower (undocumented)
Prompt expansionfast, balanced, qualitybalanced, quality
Video editingYesNo
Voice transferYesNo
Free daily generationsNo10 per day (5 on the tool page, 5 in the sandbox)
Best for2K/4K delivery, editing, reference-heavy workFast iteration, interactive tools, batch runs

Choose stock H3 when you deliver above 768p, need documented editing operations (product swaps, line replacement, relighting), want voice transfer, or run very long multi-shot prompts.

Choose H3 Max when render time governs the work: somebody is watching a spinner, you are burning nine takes to keep one, or you need per-request render timing for an instrumented pipeline.

MiniMax H3 Max Pricing

All figures are fal's published playground rates as of August 31, 2026, billed per second of output.

OutputMiniMax H3MiniMax H3 Max
480p, per second$0.05$0.05
768p, per second$0.06$0.08
2K, per second$0.13Not offered
4K, per second$0.16Not offered
5 seconds at 768p$0.30$0.40
15 seconds at 768p$0.90$1.20

Reference inputs bill differently on each model, and the schedules cross:

Reference billingMiniMax H3MiniMax H3 Max
Unit billedWhole imagesPooled tokens across images, video, audio
Free allowanceFirst 5 imagesFirst 4,096 tokens
Past the allowance$0.08 per image$0.02 per 1,000 tokens
One 1024x1024 still$0.08$0.02
One 2048x2048 still$0.08$0.08
One 3072x3072 still$0.08$0.18

The crossover sits at 2048x2048: many modest stills come out cheaper on H3 Max, while high-resolution stills cost less on stock H3. Note that on H3 Max, reference clips are priced off the resolution you generate at, not the resolution of the clip you supply.

The free tier is the quiet differentiator: H3 Max gives you five 5-second generations per day from the tool page with no account, and five more per day at up to 15 seconds each in the fal sandbox once signed in. Stock H3 bills from the first second.

How to Use H3 Max

Three doors, all on fal:

  1. Playground/tool page: point and click at fal.ai/minimax-h3-max, free daily generations included
  2. API: one SDK (@fal-ai/client) with auth, queueing, and webhooks identical across fal's video endpoints
  3. fal Agent: for agentic pipelines

A minimal text-to-video call:

import { fal } from "@fal-ai/client";

const result = await fal.subscribe("minimax/h3-max/text-to-video", {
  input: {
    prompt:
      "Slow dolly forward down a rain-slick night-market alley, steam pouring off a noodle cart under a buzzing red neon sign. Sound: wok sizzle, rain ticking on tin awnings, neon hum.",
    resolution: "768P",
    aspect_ratio: "16:9",
    duration: 5,
    prompt_expansion_mode: "balanced",
  },
});

console.log(result.data.video.url);
console.log(result.data.timings.inference); // backend render time

Swapping to stock H3 is a one-string change of the model ID.

Limitations You Should Know

  • Resolution ceiling at 768p. If the delivery spec says 2K or 4K, H3 Max is out; that territory belongs to stock H3.
  • No released weights. H3 Max is hosted-only. Teams that need local deployment, privacy control, or a custom inference stack should look at open acceleration projects like FastH3 (which reports 15 seconds of 768p in about 13 seconds on a single GPU, with a claimed 14x acceleration) or run stock open-weight H3 themselves.
  • Faster-than-real-time is a best case. API outliers under load can run much slower than the headline number.
  • Quality still varies. Community reports flag occasional anatomy issues and robotic human movement. Anatomy-critical hero shots deserve comparison against other frontier models before you commit.

Why Faster Video Matters for 3D Workflows

At Trify3D, the interesting number is not seconds per clip, it is iterations per hour. When a render takes minutes, you explore three ideas. When it takes three seconds, you explore thirty, and the bottleneck moves from waiting to deciding.

That reframes where H3 Max fits a 3D pipeline:

  • Animate your 3D renders. Export a hero frame from your image to 3D or text to 3D output, then use H3 Max's image-to-video endpoint to turn a static turntable into a cinematic product shot with synchronized sound, iterating on camera moves in seconds instead of re-rendering.
  • Concept exploration before modeling. Burn through 15 to 20 motion concepts in the time one render used to take, lock the direction, then build the asset.
  • Reference packs for consistency. The reference-to-video endpoint accepts up to 12 images, video, and audio files, which pairs naturally with character sheets and product shots you already own.

The pattern that is emerging across AI media is the same one we build around: fast, cheap drafts to find the idea, then higher-resolution passes to finish it. H3 Max is currently the strongest draft engine in video.

FAQ

What is MiniMax H3 Max?

MiniMax H3 Max is a post-trained variant of the open-weight MiniMax H3 video model, developed by fal Research and served on fal's optimized inference stack. It generates 5 to 15 second clips with native synchronized audio at 480p or 768p, producing a 5-second video in roughly 3 seconds.

Is MiniMax H3 Max free?

Not entirely, but it has a free daily tier: five 5-second generations per day from the fal tool page without an account, plus five more per day (up to 15 seconds each) in the fal sandbox when signed in. Paid generation bills per second of output.

What resolutions does MiniMax H3 Max support?

480p and 768p natively, with 768p as the default. There is no 2K or 4K option. Stock MiniMax H3 covers 480p through 4K, where 2K and 4K are upscales built on a 768p base render.

Is MiniMax H3 Max really 35x faster than H3?

Fal reports roughly 35x the throughput of the official MiniMax H3 endpoint, with a 5-second video in about 3 seconds. That figure compares fal's hosted post-trained model plus its inference stack against the official hosted endpoint. It is not a local consumer-GPU speedup, and no downloadable H3 Max checkpoint exists.

Can you run MiniMax H3 Max locally?

No official H3 Max weights have been released. You can run the open-weight base MiniMax H3 locally, but speed will depend heavily on your VRAM, precision, attention backend, and step count. For an open acceleration route, projects like FastH3 target faster local H3 inference.

Does MiniMax H3 Max generate audio?

Yes. H3 Max preserves the base model's natively synchronized audio and video, including lip-sync for spoken dialogue. Audio is composed with the picture in the same pass.

How much does MiniMax H3 Max cost?

On fal: $0.05 per second at 480p and $0.08 per second at 768p (a 5-second 768p clip is $0.40, 15 seconds is $1.20). Stock H3 costs $0.05/$0.06/$0.13/$0.16 per second at 480p/768p/2K/4K. These are fal's published playground rates as of late August 2026.

Run it yourself in Trify3D

Keep reading from this topic