MiniMax H3 Architecture & Performance

Why Shape Is the Secret Turbo Button in MiniMax H3

In standard image generation, changing resolution alters fine pixel detail. In MiniMax H3 video, aspect ratio directly dictates sequence length. Because attention cost grows as the square of video tokens, switching from 16:9 to 1:1 delivers a massive 3× speedup and drastically lower VRAM without lowering model sharpness.

3.06× Faster
1:1 Square vs 16:9
Square ($768\times768$) costs only 32.7% of the attention compute needed by 16:9 ($1344\times768$).
O(S²) Attention
Quadratic Token Scaling
Every token compares against all others. Removing 43% of tokens eliminates 67% of attention matrix operations.
32 px Rigid Grid
VAE 16× · DiT 2×
Spatial VAE compression ($\div 16$) followed by DiT patch tiling ($\div 2$) enforces $16 \times 2 = 32\text{px}$ alignment.
1.03M Pixels
Area Cap ($768\times1344$)
Any aspect ratio wider than 7:4 scales down the short edge below 768px to preserve memory safety.

Live Resolution & Cost Simulator

Test any aspect ratio and duration. Watch how adapt_canvas() derives dimensions, calculates VAE/DiT token grids, and computes relative attention cost compared to standard 16:9.

Canvas Tier Presets (Fixed 768 Height Ramp)
Aspect Ratio Presets
Continuous Aspect Ratio 1.75 : 1 (16:9)
Requested Frames (24 fps) 124 frames (5.17s)
Snaps to 17n + 5 VAE temporal grid (5, 124, 209, 311, 345, 362).
1344 × 768 1,008 tokens/frame
Output Canvas
1344×768
Latent Grid
84×48
Tokens / Frame
1,008
Snapped Length
124 f (5.17s)
Latent Frames
36
Total Video Tokens
36,288
Attention Workload vs 16:9 Standard 1.00× (100%)
✓ In Trained Family ✓ Stable Fixed Point

Why Small Token Cuts Yield Huge Speedups ($O(S^2)$)

Diffusion Transformers do not calculate pixels linearly. Self-attention requires every video token to cross-attend with every other video token. When sequence length shrinks, the attention matrix shrinks in both dimensions simultaneously.

Baseline (100%)
16:9 Landscape
1344 × 768 px
1008 × 1008
Tokens / Frame 1,008
Attention Pair-Ops 1,016,064
Relative Compute 1.00× (100%)
Tier: fast
3:2 Classic
1152 × 768 px
864 × 864
Tokens / Frame 864 (-14.3%)
Attention Pair-Ops 746,496
Relative Compute 0.73× (73.5%)
Tier: draft
4:3 Standard
1024 × 768 px
768 × 768
Tokens / Frame 768 (-23.8%)
Attention Pair-Ops 589,824
Relative Compute 0.58× (58.1%)
★ 3× Speedup
1:1 Square
768 × 768 px
576 × 576
Tokens / Frame 576 (-42.9%)
Attention Pair-Ops 331,776
Relative Compute 0.33× (32.7%)

How Pixels Become Tokens: The 32px Rule

Why can't you pick any arbitrary width like 1300 or 1000? Two compression stages multiply together, creating an architectural requirement of 32-pixel divisibility.

1
Pixel Canvas
1344 × 768 px

Raw RGB video frames generated or fed into the model. In MiniMax H3, the standard landscape canvas starts with a 768px short edge.

2
Video VAE (÷16)
84 × 48 Latent Units

The spatial Video Autoencoder compresses space by 16× in both width and height ($1344/16 = 84$, $768/16 = 48$).

3
DiT Patchify (÷2)
42 × 24 = 1,008 Tokens

The Diffusion Transformer chops the latent grid into 2×2 spatial patches (patch_size=(1, 2, 2)). $84/2 = 42$, $48/2 = 24$.

4
Attention (O(S²))
(Tokens / 1008)²

Attention computes correlations across all tokens. Because $16 \times 2 = 32$, any resolution not divisible by 32 leaves an odd latent axis that 2×2 patches cannot tile.

Interactive 32px Divisibility & Remainder Tester

Type any pixel dimension to see why 16-divisibility alone causes a latent tiling failure:

1. Raw Pixel Axis
768 px
→
2. VAE Latent (÷16)
48 units
→
3. DiT 2×2 Tile (÷2)
24 patches
Status
✓ Valid 32px Grid

The 3 Rules of adapt_canvas() & Gotchas

How MiniMax H3 turns an aspect ratio request into an exact pixel resolution, and the practical traps to avoid.

📐
The Area Cap Shrinks Ultrawide
The short edge is not always 768px. The total pixel area is capped at $768 \times 1344 = 1,032,192\text{px}$. If you request 21:9, maintaining 768px height would exceed the cap, so the algorithm scales the frame down to 1536×672.
🔄
Portrait Mirror Symmetry
Portrait resolutions cost the exact same as landscape. $768 \times 1344$ (9:16) produces $24 \times 42 = 1,008$ tokens, identical to $1344 \times 768$ (16:9). Token arithmetic is commutative: $(W/32) \times (H/32) = (H/32) \times (W/32)$.
⚠️
The "Rounding Tax" Anomaly
Rounding to 32px can push weird aspect ratios above 1.00x cost! Requesting 23:7 gives 1856×576 ($1,044$ tokens/frame = 1.073× attention cost). You pay 7.3% more VRAM and compute for zero extra pixel detail! Stick to standard ratios.
🚫
No "4K / 720p" Quality Dial
Typing 4K or 720p in adapt_canvas() returns the exact same 1344×768 resolution. H3 has no resolution quality slider—only an aspect ratio shape choice.
⚡

Important Caveat: Sol-Attn Token Floor (~60,000 Tokens)

Sparse attention algorithms like Sol-Attn require roughly 60,000 total tokens before sparse pruning yields measurable speedups. At 243 frames: fast (1152×768) has 62,208 tokens (works), while draft (1024×768) has 55,296 tokens (falls below the floor). Use draft to verify pipeline wiring, but benchmark Sol-Attn on fast or higher.

The 14 Standard Aspects & 48 Trained Landscape Canvases

Every resolution produced by adapt_canvas() within the legal $1/4 \dots 4$ aspect range.

Aspect Ratio Resolution (px) Latent Grid Tokens / Frame Attention Cost Canvas Tier Stability