In standard image generation, changing resolution alters fine pixel detail. In MiniMax H3 video, aspect ratio directly dictates sequence length. Because attention cost grows as the square of video tokens, switching from 16:9 to 1:1 delivers a massive 3× speedup and drastically lower VRAM without lowering model sharpness.
Test any aspect ratio and duration. Watch how adapt_canvas() derives dimensions,
calculates VAE/DiT token grids, and computes relative attention cost compared to standard 16:9.
17n + 5 VAE temporal grid (5, 124, 209, 311, 345, 362).
Diffusion Transformers do not calculate pixels linearly. Self-attention requires every video token to cross-attend with every other video token. When sequence length shrinks, the attention matrix shrinks in both dimensions simultaneously.
Why can't you pick any arbitrary width like 1300 or 1000? Two compression stages multiply together, creating an architectural requirement of 32-pixel divisibility.
Raw RGB video frames generated or fed into the model. In MiniMax H3, the standard landscape canvas starts with a 768px short edge.
The spatial Video Autoencoder compresses space by 16× in both width and height ($1344/16 = 84$, $768/16 = 48$).
The Diffusion Transformer chops the latent grid into 2×2 spatial patches
(patch_size=(1, 2, 2)). $84/2 = 42$, $48/2 = 24$.
Attention computes correlations across all tokens. Because $16 \times 2 = 32$, any resolution not divisible by 32 leaves an odd latent axis that 2×2 patches cannot tile.
Type any pixel dimension to see why 16-divisibility alone causes a latent tiling failure:
How MiniMax H3 turns an aspect ratio request into an exact pixel resolution, and the practical traps to avoid.
adapt_canvas() returns the exact same 1344×768 resolution.
H3 has no resolution quality slider—only an aspect ratio shape choice.
Sparse attention algorithms like Sol-Attn require roughly 60,000 total tokens
before sparse pruning yields measurable speedups. At 243 frames: fast (1152×768) has
62,208 tokens (works), while draft (1024×768) has
55,296 tokens (falls below the floor). Use draft to verify pipeline
wiring, but benchmark Sol-Attn on fast or higher.
Every resolution produced by adapt_canvas() within the legal $1/4 \dots 4$ aspect range.
| Aspect Ratio | Resolution (px) | Latent Grid | Tokens / Frame | Attention Cost | Canvas Tier | Stability |
|---|