Will it run?
Models

MiniMax opens VTP and challenges core assumption in latent diffusion models

By Rae Whitlock Clawpit staff
MiniMax opens VTP and challenges core assumption in latent diffusion models

The video team at MiniMax, also known under the Hailuo brand, announced that it is releasing Visual Tokenizer Pre-training (VTP) to the community. VTP is a tiered pre-training framework for visual tokenizers built for next-generation generative models. The announcement contests a central assumption in latent diffusion models: that scaling the visual tokenizer in the first stage—whether by increasing data volume, compute power, or model size—does not improve the quality of generated frames in the second stage.

The team claims its research refutes this assumption. The method it developed combines representation learning with compression and reconstruction, and for the first time presents a scaling curve linking tokenizer enlargement to performance gains in diffusion transformers. The core claim is that frame-quality improvement is achieved solely by allocating more resources to tokenizer training, without adding computational load to the generative model at inference time. In other words, a richer spatial representation of visual information can be obtained before the final model runs, so the generator does not incur a runtime performance penalty. The release arrived as an early Christmas greeting to the community.

The claims are bold, but supporting evidence has not yet been provided.