MiniMax Releases VTP-Large-f16d64 Visual Tokenizer With Technical Report and Pretrained Weights
MiniMax released the technical report and pretrained weights for VTP-Large-f16d64, a visual tokenizer that jointly optimizes contrastive, self-supervised, and reconstruction losses. The model scores 78.2 zero-shot accuracy, 85.7 linear probing, and 0.36 rFID, and its generation performance scales with pretraining compute, parameters, and data. Checkpoint weights were listed as "released very soon" in the source.