MiniMax Explains Why Its Model Failed to Output Certain Rare Chinese Tokens
Original titleWhy Can't the MiniMax LLM Say "Ma Jiaqi"? Internal Investigation of Sparse Token Forgetting
MiniMax says the M2 series could not generate the rare token "嘉祺" in names like Ma Jiaqi, and its investigation traced the cause to post-training.
The company found the token was learned in pretraining, but low coverage of rare tokens in post-training data caused lm_head vectors to drift.
Adding synthetic full-vocabulary repetition data restored generation for these tokens and reduced Japanese-to-Russian mixing from 47% to 1%.
The post traces a specific token failure through tokenizer, embedding, and lm_head checks, showing a reusable way to diagnose post-training generation problems.
Source: MiniMax Blog · minimax.ioPublished · added here