Skip to content
View original post on X: Stability AI· 30/100AI score30/100
AISummary

Stability AI's Interactive Research team introduced SemanTok, which makes early video tokens more semantically meaningful so the representation is easier to predict. According to the post, a model using SemanTok matches or beats the performance of a model more than three times its size. The approach targets more efficient autoregressive video generation.

View original post on X x.com
Full text

What if making video-based world models smaller isn't just about better compression, but about making their representations easier to predict?

Recent approaches to video generation explore building scenes from coarse to fine. The first few tokens capture the big picture, like a person playing a guitar, while later tokens progressively add visual details.

Our Interactive Research team just published SemanTok, which takes this idea further. By making those early tokens more semantically meaningful, we give the model a clearer understanding of what's happening in a scene, making the representation easier to predict and video generation more efficient.

The result: a model using SemanTok matches or beats the performance of a model more than three times its size.

Read the full paper: https://stability.ai/research/semantok-predictable-semantic-tokens-for-efficient-autoregressive-video-generation

Source: Stability AI · x.comPublished · added here