Skip to content
View original post on X: Google Gemma· 22/100AI score22/100

DiffusionGemma runs as a parallel decision model, faster than autoregressive generation

AISummary

Google Gemma's account says DiffusionGemma, running in a Jev-style decision setup, denoises an open canvas in one step rather than generating tokens sequentially, taking about 0.2 seconds on a DGX Spark.

It says full bidirectional attention lets every option attend to the full context at once, and that the model inherits Gemma 4's spatial vision capabilities for visual and text decisions.

Post on XView on X
@googlegemma

"DiffusionGemma as Jev" showcases the power of non-autoregressive architectures.

While Jev demonstrates the value of rapid decision models, running DiffusionGemma in this paradigm leverages canvas diffusion to evaluate structured choices in a single parallel pass:

⚡ ️Massive Parallelism: Denoises across an open canvas in a single step instead of sequential autoregressive token generation (~0.2s on a DGX spark).

🧠 Full Bidirectional Attention: Allows every option to attend to the full context concurrently, yielding well-calibrated decision distributions.

👁️ Multimodal Grounding: Inherits Gemma 4's spatial vision capabilities for complex visual and text decisions.

Read more about this approach here:
https://github.com/vllm-project/vllm/pull/57250
https://x.com/mmastrac/status/2100373761195401724
https://x.com/mmastrac/status/2100626193943052784

Matt Mastracci@mmastrac
I ran some real, live evals on Jev vs DiffusionGemma-as-Jev (my patch for vLLM!) DiffusionGemma comes out as the winner, I think. Headlines: Is Jev faster than DiffusionGemma? No ❌ (API vs DGX Spark) Is Jev smarter than DiffusionGemma? No ❌ (they're roughly tied!)
View quoted post on X

Source: Google Gemma · x.comPublished · added here