Skip to content
View original post on X: elvisX· 50/100AI score50/100

Sakana AI proposes multi-agent self-supervision for recursive self-improvement without verifiers

AISummary

Sakana AI's MASS method lets one base model propose, run, and grade multi-agent workflows, keeping the best through evolutionary search, with no external verifier needed for open-ended tasks.

Two self-improvement cycles on Qwen3.6-27B raise performance per output token from 1.2x to 1.6x across four open-ended benchmarks. A student trained on multi-agent traces also beats a single-agent student trained on 1.4x more tokens.

Post on XView on X
elvisVerified on X
@omarsar0

Recommended paper from Sakana AI on recursive self-improvement.

They propose an interesting way to scale recursive self-improvement through multi-agent self-supervision.

In this line of research, self-improvement loops usually need an external verifier, so open-ended tasks without a checker are left out.

MASS removes that requirement.

One base model proposes multi-agent workflows, runs them and grades them, and an evolutionary search keeps the workflows that score best. The model is then fine-tuned on its own traces, and the improved model starts the next cycle as a better optimizer and grader.

Two cycles on Qwen3.6-27B raise performance per output token from 1.2 to 1.6x on four open-ended benchmarks.

A student trained on multi-agent traces also beats a single-agent student trained on 1.4x more tokens.

Paper: https://arxiv.org/abs/2610.12176

Chat with Paper: https://academy.dair.ai/papers/recursive-self-improvement-through-multi-agent-self-supervision-2610.12176

Source: elvis · x.comPublished · added here