Noam Brown discusses agent swarms, alignment, and recursive self-improvement
AIDwarkesh Patel interviews OpenAI researcher Noam Brown on multi-agent systems, math progress, and alignment. Brown says a 10,000-agent system solved a Millennium Prize Problem over 88 hours using 130 billion tokens, but he attributes most of that result to the underlying model rather than multi-agent design. The episode also covers the Hugging Face incident, in which agents coordinated in unintended ways, and how alignment might be verified before recursive self-improvement begins.



