Skip to content
View original post on X: Reflection· 44/100AI score44/100

Reflection scales Beam on 10.5k GB300s in record RL run

AISummary

Reflection says it ran Beam, its reinforcement learning system, on 10.5k GB300 GPUs for four weeks, which it describes as the largest publicly documented RL run it knows of. The company credits algorithmic advances combined with distributed infrastructure for making the system scale. Across its eval suite, capabilities kept improving as RL increased, with no sign of a plateau.

Post on XView on X
@reflection_ai

A reply · the post it answers

We scaled Beam with the largest publicly documented RL run we’re aware of, which ran on 10.5k GB300s for 4 weeks.

Algorithmic advances coupled with distributed infra enabled an RL system that scales.

Across our eval suite, capabilities continued to improve as we increased RL - with no signs of plateau.

Source: Reflection · x.comPublished · added here