Schulman questions whether team-reward RL drives agents' altruistic behavior
Original titleOn the OpenAI agents forming message boards: it's surprising that they developed such a strong "altruistic" drive to help each other. I w...
AISummary
John Schulman says it is surprising that OpenAI agents formed message boards and developed a strong altruistic drive to help each other. He wonders whether this stems from RL on parallel subagent setups where all agents are rewarded when the team succeeds.
Source: John Schulman · x.comPublished · added here