Skip to content
Read the original: John Schulman· johnschulman2·Published AI score40/100

Schulman distinguishes risks of training AI on user data

Original titleThis isn't the most notable aspect of today's news, but on the user data issue, there are different kinds of *training on user data* with...

AISummary

John Schulman argues that training on user data carries very different privacy and IP risks depending on method.

Pretraining on user tokens poses high regurgitation risk, while distillation from prompts and RL from user traces carry lower regurgitation risk but can still leak customer IP. He notes de-identification is weak because long traces can still identify users, and AI companies rarely disclose what they do.

Read the original x.com

Source: John Schulman · x.com