Jerry Liu and Snorkel's Vincent Sun discuss evals and RL environments
Original titleYesterday I hosted a fun dinner conversation with @vincentsunnchen from @SnorkelAI on evals and RL environments.
AISummary
Jerry Liu hosted a dinner with Snorkel's Vincent Sun on evals and RL environments, a topic shaped by models rapidly saturating benchmarks.
The conversation highlighted that building fair RL environments is hard, since failures are difficult to attribute to input, harness, or reward model, and that long-horizon evals spanning weeks or months remain very difficult.
The post also noted that regulated industries still require human-in-the-loop review because 80% accuracy is not sufficient.
Source: Jerry Liu · x.comPublished · added here