Skip to content
Read the original: Jerry Liu· Published 22/100AI score22/100

Jerry Liu and Snorkel's Vincent Sun discuss evals and RL environments

Original titleYesterday I hosted a fun dinner conversation with @vincentsunnchen from @SnorkelAI on evals and RL environments.

AISummary

Jerry Liu hosted a dinner with Snorkel's Vincent Sun on evals and RL environments, a topic shaped by models rapidly saturating benchmarks.

The conversation highlighted that building fair RL environments is hard, since failures are difficult to attribute to input, harness, or reward model, and that long-horizon evals spanning weeks or months remain very difficult.

The post also noted that regulated industries still require human-in-the-loop review because 80% accuracy is not sufficient.

Read the original x.com

Source: Jerry Liu · x.comPublished · added here