Skip to content
Read the original: Eugene Yan· Published 36/100AI score36/100

Eugene Yan argues evals should weigh tail tasks, not median performance

Original titleWhen evaling models, we anchor on the median task. But this is like how devs estimate the median task accurately but underestimate the me...

AISummary

Eugene Yan argues that model evals anchor on median tasks, but tail tasks determine project completion, making reliable models like Fable and Opus the difference between success and failure.

He recommends treating models as collaborators who handle multi-hour or multi-day work with intent and success criteria, not as narrow-spec tools. Steve Yegge adds that Fable's carefulness is the dimension that matters most for production work.

Read the original x.com

Source: Eugene Yan · x.comPublished · added here