Eugene Yan argues evals should weigh tail tasks, not median performance
Original titleWhen evaling models, we anchor on the median task. But this is like how devs estimate the median task accurately but underestimate the me...
AISummary
Eugene Yan argues that model evals anchor on median tasks, but tail tasks determine project completion, making reliable models like Fable and Opus the difference between success and failure.
He recommends treating models as collaborators who handle multi-hour or multi-day work with intent and success criteria, not as narrow-spec tools. Steve Yegge adds that Fable's carefulness is the dimension that matters most for production work.
Source: Eugene Yan · x.comPublished · added here