Using the same model for task and evaluation: validate judge against humans
Original titleQ: Can I use the same model for both the main task and evaluation?
AISummary
Hamel Husain answers that the same model can usually serve as both the main task model and the evaluation judge. He advises testing whether the judge agrees with human labels on a held-out test set.
Source: Hamel Husain · x.comPublished · added here