Skip to content
Read the original: Amazon Science·Published AI score46/100

Dependence-Aware Aggregation Improves LLM-as-a-Judge Accuracy by 9% to 14%

Original titleWhen LLM judges agree, should we believe them?

AISummary

Amazon researchers proposed a dependence-aware method for aggregating LLM judges' votes, using an Ising model to account for correlated errors among judges.

The approach outperformed a weighted majority-vote baseline by 9% to 14% on standard metrics across three binary tasks, including relevance classification, where it reached 0.912 accuracy versus 0.820.

The method is unsupervised, learning from judge outputs without human reference labels.

Read the original amazon.science

Source: Amazon Science · amazon.science