LLM-as-a-Judge
Also called Model judge
LLM-as-a-judge uses a language model to assess outputs against instructions or a rubric. It can produce scores, comparisons, or written judgments to support evaluation.
[Zheng et al.]In practice · hypothetical example
A team asks a judge model to compare two summaries using a stated coverage rubric, then checks agreement with human reviewers.
[Zheng et al.]A little deeper
Research on model judges identifies biases such as preference for answer position or verbosity. Agreement with human judgment must be checked for the intended task. [Zheng et al.]
A common mix-up
A model judge is an unbiased ground truth.
The judge can be inconsistent or biased and needs validation. [Zheng et al.]
Why compare judge ratings with human reviews?
Sources & editorial notes
Evidence: supported. Primary-source support for this scoped entry; publication approved by the project owner.
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena ↗ (opens in new tab)Zheng et al. · Publication date unknown
Relevant section: Abstract
Last editorial review: 2026-09-13 by project-owner.
First observed in this corpus: Unknown.
Revision history
Revision 2 · Created 2026-09-13 · Updated 2026-09-13
Project owner approved the current content for publication. Existing evidence scope and limitations remain applicable.