← Back to glossary

Meta-evaluation

Evaluation of the evaluation method itself. It checks whether a metric, rubric, human evaluator, or LLM-as-judge separates good and bad cases according to the criterion it claims to represent. Agreement with experts, repeatability, sensitivity to perturbations, bias, and stability across versions are central signals.

A judge may produce precise scores while ranking answers poorly. Meta-evaluation uses annotated samples outside the calibration set, adversarial cases, and disagreement analysis to detect this drift. A change in judge model, prompt, or rubric requires renewed validation before historical results can be compared.