← Back to glossary

AI Benchmark

A versioned set of tasks, data, execution protocols, and metrics used to compare models or systems under the same conditions. A published number only has meaning alongside the evaluated population, inference configuration, scoring method, and benchmark version. Changing any of these parts changes the comparison contract.

Public benchmarks support comparability but may be far from real traffic or contaminated in training data. An internal benchmark should cover segments, edge cases, and failures with operational consequences while preserving a stable regression set. Combining complementary metrics reduces blind spots; an aggregate result without an error distribution can hide the exact group where the system fails.