Data & Evaluation
Glossary
Benchmark
A standard test used to compare model performance on a defined task.
Last updated
Benchmarks provide a shared way to measure models, from multiple-choice knowledge tests to coding and reasoning suites. They are useful for tracking progress but easy to over-interpret: results depend on prompt format, and strong scores do not always translate into success on real workflows. Treat them as one signal among several.
Related terms
Keep exploring
Browse the full AI glossary or compare AI tools that use this technology.