Skip to content

Benchmark

A standard test used to compare model performance on a defined task.

Last updated

Benchmarks provide a shared way to measure models, from multiple-choice knowledge tests to coding and reasoning suites. They are useful for tracking progress but easy to over-interpret: results depend on prompt format, and strong scores do not always translate into success on real workflows. Treat them as one signal among several.

Related terms

Keep exploring

Browse the full AI glossary or compare AI tools that use this technology.

Report an issue with this page