Skip to content
Fundamentals Glossary

Inference

Running a trained model to produce an output, as opposed to training it.

Last updated

Inference is the act of using a model: feeding it input and getting a prediction or generated text back. It is what happens every time a user sends a message to an AI assistant. Inference cost and latency depend on model size, hardware and how many tokens are processed, which is why providers optimise it heavily.

Related terms

Keep exploring

Browse the full AI glossary or compare AI tools that use this technology.

Report an issue with this page