Skip to content
Fundamentals Glossary

Latency

The time between sending a request and receiving a model's response.

Last updated

Latency measures responsiveness. With language models it includes queueing, prefill of the prompt and the time to generate each token. Long prompts, large models and distant servers all add delay. Streaming helps perceived speed by showing tokens as they arrive, while caching and smaller models cut real latency.

Related terms

Keep exploring

Browse the full AI glossary or compare AI tools that use this technology.

Report an issue with this page