Fundamentals
Glossary
Latency
The time between sending a request and receiving a model's response.
Last updated
Latency measures responsiveness. With language models it includes queueing, prefill of the prompt and the time to generate each token. Long prompts, large models and distant servers all add delay. Streaming helps perceived speed by showing tokens as they arrive, while caching and smaller models cut real latency.
Related terms
Keep exploring
Browse the full AI glossary or compare AI tools that use this technology.