Skip to content

Transformer

A neural network architecture that uses self-attention to weigh the importance of every word in a sequence relative to every other word.

Last updated

Introduced in the 2017 paper "Attention Is All You Need", the transformer replaced recurrence with self-attention. That lets a model process all tokens in parallel and learn long-range relationships, which made training far more efficient on modern hardware. Nearly every leading language and multimodal model today is a transformer or a close variant of it.

Related terms

Read more about Transformer

Keep exploring

Browse the full AI glossary or compare AI tools that use this technology.

Report an issue with this page