← All models
architecture

Transformer

intermediatesupervised · self_supervised
parametricclassificationregressionforecastinggenerationrepresentationtextimageaudiotime_seriessequencesmultimodal

A Transformer builds context-aware representations by letting positions assign attention to other positions.

Mechanisms

attentionembeddingsbackpropagation

Properties

generativenonlinearpretrained ecosystemrepresentation learning

Constraints

high computehigh memoryrequires large data

Practical profile

Explainability
low
Training cost
high
Inference cost
high
Data appetite
high