← All modelsarchitecture
Transformer
intermediatesupervised · self_supervised
parametricclassificationregressionforecastinggenerationrepresentationtextimageaudiotime_seriessequencesmultimodal
A Transformer builds context-aware representations by letting positions assign attention to other positions.
Mechanisms
attentionembeddingsbackpropagation
Properties
generativenonlinearpretrained ecosystemrepresentation learning
Constraints
high computehigh memoryrequires large data
Practical profile
- Explainability
- low
- Training cost
- high
- Inference cost
- high
- Data appetite
- high