Architecture / Building blocks
Self-attention
Every token looks at every other token and decides what matters to it.
Each token asks a question (a query), advertises what it contains (a key) and offers information (a value). It then gathers the values of the tokens whose keys best match its query.
- Weights = softmax(Q·Kᵀ / √d): one row of weights per token, summing to 1.
- Output = weights · V: each token becomes a weighted mix of the others. This is how one region of a body can be informed by a distant one, in a single layer.
- Unlike message passing, which reaches one hop per layer, attention is global from the first layer.
- The cost is the catch: N tokens means N² weights. 100,000 points is already 10 billion pairs.
Industrial example: Why a race car is designed as one system
On a race car, the front wing shapes the air that reaches the floor and the rear wing. Change one part and the others feel it. Self-attention lets every region of the car consult every other region, and learn how strongly they influence each other.
References
- Attention Is All You Need (Vaswani et al., NeurIPS 2017)
- Transformer for PDEs’ Operator Learning (OFormer) (Li et al., TMLR 2023)
Related
- Cross-attention: A small set of queries reads from a large set of points, so cost stays linear.
- Scaling to millions of points: Attention grows with the square of the points; latent and linear schemes keep it tractable.
- Tokens: Each mesh point becomes a vector of numbers describing where it is and what it carries.