Architecture / Building blocks
Cross-attention
A small set of queries reads from a large set of points, so cost stays linear.
In cross-attention, the queries come from one set of tokens and the keys and values from another. Surrogates use it twice: to read the geometry in, and to read the field out.
- Encoder: a small set of learned latent queries attends over all mesh points (Perceiver-style), summarising the geometry and the flow condition.
- Decoder: any output location becomes a query that attends to the latent tokens. Because queries are just coordinates, the model is discretisation-invariant: train on one mesh, predict on another.
- Conditioning: extra inputs such as inflow speed or design parameters can be injected as additional keys and values.
Industrial example: A probe you can place anywhere
In a wind tunnel, you move a pitot probe to wherever you want a reading. A cross-attention decoder works the same way: you ask for any location, and it reads the answer from the model's internal state, regardless of the mesh it was trained on.
References
- Perceiver: General Perception with Iterative Attention (Jaegle et al., ICML 2021)
- Transformer for PDEs’ Operator Learning (OFormer) (Li et al., TMLR 2023)
- GNOT: A General Neural Operator Transformer (Hao et al., ICML 2023)
- Universal Physics Transformers (Alkin et al., NeurIPS 2024)
Related
- Self-attention: Every token looks at every other token and decides what matters to it.
- Latent space: Compress millions of points into a few hundred learned slots that hold the physics.
- Decoders: Turn latent tokens into surface pressure, shear or volume flow, with physics built into the output.