Advanced and frontier / Advanced methods
Scaling to millions of points
Attention grows with the square of the points; latent and linear schemes keep it tractable.
Plain attention is quadratic. Recent architectures change what attends to what, so cost grows linearly with mesh size.
- Physics slices (Transolver): points are softly grouped into M learned slices of similar physical state; attention runs between the M slice tokens. Cost ~ N·M.
- Million-scale (Transolver++): locally adaptive slicing and parallelism across GPUs push single-GPU inputs past a million points.
- Linear attention (GNOT, OFormer): normalisation tricks that avoid forming the N×N matrix.
- Multiscale graph encoders + transformer processor (GAOT): local geometry is captured by attentional graph operators, global coupling by a transformer on fewer tokens.
Industrial example: Thinking in flow zones, not cells
No engineer tracks 50 million cells one by one. They reason in zones: boundary layer, underbody, wake, freestream. Physics-slice attention does the same: points in a similar flow state are grouped, and the groups exchange information.
References
- Transolver: A Fast Transformer Solver for PDEs on General Geometries (Wu et al., ICML 2024)
- Transolver++: PDEs on Million-Scale Geometries (Luo et al., ICML 2025)
- GNOT: A General Neural Operator Transformer (Hao et al., ICML 2023)
- Geometry Aware Operator Transformer (GAOT) (Wen et al., NeurIPS 2025)
Related
- Self-attention: Every token looks at every other token and decides what matters to it.
- Cross-attention: A small set of queries reads from a large set of points, so cost stays linear.
- Recent papers: Recent field-surrogate architectures from NeurIPS, ICML and ICLR, 2023 to 2026, in order.