Reference and tools / Catalogues
Technique library
Thirty techniques that make surrogates work in practice, from basics to the frontier.
- Design of experiments (Basics, Data & sampling): Latin hypercube, Sobol sequences or maximin designs spread simulations evenly over the design space. Every simulation is expensive; space-filling designs give the most information per run.
- Response surfaces & kriging (Basics, Architecture): Polynomial regression and Gaussian-process interpolation over a few design parameters. Data-efficient, interpretable, and kriging gives uncertainty for free.
- Input & output normalisation (Basics, Training): Scale coordinates, parameters and fields to unit range; predict non-dimensional quantities (e.g. a pressure coefficient rather than pascals). Stabilises training and makes models transfer across scales and units.
- Hold-out & cross-validation (Basics, Uncertainty & trust): Measure RMSE, relative L2 and R² on designs never used in training, by k-fold or a fixed hold-out set. Training error says nothing about new designs.
- Signed distance & normals (Basics, Geometry & inputs): Give each point its distance to the wall and the surface normal as input features. Tells the network where the body is: the physics changes fastest near walls.
- Encode–process–decode (Basics, Architecture): Lift inputs to an embedding, process with a stack of blocks, read out the field. The template almost every modern surrogate follows.
- Fourier features (Core, Geometry & inputs): Map coordinates through sines and cosines at several frequencies before the first layer. Networks are biased towards smooth functions; this lets them resolve sharp gradients.
- Message passing (Core, Architecture): Each node updates from aggregated messages of its neighbours, one hop per layer. Local, mesh-native, and matches how PDE stencils couple neighbouring cells.
- Spectral convolution (Core, Architecture): Learned filters on the lowest Fourier modes of the field. Global coupling in one layer, evaluated at any resolution.
- Multiscale hierarchies (Core, Architecture): Coarsen the grid, graph or point set, process, and refine with skip connections. Captures both far-field coupling and near-wall detail at manageable cost.
- Self- & cross-attention (Core, Architecture): Tokens weight each other by query–key similarity; cross-attention reads one set from another. Global, data-dependent coupling and resolution-free decoding.
- Noise injection for rollout (Core, Training): Corrupt inputs with noise during training of time-stepping models. The model learns to correct its own errors, so long rollouts stay stable.
- Reduced-order models (POD) (Core, Architecture): Compress snapshots onto a few dominant modes and regress the mode coefficients. Very fast and interpretable on one fixed mesh.
- Latent tokens (Advanced, Architecture): Encode N mesh points into M ≪ N latent tokens; process only those. Processing cost becomes independent of mesh size.
- Physics slices (Advanced, Architecture): Softly group points by learned physical state and attend between groups. Linear cost in mesh size while keeping global attention.
- Anchors & branched decoders (Advanced, Architecture): Full attention among a few anchor points; millions of queries attend only to anchors; separate surface and volume branches. High-resolution outputs at 100M-cell scale.
- Rotary & relative position encodings (Advanced, Geometry & inputs): Encode relative positions inside attention instead of adding absolute coordinates. Better geometric generalisation on large vehicle meshes.
- Spectral graph embeddings (Advanced, Geometry & inputs): Laplacian eigenvectors of the mesh give each node a connectivity-aware position. Distinguishes parts that are close in space but not connected.
- Global + local processing (Advanced, Architecture): Combine global latent exchange with local refinement around each query point. Global models blur boundary layers; local ones miss far-field coupling.
- Hard physics constraints (Advanced, Physics): Build conservation into the output, for example a divergence-free velocity from a predicted potential. Guarantees physical consistency instead of hoping the loss enforces it.
- Physics-informed losses (Advanced, Physics): Add PDE residuals or boundary conditions to the training loss. Fewer data needed and fewer unphysical predictions, at extra training cost.
- Multi-fidelity learning (Advanced, Data & sampling): Combine many cheap low-fidelity runs with a few expensive high-fidelity ones (co-kriging, residual learning). Most of the data cost moves to cheap solvers.
- Ensembles & MC dropout (Advanced, Uncertainty & trust): Several networks, or dropout at inference, give a spread of predictions. A practical epistemic-uncertainty estimate for neural surrogates.
- Pretrained physics foundation models (Frontier, Training): Pretrain on many PDEs and simulators, then fine-tune on the task. Order-of-magnitude fewer task-specific samples.
- Transfer to new geometry families (Frontier, Training): Adapt a surrogate trained on one vehicle family to another with a small new dataset. Products change; retraining from scratch every programme is too costly.
- Joint-embedding prediction (JEPA) (Frontier, Architecture): Predict the latent representation of the flow from the latent of the geometry and conditions; compare in latent space, decode only on demand. Decouples prediction cost from field resolution and yields a latent space organised by physics.
- Anti-collapse regularisation (Frontier, Training): Regularisers such as SIGReg or VICReg keep latent tokens spread out and informative. Without them a latent-prediction model can trivially map everything to one point.
- Latent probing & design (Frontier, Design use): Fit linear probes on the latent space to read lift, drag or settings; interpolate or optimise directly in latent space. Turns the surrogate’s internal state into an interface for analysis and early design exploration.
- Generative surrogates (Frontier, Uncertainty & trust): Diffusion or flow-matching models sample fields instead of regressing a mean. Capture uncertainty and multimodal flows; stable long rollouts.
- Conformal prediction (Frontier, Uncertainty & trust): Calibrate error bars on held-out data so they cover the truth with a guaranteed probability. Engineers need intervals, not point estimates, to sign off a design.
- Active learning & Bayesian optimisation (Frontier, Design use): Let the surrogate choose the next simulation, e.g. by expected improvement or uncertainty. Spends the simulation budget where it changes decisions.
- Surrogate-assisted optimisation (Frontier, Design use): Evolutionary or gradient-based optimisers query the surrogate, with periodic high-fidelity checks. Thousands of evaluations become affordable; gradients come for free from differentiable models.
- Sensitivity analysis (Frontier, Design use): Sobol indices and variance propagation computed through the surrogate. Which parameters matter, and how tolerances propagate, without millions of solves.
Related
- Model atlas: Ten surrogate families at a glance, grouped by the data they consume.
- Limits and trust: A surrogate is only as good as its data: know the domain, quantify error, verify the winner.
- Training: Fit the weights once on the training simulations and watch the validation curve.