REVIEW 4 cited by
CViT: Continuous Vision Transformer for Operator Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Operator learning, which aims to approximate maps between infinite-dimensional function spaces, is an important area in scientific machine learning with applications across various physical domains. Here we introduce the Continuous Vision Transformer (CViT), a novel neural operator architecture that leverages advances in computer vision to address challenges in learning complex physical systems. CViT combines a vision transformer encoder, a novel grid-based coordinate embedding, and a query-wise cross-attention mechanism to effectively capture multi-scale dependencies. This design allows for flexible output representations and consistent evaluation at arbitrary resolutions. We demonstrate CViT's effectiveness across a diverse range of partial differential equation (PDE) systems, including fluid dynamics, climate modeling, and reaction-diffusion processes. Our comprehensive experiments show that CViT achieves state-of-the-art performance on multiple benchmarks, often surpassing larger foundation models, even without extensive pretraining and roll-out fine-tuning. Taken together, CViT exhibits robust handling of discontinuous solutions, multi-scale features, and intricate spatio-temporal dynamics. Our contributions can be viewed as a significant step towards adapting advanced computer vision architectures for building more flexible and accurate machine learning models in the physical sciences.
Forward citations
Cited by 4 Pith papers
-
Adaptive Physics Transformer with Fused Global-Local Attention for Subsurface Energy Systems
APT, a mesh-agnostic neural operator fusing graph-based local features with global attention, is claimed to be the first architecture trained directly on adaptive-mesh-refinement simulations and outperforms state-of-t...
-
Structure-Preserving Learning Improves Geometry Generalization in Neural PDEs
A geometry-conditioned Whitney-form neural network that solves a learned discrete conservation law improves out-of-distribution geometry generalization for steady-state PDEs compared with regression-based neural operators.
-
Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics
MANO replaces quadratic self-attention with multiscale windowed attention over progressively downsampled grids, keeping complexity linear and reporting competitive accuracy on vision and PDE benchmarks.
-
Equivariant Eikonal Neural Networks: Grid-Free, Scalable Travel-Time Prediction on Homogeneous Spaces
E-NES uses Lie-group point-cloud conditioning and equivariant neural fields to make grid-free eikonal travel-time prediction steerable under rotations and translations, with complete invariant features and competitive...
Discussion (0). Continue with ORCID to comment.