REVIEW 2 major objections 1 minor 43 references
CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval
T0 review · 2 major / 1 minor · reviewed 2026-06-28 · grok-4.3
Pith's one-line read CR-JEPA uses modality-specific stems, a shared predictive trunk, and decoupled heads to raise cross-modal remote sensing retrieval accuracy by more than 14 points while matching same-modal performance with fewer parameters.
desk verdict CR-JEPA applies existing JEPA ideas to cross-modal RS retrieval and reports cross-modal gains on BEN-14K, but the abstract leaves the no-trade-off claim on same-modal performance unquantified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Modality-specific stems plus shared transformer trunk trained with within- and across-modality JEPA predictive objectives, stabilized by sketched isotropic Gaussian regularization on the retrieval projections and routed through decoupled same-modal and cross-modal heads.
What would settle it
An experiment on BEN-14K in which cross-modal accuracy falls back to the X-JEPA level or same-modal accuracy drops when the sketched isotropic Gaussian regularization is removed would falsify the central claim.
Extended reading notes
Core claim
CR-JEPA employs modality-specific stems feeding a shared transformer trunk that is trained with JEPA-style predictive objectives to reconstruct masked latent targets within and across modalities. Sketched isotropic Gaussian regularization is applied to the raw retrieval projections, and a decoupled-head design supplies one unified head for same-modal retrieval and a separate head for cross-modal search. On BEN-14K this produces S1-to-S2 retrieval of 75.82 percent and S2-to-S1 retrieval of 75.40 percent, against 61.23 percent and 63.73 percent for the prior X-JEPA baseline, while same-modal retrieval stays competitive and total parameter count is lower.
Load-bearing premise
The predictive objectives, regularization, and decoupled heads together produce cross-modal alignment and same-modal neighborhood preservation at the same time without collapse or accuracy trade-offs on heterogeneous remote-sensing data.
Editorial extensions
If this is right
- Cross-modal retrieval between Sentinel-1 and Sentinel-2 scenes improves by more than 14 percentage points.
- Same-modal retrieval accuracy remains competitive without increasing model size.
- The architecture jointly optimizes semantic alignment across modalities and local neighborhood preservation within each modality.
- The same design yields gains on three separate remote-sensing retrieval benchmarks.
Reading between the lines
- The decoupled-head pattern could be tested on other pairs of imaging modalities that differ in physics.
- The regularization step may be the main stabilizer when predictive objectives are applied to high-variance sensor data.
- If the gains hold on larger or noisier collections, the method could reduce the need for separate retrieval systems per sensor type.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CR-JEPA, a Cross-modal Retrieval Joint-Embedding Predictive Architecture for dual-modality remote sensing image retrieval. It employs modality-specific stems, a shared transformer trunk, JEPA-style predictive objectives to estimate masked latent target features within and across modalities, sketched isotropic Gaussian regularization on raw retrieval projections, and a decoupled-head design with a unified retrieval head for same-modal tasks and a cross-modal retrieval head. Evaluations on BEN-14K, CBRSIR_VS, and DSRSID report cross-modal gains on BEN-14K (S1 to S2 retrieval from 61.23% to 75.82%; S2 to S1 from 63.73% to 75.40% over X-JEPA) while claiming competitive same-modal retrieval with fewer parameters.
Significance. If the empirical results hold under full experimental scrutiny, the work could provide a practical framework for cross-modal retrieval in remote sensing by combining predictive objectives with regularization to achieve semantic alignment across heterogeneous modalities without collapse. The decoupled-head design and use of JEPA objectives represent architectural strengths for jointly handling alignment and neighborhood preservation. The reported cross-modal improvements suggest potential utility in multi-sensor applications, though significance hinges on verifying the no-trade-off claim for same-modal performance.
major comments (2)
- [Abstract] Abstract (final sentence): The central claim that CR-JEPA achieves 'competitive same-modal retrieval with fewer parameters' provides no quantitative metrics, tables, or direct comparisons for same-modal performance. This is load-bearing for the contribution, as the abstract and introduction stress that a single objective may fail to jointly support cross-modal alignment and same-modal neighborhood preservation, yet only cross-modal deltas are quantified.
- [Abstract] Abstract and likely §4 (Experimental results): Specific percentage gains are stated (e.g., 61.23% to 75.82%) but without any mention of experimental protocol, baseline implementation details beyond X-JEPA, error bars, statistical tests, dataset splits, or number of runs. This prevents assessment of whether the numbers reliably support the claims.
minor comments (1)
- [Abstract] The abstract packs many technical components into one paragraph; a brief clarification of how the sketched isotropic Gaussian regularization interacts with the JEPA objectives would aid readability.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback on the abstract. We address the major comments point by point below and will make revisions to strengthen the quantitative support and experimental transparency.
read point-by-point responses
-
Referee: [Abstract] Abstract (final sentence): The central claim that CR-JEPA achieves 'competitive same-modal retrieval with fewer parameters' provides no quantitative metrics, tables, or direct comparisons for same-modal performance. This is load-bearing for the contribution, as the abstract and introduction stress that a single objective may fail to jointly support cross-modal alignment and same-modal neighborhood preservation, yet only cross-modal deltas are quantified.
Authors: We agree that the abstract would be strengthened by including quantitative support for the same-modal claim. Section 4 of the manuscript contains tables reporting same-modal mAP and parameter counts demonstrating competitive performance with fewer parameters than X-JEPA. We will revise the abstract's final sentence to reference these specific metrics and tables. revision: yes
-
Referee: [Abstract] Abstract and likely §4 (Experimental results): Specific percentage gains are stated (e.g., 61.23% to 75.82%) but without any mention of experimental protocol, baseline implementation details beyond X-JEPA, error bars, statistical tests, dataset splits, or number of runs. This prevents assessment of whether the numbers reliably support the claims.
Authors: Dataset splits, baseline reproduction details, and evaluation procedures are described in Sections 3.3 and 4.1. However, the abstract omits reference to them, and the current version lacks error bars and statistical tests. We will add a concise protocol note to the abstract, report error bars from multiple runs, and include statistical significance in the revised §4. revision: partial
Circularity Check
No significant circularity; empirical architecture claims rest on reported metrics, not self-referential reductions.
full rationale
The manuscript proposes CR-JEPA via modality-specific stems, shared trunk, JEPA predictive objectives, sketched isotropic Gaussian regularization, and decoupled heads. All performance claims are quantified through standard retrieval metrics (e.g., BEN-14K cross-modal gains over X-JEPA) rather than any derivation that reduces a prediction to a fitted parameter or self-citation by construction. No equations appear that equate outputs to inputs tautologically, no uniqueness theorems are imported from the same authors, and no ansatz is smuggled via prior self-work. The design is motivated by external references (LeJEPA) and evaluated on external benchmarks, rendering the central claims self-contained against falsifiable data.
Assumptions & free parameters
Cite this review
Pith. "Pith review of CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval." pith.science (2026). https://pith.science/paper/X4M6VM74
@misc{pith2026260600706,
author = {Pith},
title = {Pith review of: CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval},
year = {2026},
howpublished = {\url{https://pith.science/paper/X4M6VM74}},
note = {Machine review of arXiv:2606.00706}
}
read the original abstract
Cross-modal remote sensing image retrieval aims to retrieve semantically related scenes across heterogeneous sensing modalities. This remains challenging because paired observations may differ substantially in imaging physics, spatial resolution, spectral configuration, and visual appearance. Moreover, a single retrieval projection trained with one objective may be insufficient to jointly support cross-modal semantic alignment and same-modal neighbourhood preservation. We propose CR-JEPA, a Cross-modal Retrieval Joint-Embedding Predictive Architecture for dual-modality remote sensing retrieval. The model uses modality-specific stems, a shared transformer trunk, and JEPA-style predictive objectives to estimate masked latent target features within and across modalities. Inspired by LeJEPA, we apply Sketched Isotropic Gaussian Regularization to raw retrieval projections to stabilize embeddings and mitigate collapse. CR-JEPA further employs a decoupled-head design with a unified retrieval head for same-modal retrieval and a cross-modal retrieval head for cross-modal search. We evaluate CR-JEPA on BEN-14K, CBRSIR_VS, and DSRSID. On BEN-14K, CR-JEPA improves S1 to S2 retrieval from 61.23% to 75.82% and S2 to S1 retrieval from 63.73% to 75.40% over X-JEPA, while also achieving competitive same-modal retrieval with fewer parameters.
Figures
Reference graph
Works this paper leans on
-
[1]
Self-supervised learning from im- ages with a joint-embedding predictive architecture
Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas. Self-supervised learning from im- ages with a joint-embedding predictive architecture. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15619–15629, 2023
2023
-
[2]
Anysat: One earth observation model for many resolutions, scales, and modalities
Guillaume Astruc, Nicolas Gonthier, Clément Mallet, and Loïc Landrieu. Anysat: One earth observation model for many resolutions, scales, and modalities. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025
2025
-
[3]
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
Randall Balestriero and Yann LeCun. Lejepa: Provable and scalable self-supervised learning without the heuristics.arXiv preprint arXiv:2511.08544, 2025
work page Pith review arXiv 2025
-
[4]
VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning
Adrien Bardes, Jean Ponce, and Yann LeCun. Vicreg: Variance-invariance-covariance regularization for self-supervised learning.arXiv preprint arXiv:2105.04906, 2021
work page Pith review arXiv 2021
-
[5]
Mul- tilabel remote sensing image retrieval using a semisupervised graph-theoretic method
Bindita Chaudhuri, Begüm Demir, Subhasis Chaudhuri, and Lorenzo Bruzzone. Mul- tilabel remote sensing image retrieval using a semisupervised graph-theoretic method. IEEE Transactions on Geoscience and Remote Sensing, 56(2):1144–1158, 2018
2018
-
[6]
Cmir-net: A deep learning based model for cross-modal retrieval in remote sensing.Pattern Recognition Letters, 131:456–462, 2020
Ushasi Chaudhuri, Biplab Banerjee, Avik Bhattacharya, and Mihai Datcu. Cmir-net: A deep learning based model for cross-modal retrieval in remote sensing.Pattern Recognition Letters, 131:456–462, 2020
2020
-
[7]
Boosting cross-modal re- trieval in remote sensing via a novel unified attention network.Neural Networks, 180: 106718, 2024
Shabnam Choudhury, Devansh Saini, and Biplab Banerjee. Boosting cross-modal re- trieval in remote sensing via a novel unified attention network.Neural Networks, 180: 106718, 2024
2024
-
[8]
Rejepa: A novel joint-embedding predictive architecture for efficient remote sensing image re- trieval
Shabnam Choudhury, Yash Salunkhe, Sarthak Mehrotra, and Biplab Banerjee. Rejepa: A novel joint-embedding predictive architecture for efficient remote sensing image re- trieval. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 2398–2407, 2025. 16HOSSAIN ET AL.: CR-JEPA: CROSS-MODAL REMOTE SENSING IMAGE RETRIEV AL
2025
Show all 43 references
-
[9]
X-jepa: A novel joint learning cross-modal predictive alignment framework for remote sensing image retrieval
Shabnam Choudhury, Yash Salunkhe, Vaibhav Rajan, Subhasis Chaudhuri, and Biplab Banerjee. X-jepa: A novel joint learning cross-modal predictive alignment framework for remote sensing image retrieval. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer V...
2026
-
[10]
Satmae: Pre-training transformers for temporal and multi-spectral satellite imagery
Yezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu, Erik Rozi, Yutong He, Mar- shall Burke, David Lobell, and Stefano Ermon. Satmae: Pre-training transformers for temporal and multi-spectral satellite imagery. InAdvances in Neural Information Pro- cessing Systems, volume 35,...
2022
-
[11]
From alignment to prediction: A study of self-supervised learning and predictive representation learning.arXiv preprint arXiv:2604.13518, 2026
Mintu Dutta, Ritesh Vyas, and Mohendra Roy. From alignment to prediction: A study of self-supervised learning and predictive representation learning.arXiv preprint arXiv:2604.13518, 2026
2026 arXiv
-
[12]
Croma: Remote sensing repre- sentations with contrastive radar-optical masked autoencoders
Anthony Fuller, Koreen Millard, and James Green. Croma: Remote sensing repre- sentations with contrastive radar-optical masked autoencoders. InAdvances in Neural Information Processing Systems, volume 36, 2023
2023
-
[13]
Skysense: A multi-modal remote sens- ing foundation model towards universal interpretation for earth observation imagery
Xin Guo, Jiangwei Lao, Bo Dang, Yingying Zhang, Lei Yu, Lixiang Ru, Liheng Zhong, Ziyuan Huang, Kang Wu, Dingxiang Hu, et al. Skysense: A multi-modal remote sens- ing foundation model towards universal interpretation for earth observation imagery. InProceedings of the IEEE/CVF...
2024
-
[14]
X-clpa: A contrastive learning and pro- totypical alignment-based crossmodal remote sensing image retrieval.Expert Systems with Applications, 308:131169, 2026
Aparna H, Biplab Banerjee, and Avik Hati. X-clpa: A contrastive learning and pro- totypical alignment-based crossmodal remote sensing image retrieval.Expert Systems with Applications, 308:131169, 2026
2026
-
[15]
Csmoe: An efficient remote sens- ing foundation model with soft mixture-of-experts.arXiv preprint arXiv:2509.14104, 2025
Leonard Hackel, Tom Burgert, and Begüm Demir. Csmoe: An efficient remote sens- ing foundation model with soft mixture-of-experts.arXiv preprint arXiv:2509.14104, 2025
2025
-
[16]
Explor- ing masked autoencoders for sensor-agnostic image retrieval in remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 2024
Jakob Hackstein, Gencer Sumbul, Kai Norman Clasen, and Begüm Demir. Explor- ing masked autoencoders for sensor-agnostic image retrieval in remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 2024
2024
-
[17]
Geomeld: Toward semantically grounded foundation models for remote sensing
Maram Hasan, Md Aminur Hossain, Savitra Roy, Souparna Bhowmik, Ayush V Patel, Mainak Singha, Subhasis Chaudhuri, Muhammad Haris Khan, and Biplab Banerjee. Geomeld: Toward semantically grounded foundation models for remote sensing. In Proceedings of the IEEE/CVF Conference on C...
2026
-
[18]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022
2022
-
[19]
OT-PFCNet: A generative pre-alignment framework for cross-domain remote sensing retrieval
Jirui Huang and Dongjie Zhao. OT-PFCNet: A generative pre-alignment framework for cross-domain remote sensing retrieval. InIGARSS 2025 - 2025 IEEE International Geoscience and Remote Sensing Symposium, pages 6620–6624, 2025. HOSSAIN ET AL.: CR-JEPA: CROSS-MODAL REMOTE SENSING ...
2025
-
[20]
Raffaele Imbriaco, Clint Sebastian, Egor Bondarev, and Peter H. N. de With. Toward multilabel image retrieval for remote sensing.IEEE Transactions on Geoscience and Remote Sensing, 60:1–14, 2022
2022
-
[21]
Yansheng Li, Yongjun Zhang, Xin Huang, and Jiayi Ma. Learning source-invariant deep hashing convolutional neural networks for cross-source remote sensing image retrieval.IEEE Transactions on Geoscience and Remote Sensing, 56(11):6521–6536, 2018
2018
-
[22]
Image retrieval from remote sensing big data: A survey.Information Fusion, 67:94–115, 2021
Yansheng Li, Jiayi Ma, and Yongjun Zhang. Image retrieval from remote sensing big data: A survey.Information Fusion, 67:94–115, 2021
2021
-
[23]
Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017
2017 arXiv
-
[24]
Cross-source image retrieval based on ensemble learning and knowledge distillation for remote sensing images
Jingjing Ma, Duanpeng Shi, Xu Tang, Xiangrong Zhang, Xiao Han, and Licheng Jiao. Cross-source image retrieval based on ensemble learning and knowledge distillation for remote sensing images. InIGARSS 2021 - 2021 IEEE International Geoscience and Remote Sensing Symposium, pages...
2021
-
[25]
Rethinking transformers pre-training for multi-spectral satellite imagery
Mubashir Noman, Muzammal Naseer, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan, and Fahad Shahbaz Khan. Rethinking transformers pre-training for multi-spectral satellite imagery. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 2...
2024
-
[26]
Reed, Ritwik Gupta, Shufan Li, Sarah Brockman, Christopher Funk, Brian Clipp, Kurt Keutzer, Salvatore Candido, Matt Uyttendaele, and Trevor Darrell
Colorado J. Reed, Ritwik Gupta, Shufan Li, Sarah Brockman, Christopher Funk, Brian Clipp, Kurt Keutzer, Salvatore Candido, Matt Uyttendaele, and Trevor Darrell. Scale- mae: A scale-aware masked autoencoder for multiscale geospatial representation learn- ing. InProceedings of t...
2023
-
[27]
Performance evaluation of single-label and multi-label remote sensing image retrieval using a dense labeling dataset.Remote Sensing, 10(6):964, 2018
Zhenfeng Shao, Ke Yang, and Weixun Zhou. Performance evaluation of single-label and multi-label remote sensing image retrieval using a dense labeling dataset.Remote Sensing, 10(6):964, 2018
2018
-
[28]
Multilabel remote sensing image retrieval based on fully convolutional network.IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 13: 318–328, 2020
Zhenfeng Shao, Weixun Zhou, Xueqing Deng, Maoding Zhang, and Qimin Cheng. Multilabel remote sensing image retrieval based on fully convolutional network.IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 13: 318–328, 2020
2020
-
[29]
S. K. Sudha and S. Aji. A review on recent advances in remote sensing image retrieval techniques.Journal of the Indian Society of Remote Sensing, 47(12):2129–2139, 2019
2019
-
[30]
Bigearthnet: A large-scale benchmark archive for remote sensing image understanding
Gencer Sumbul, Marcela Charfuelan, Begüm Demir, and V olker Markl. Bigearthnet: A large-scale benchmark archive for remote sensing image understanding. InProceedings of the IEEE International Geoscience and Remote Sensing Symposium, pages 5901– 5904, 2019
2019
-
[31]
Gencer Sumbul, Arne De Wall, Tristan Kreuziger, Filipe Marcelino, Hugo Costa, Pe- dro Benevides, Mário Caetano, Begüm Demir, and V olker Markl. Bigearthnet-mm: A large-scale, multimodal, multilabel benchmark archive for remote sensing image classi- fication and retrieval.IEEE ...
2021
-
[32]
A novel self-supervised cross- modal image retrieval method in remote sensing
Gencer Sumbul, Markus Müller, and Begüm Demir. A novel self-supervised cross- modal image retrieval method in remote sensing. InProceedings of the IEEE Interna- tional Conference on Image Processing, pages 2426–2430, 2022
2022
-
[33]
Yuxi Sun, Shanshan Feng, Yunming Ye, Xutao Li, Jian Kang, Zhichao Huang, and Chuyao Luo. Multisensor fusion and explicit semantic preserving-based deep hashing for cross-modal remote sensing image retrieval.IEEE Transactions on Geoscience and Remote Sensing, 60:1–14, 2022
2022
-
[34]
Cross-scale mae: A tale of multiscale exploitation in remote sensing
Maofeng Tang, Andrei Cozma, Konstantinos Georgiou, and Hairong Qi. Cross-scale mae: A tale of multiscale exploitation in remote sensing. InAdvances in Neural Infor- mation Processing Systems, volume 36, 2023
2023
-
[35]
A survey of deep learning based remote sensing image retrieval.Remote Sensing, 14(5):1323, 2022
Xi Tong, Gui-Song Xia, Qikai Lu, and Weimin Shen. A survey of deep learning based remote sensing image retrieval.Remote Sensing, 14(5):1323, 2022
2022
-
[36]
Albrecht, and Xiao Xiang Zhu
Yi Wang, Nassim Ait Ali Braham, Zhitong Xiong, Chenying Liu, Conrad M. Albrecht, and Xiao Xiang Zhu. Ssl4eo-s12: A large-scale multi-modal, multi-temporal dataset for self-supervised learning in earth observation.IEEE Geoscience and Remote Sensing Magazine, 11(3):98–106, 2023
2023
-
[37]
Albrecht, Nassim Ait Ali Braham, Chenying Liu, Zhitong Xiong, and Xiao Xiang Zhu
Yi Wang, Conrad M. Albrecht, Nassim Ait Ali Braham, Chenying Liu, Zhitong Xiong, and Xiao Xiang Zhu. Decoupling common and unique representations for multimodal self-supervised learning. InEuropean Conference on Computer Vision, 2024
2024
-
[38]
Robust cross-modal remote sensing image retrieval via maximal correlation augmentation.IEEE Transactions on Geoscience and Remote Sensing, 62:1–17, 2024
Zhuoyue Wang, Xueqian Wang, Gang Li, and Chengxi Li. Robust cross-modal remote sensing image retrieval via maximal correlation augmentation.IEEE Transactions on Geoscience and Remote Sensing, 62:1–17, 2024
2024
-
[39]
Wei Xiong, Zhenyu Xiong, Yaqi Cui, and Yafei Lv. A discriminative distillation net- work for cross-source remote sensing image retrieval.IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 13:1234–1247, 2020
2020
-
[40]
Frequency-guided mutual distillation for lightweight cross-modal remote sensing im- age retrieval.IEEE Transactions on Geoscience and Remote Sensing, 2026
Lixue Xu, Mingyang Li, Yimeng Fan, Tianle Yang, Guoxuan Qin, and Wei Zhang. Frequency-guided mutual distillation for lightweight cross-modal remote sensing im- age retrieval.IEEE Transactions on Geoscience and Remote Sensing, 2026
2026
-
[41]
Hongfeng Yu, Chubo Deng, Liangjin Zhao, Lingxiang Hao, Xiaoyu Liu, Wanxuan Lu, and Hongjian You. A light-weighted hypergraph neural network for multimodal remote sensing image retrieval.IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 16:2690–2...
2023
-
[42]
Weixun Zhou, Haiyan Guan, Ziyu Li, Zhenfeng Shao, and Mahmoud R. Delavar. Re- mote sensing image retrieval in the past decade: Achievements, challenges, and future directions.IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 16:1447–1473, 2023
2023
-
[43]
Lilu Zhu, Yang Wang, Yanfeng Hu, Xiaolu Su, and Kun Fu. Cross-modal contrastive learning with spatiotemporal context for correlation-aware multiscale remote sensing image retrieval.IEEE Transactions on Geoscience and Remote Sensing, 62:1–21, 2024. HOSSAIN ET AL.: CR-JEPA: CROS...
2024
Reviewed June 28, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.