REVIEW 5 major objections 6 minor 140 references
FrogDogNet: Fourier frequency Retained visual prompt Output Guidance for Domain Generalization of CLIP in Remote Sensing
T0 review · 5 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Filtering CLIP's image features to their strongest low-frequency Fourier components improves domain generalization in remote-sensing scene classification.
desk verdict A useful prompt-learning recipe for remote sensing DG with consistently strong numbers, but the core Fourier-filtering claim is never directly tested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Fourier Filter Block (FFB), which applies a fast Fourier transform to the processed image embedding, keeps the top k=350 low-frequency coefficients out of 512, and applies an inverse transform to reconstruct a filtered embedding. A projection network and a scaled self-attention module with residual connections refine the raw CLIP feature before filtering, and a lightweight Meta-Net then converts the filtered embedding into M visual tokens that are added to learnable text tokens. A Remote Sensing Prompt Alignment loss, minimized jointly with cross-entropy, pulls the learned text prompts toward CLIP's own remote-sensing text embeddings. The FFB's job is to make the visual prompt tokens depend on class-redundant low-frequency structure, which the authors identify as the invariant content that survives domain shift.
What would settle it
Replace the Fourier Filter Block in the published pipeline with a random mask that keeps 350 of the 512 coordinates, with a PCA projection to 350 dimensions, and with a version where the FFT is applied after permuting the feature dimensions. If accuracy on the three domain-generalization tasks stays within noise across these variants, the specific frequency ordering is not doing the work.
Extended reading notes
Core claim
The paper's central discovery is that truncating the discrete Fourier transform of CLIP's processed image features to the top 350 of 512 low-frequency coefficients—instead of using the full feature vector—yields visual prompt tokens that generalize better across remote-sensing domains. The authors argue that full-image features carry within-class variation from backgrounds, sensors, and atmospheric conditions, and that these artifacts concentrate in the high-frequency components of the embedding. Removing them before a lightweight Meta-Net generates prompt tokens, together with a self-attention stage that preserves boundary-level local cues, produces a prompt-learning pipeline that outperforms prior methods on base-to-new class generalization, cross-dataset transfer, and single-source multi-target domain generalization. The reported cost of the added machinery is small: 192.361 GFLOPS, nearly identical to CoCoOp and below APPLeNet.
Load-bearing premise
The load-bearing premise is that cutting a CLIP feature vector's Fourier transform down to its 350 largest low-frequency coefficients keeps class-relevant structure and discards noise and background artifacts; if an embedding's FFT has no meaningful frequency ordering, the filter is not the source of the gains.
Editorial extensions
If this is right
- If the reported numbers hold, the method reaches 76.64 average harmonic mean for base-to-new classes across the four datasets, 3.91 points above StyLIP and roughly 7.8 points above CoOp.
- Cross-dataset transfer from PatternNet improves over StyLIP by 6.98, 6.41, and 5.36 points on RSICD, RESISC45, and MLRSNet, respectively.
- Single-source multi-target generalization improves over StyLIP by 1.83, 2.34, and 5.04 points on the three target datasets while remaining within 0.05 points of StyLIP on the source.
- Prompt alignment is particularly valuable when train and test prompts differ: adding the Remote Sensing Prompt Alignment loss gives a 9.3% average gain in the cross-dataset setup with mismatched prompts.
- At 16 shots on PatternNet, the method reports an 85.63 harmonic mean, but its 32-shot result (77.28) is lower, so the benefit is not monotonic in the number of training examples.
Reading between the lines
- If the frequency hypothesis transfers, the same Fourier Filter Block recipe should improve other frozen vision-language encoders such as remote-sensing-specific CLIP variants; that is a direct test the paper does not run.
- The cutoff of 350 out of 512 coefficients is tuned on these four datasets; a practical deployment would need to verify whether that single cutoff survives new sensors or whether per-domain tuning is required.
- Because the prompt-alignment loss alone yields a 9.3% cross-dataset gain when prompts differ, an ablation that removes the FFB while keeping that loss would separate the contributions of frequency filtering from prompt alignment.
- The t-SNE comparison suggests better class separation, but that visual claim could be quantified with linear-probe accuracy or nearest-centroid distances on the filtered embeddings.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes FrogDogNet, a prompt-learning framework for CLIP-based remote sensing scene classification and domain generalization. The pipeline extracts CLIP visual features, refines them through a projection network and self-attention with residual connections, then passes them through a Fourier Filter Block (FFB) that retains the top k=350 low-frequency FFT coefficients of the 512-dimensional feature vector and applies an inverse FFT. A Meta-Net converts the filtered features into visual prompt tokens, and a Remote Sensing Prompt Alignment (RPA) loss aligns learned text prompts with RS-specific prompt initializations. Experiments on PatternNet, RSICD, RESISC45, and MLRSNet report state-of-the-art results on base-to-new class generalization, cross-dataset transfer, and single-source multi-target domain generalization, with an average base-to-new harmonic mean of 76.64 versus 72.73 for StyLIP.
Significance. If the reported gains are real and attributable to the FFB, the paper would make a useful contribution: it provides a comprehensive benchmark comparison across four RS datasets and three domain-generalization protocols, it includes a public code link, and its results are consistently above strong prompt-learning baselines such as StyLIP, APPLeNet, and MaPLe. The central hypothesis that low-frequency feature retention improves invariance to domain shift is interesting and testable. However, the current evidence does not yet isolate the FFB as the cause of the gains, the FFT-on-features mechanism lacks a clear semantic interpretation, and the reported results lack variance estimates and show non-monotonic behavior in the shot-response curve. These issues are fixable but need to be addressed before the central claim can be accepted.
major comments (5)
- [§4.2, Tables 1–3] The Fourier Filter Block is never ablated. Section 4.2 ablates the RPA loss, context length, Λ, backbone, and shot count, but no condition removes or replaces the FFB. Because the final model also contains the projection network, self-attention, residual scaling, Meta-Net, and RPA loss, the reported gains over StyLIP (3.91 average harmonic mean in Table 1, and 5.36–6.98 points in Table 2) cannot be attributed to frequency filtering. Please add an ablation with the FFB replaced by an identity mapping and by a fixed random mask of the same size, keeping everything else fixed.
- [§3.2.3, Figure 3] The conceptual basis of the FFB is underspecified. The DFT formulas in Eqs. (4)–(5) are given for image-like spatial inputs, but the block operates on the 512-dimensional CLIP visual embedding fv(x); it is not explained how low-frequency coefficients of this learned feature vector correspond to spatial or semantic frequency, why a hard threshold at k=350 is appropriate, or how the FFT is taken (real versus complex, and along which axis). The claim that low-frequency retention removes background artifacts and preserves class-relevant structure is an untested assumption; it is equally plausible that the operation acts as a fixed regularizing channel mask.
- [Figure 1(b), §4] The choice k=350 is made from a sensitivity analysis on the same generalization tasks and datasets used for the final comparison in Tables 1–3, with supporting details deferred to a missing supplementary. Selecting a hyperparameter on the evaluation tasks can inflate reported performance. Please provide a nested or held-out selection procedure, or report k sensitivity separately for each dataset and task so the reader can gauge the stability of the choice.
- [Table 4] The shot-response curve is non-monotonic: the harmonic mean rises to 85.63 at 16 shots and falls to 77.28 at 32 shots. This is surprising because more training data should not degrade performance, and it suggests that the reported operating point is not a stable optimum. Please investigate and report whether this drop persists with different seeds, schedulers, or the final hyperparameter settings used for Tables 1–3.
- [§4.1, Tables 1–3] All tables report averages over three random seeds but no standard deviations or confidence intervals. Given the small margins in several comparisons (for example, Table 1 RESISC45 base accuracy is 0.27 below StyLIP, and Table 3 source accuracy is 0.05 below StyLIP), the reader cannot assess whether the differences are significant. Please add variance estimates or per-seed results.
minor comments (6)
- [§3.3, Eqs. (18) and (20)] The argmin notation inside the definitions of Lce and Ltotal is not a loss definition; these equations should be written as minimization objectives over the parameters.
- [§3.2.2, Eq. (13)] The expression Xfinal = Xout/X + X divides by a matrix without defining the operation; please specify elementwise division and check the dimensions.
- [§4.2] Several analyses are deferred to a supplementary that is not included in the arXiv v1 manuscript (Figure 1(b) details, dataset splits, and hyperparameter trends on the remaining datasets); please include the supplementary or move the essential details into the main text.
- [Abstract and §1] The paper claims that full-image features introduce noise and background artifacts and that the FFB removes them, but no quantitative evidence directly supports this causal story beyond the t-SNE visualization; consider adding a direct comparison of filtered versus unfiltered features on a controlled noise or background perturbation.
- [Figure 1(a) vs Figure 3] Figure 1(a) illustrates frequency filtering on an image, whereas the FFB operates on the CLIP feature vector; the relationship between these two levels should be clarified to avoid confusing the reader about where the filtering actually happens.
- [Table 5] The GFLOPS comparison is only against CoCoOp and APPLeNet, and reporting a percentage difference with two decimals for a 0.009 GFLOPS gap is misleading without noting that such differences are likely within measurement noise.
Circularity Check
Partially circular validation: k=350 and other key hyperparameters are tuned on the same benchmark that is later reported as the headline result.
-
fitted input called prediction
[Section 1 (Figure 1(b) caption); Section 4 'Architecture'; Section 3.2.3]
"Our analysis in Figure 1(b) indicates that retaining 350 out of 512 LFCs yields the best average performance across multiple RS datasets and generalization tasks. We fixk = 350 out of 512 image features (based on Figure 1(b), detailed explanation is presented in supplementary material)."
The choice k=350 is not derived from a theory; it is the argmax of the same generalization tasks (B2N, CD, DG) and the same datasets (PatternNet, RSICD, RESISC45, MLRSNet) that Tables 1-3 use as the headline evidence. The statement that retaining 350 LFCs 'yields the best average performance' is therefore a selection summary, not a prediction. Re-reporting the same benchmark with k fixed at the selected value does not independently confirm that frequency filtering is the cause of the gains; the reported numbers are partly a selected maximum.
-
fitted input called prediction
[Section 4 'Training and evaluation'; Figures 5, 7, 8; Table 4]
"Based on extensive experiments, we set λ = 0.3 in Equation 12 and we set Λ = 0.5 in Equation 20, analysis shown in Figure 7. Following Figure 8, we use ViT-B/16 as the image encoder and train with 16 shots per class."
The remaining configuration (balancing weights λ, Λ, context length M=4, ViT-B/16 backbone, and the 16-shot setting) is selected from sensitivity analyses performed on the same PatternNet B2N task that supplies the headline harmonic-mean numbers in Table 1. The final architecture is thus tuned on the same benchmark used for the state-of-the-art comparison, making the configuration part of the benchmark fit rather than an independently derived design choice. This compounds the k=350 circularity.
full rationale
FrogDogNet contains no formal derivation chain; it is an empirical pipeline, so circularity must be assessed through its validation logic. The principal circular step is that the design-defining hyperparameter k (number of retained low-frequency coefficients) is selected from Figure 1(b), a sensitivity sweep over the same RS datasets and generalization tasks that Tables 1-3 later use as the headline evidence. The claim that 'keeping 350 out of 512 LFCs achieves the highest average generalization performance' is a selection, not a prediction; reporting performance on the same benchmark after conditioning on that selection does not independently confirm the frequency-filtering hypothesis. The same pattern applies to λ=0.3, Λ=0.5, M=4, the ViT-B/16 backbone, and 16 shots, all chosen from analyses on the same PatternNet B2N task. I did not find load-bearing self-citation: StyLIP and APPLeNet appear as baselines, and the Meta-Net and the RPA loss cite external prior work. A separate, non-circular but important support gap is that Section 4.2 ablates RPA loss, context length, Λ, backbone, and shot count but never ablates or replaces the Fourier Filter Block, so even the non-circular parts of the comparison cannot attribute the gains specifically to frequency filtering. The supporting details for Figure 1(b) are deferred to a missing supplementary, limiting auditability. Because the architecture still contains trainable components and the SOTA comparison is not logically forced by k=350 alone, this is partial circularity rather than a complete reduction, hence a score of 6 rather than 8-10.
Assumptions & free parameters
free parameters (5)
- k (number of retained low-frequency components) =
350 out of 512
- lambda (mixing weight in Eq. 12) =
0.3
- Lambda (weight of LRPA in Eq. 20) =
0.5
- Z (number of RS prompt variants) =
4
- M (prompt context length) =
4
assumptions (3)
- domain assumption CLIP's frozen image and text encoders provide semantically meaningful representations for remote sensing scene classification.
- ad hoc to paper Applying the DFT to the CLIP visual feature vector fv(x) and keeping the top 350 of 512 low-frequency coefficients removes noise and background artifacts while preserving class-relevant structure.
- domain assumption The Meta-Net design from CoCoOp, which maps a single visual feature vector to M prompt tokens, remains effective when its input is replaced by the filtered features psi.
Cite this review
Pith. "Pith review of FrogDogNet: Fourier frequency Retained visual prompt Output Guidance for Domain Generalization of CLIP in Remote Sensing." pith.science (2026). https://pith.science/paper/J7PMF2R7
@misc{pith2026250416433,
author = {Pith},
title = {Pith review of: FrogDogNet: Fourier frequency Retained visual prompt Output Guidance for Domain Generalization of CLIP in Remote Sensing},
year = {2026},
howpublished = {\url{https://pith.science/paper/J7PMF2R7}},
note = {Machine review of arXiv:2504.16433}
}
read the original abstract
In recent years, large-scale vision-language models (VLMs) like CLIP have gained attention for their zero-shot inference using instructional text prompts. While these models excel in general computer vision, their potential for domain generalization in remote sensing (RS) remains underexplored. Existing approaches enhance prompt learning by generating visual prompt tokens but rely on full-image features, introducing noise and background artifacts that vary within a class, causing misclassification. To address this, we propose FrogDogNet, a novel prompt learning framework integrating Fourier frequency filtering and self-attention to improve RS scene classification and domain generalization. FrogDogNet selectively retains invariant low-frequency components while eliminating noise and irrelevant backgrounds, ensuring robust feature representation across domains. The model first extracts significant features via projection and self-attention, then applies frequency-based filtering to preserve essential structural information for prompt learning. Extensive experiments on four RS datasets and three domain generalization tasks show that FrogDogNet consistently outperforms state-of-the-art prompt learning methods, demonstrating superior adaptability across domain shifts. Our findings highlight the effectiveness of frequency-based invariant feature retention in generalization, paving the way for broader applications. Our code is available at https://github.com/HariseetharamG/FrogDogNet
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Abdullah, Y
T. Abdullah, Y . Bazi, M. M. Al Rahhal, M. L. Mekhalfi, L. Rangarajan, and M. Zuair. Textrs: Deep bidirectional triplet network for matching text to remote sensing images. Remote Sensing, 12(3):405, 2020. 3
2020
-
[2]
Alhichri, Nassim Ammour, and Naif Alajlan
Dalal Alajaji, Haikel S. Alhichri, Nassim Ammour, and Naif Alajlan. Few-shot learning for remote sensing scene classification. In 2020 Mediterranean and Middle-East Geoscience and Remote Sensing Symposium (M2GARSS) , pages 81–84, 2020. 2
2020
-
[3]
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision , pages 2425–2433, 2015. 3
2015
-
[4]
Metareg: Towards domain generalization using meta-regularization
Yogesh Balaji, Swami Sankaranarayanan, and Rama Chel- lappa. Metareg: Towards domain generalization using meta-regularization. Advances in neural information pro- cessing systems, 31, 2018. 2
2018
-
[5]
Y . Bazi, M. M. A. Rahhal, M. L. Mekhalfi, M. A. Al Zuair, and F. Melgani. Bi-modal transformer-based approach for visual question answering in remote sensing imagery.IEEE Transactions on Geoscience and Remote Sensing, 60:1–11,
-
[6]
C-saw: Self-supervised prompt learning for image generalization in remote sensing
Avigyan Bhattacharya, Mainak Singha, Ankit Jha, and Bi- plab Banerjee. C-saw: Self-supervised prompt learning for image generalization in remote sensing. In Proceedings of the Fourteenth Indian Conference on Computer Vision, Graphics and Image Processing, pages 1–10, 2023. 2, 3, 7
2023
-
[7]
StyLIP: Multi-Scale Style-Conditioned Prompt Learning for CLIP-based Domain Generalization
Shirsha Bose, Ankit Jha, Enrico Fini, Mainak Singha, Elisa Ricci, and Biplab Banerjee. Stylip: Multi-scale style- conditioned prompt learning for clip-based domain general- ization. arXiv preprint arXiv:2302.09251, 2024. Accepted in W ACV 2024. 2, 3, 7
work page Pith review arXiv 2024
-
[8]
Domain-controlled prompt learning
Qinglong Cao, Zhengqin Xu, Yuntian Chen, Chao Ma, and Xiaokang Yang. Domain-controlled prompt learning. In Proceedings of the AAAI Conference on Artificial Intelli- gence, pages 936–944, 2024. 3
2024
Show all 140 references
-
[9]
Domain prompt learning with quaternion networks
Qinglong Cao, Zhengqin Xu, Yuntian Chen, Chao Ma, and Xiaokang Yang. Domain prompt learning with quaternion networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 26637– 26646, 2024. 3
2024
-
[10]
Chappuis, V
C. Chappuis, V . Mendez, E. Walt, S. Lobry, B. Le Saux, and D. Tuia. Language transformers for remote sensing visual question answering. In IGARSS 2022-2022 IEEE Interna- tional Geoscience and Remote Sensing Symposium , pages 4855–4858. IEEE, 2022. 3
2022
-
[11]
Learning to balance specificity and invariance for in and out of domain generalization
Prithvijit Chattopadhyay, Yogesh Balaji, and Judy Hoff- man. Learning to balance specificity and invariance for in and out of domain generalization. In European Conference on Computer Vision, pages 301–318. Springer, 2020. 2
2020
-
[12]
Rsprompter: Learning to prompt for remote sensing instance segmenta- tion based on visual foundation model
Keyan Chen, Chenyang Liu, Hao Chen, Haotian Zhang, Wenyuan Li, Zhengxia Zou, and Zhenwei Shi. Rsprompter: Learning to prompt for remote sensing instance segmenta- tion based on visual foundation model. IEEE Transactions on Geoscience and Remote Sensing, 2024. 3
2024
-
[13]
Meta relational learning for few- shot link prediction in knowledge graphs
Mingyang Chen, Wen Zhang, Wei Zhang, Qiang Chen, and Huajun Chen. Meta relational learning for few- shot link prediction in knowledge graphs. arXiv preprint arXiv:1909.01515, 2019. 2
1909 arXiv
-
[14]
Y . Chen, C. Wei, D. Wang, C. Ji, and B. Li. Semi- supervised contrastive learning for few-shot segmentation of remote sensing images. Remote Sensing, 14(17):4254,
-
[15]
Remote sens- ing image scene classification: Benchmark and state of the art
Gong Cheng, Junwei Han, and Xiaoqiang Lu. Remote sens- ing image scene classification: Benchmark and state of the art. Proceedings of the IEEE, 105(10):1865–1883, 2017. 6
2017
-
[16]
Cheng, Y
Q. Cheng, Y . Zhou, P. Fu, Y . Xu, and L. Zhang. A deep se- mantic alignment network for the cross-modal image-text retrieval in remote sensing. IEEE Journal of Selected Top- ics in Applied Earth Observations and Remote Sensing, 14: 4284–4297, 2021. 3
2021
-
[17]
J. W. Cooley and J. W. Tukey. An algorithm for the ma- chine calculation of complex fourier series. Mathematics of Computation, 19(90):297–301, 1965. 4
1965
-
[18]
Multimodal artificial intelligence foundation models: Unleashing the power of remote sensing big data in earth observation
MRSB DATA. Multimodal artificial intelligence foundation models: Unleashing the power of remote sensing big data in earth observation. Innovation, 2(1):100055, 2024. 1
2024
-
[20]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 3
2018 arXiv
-
[21]
S. Dong, L. Wang, B. Du, and X. Meng. Changeclip: Remote sensing change detection with multimodal vision- language representation learning. ISPRS Journal of Pho- togrammetry and Remote Sensing, 208:53–69, 2024. 2
2024
-
[23]
An im- age is worth 16x16 words: Transformers for image recog- nition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An im- age is worth 16x16 words: Transformers for image recog- nitio...
2010 arXiv
-
[24]
Domain generalization via model- agnostic learning of semantic features
Qi Dou, Daniel Coelho de Castro, Konstantinos Kamnit- sas, and Ben Glocker. Domain generalization via model- agnostic learning of semantic features. Advances in Neural Information Processing Systems, 32, 2019. 2
2019
-
[25]
Domain generalized object detection for remote sensing images
Efkan Duraklı and Erchan Aptoula. Domain generalized object detection for remote sensing images. In 2023 31st Signal Processing and Communications Applications Con- ference (SIU), pages 1–4, 2023. 1
2023
-
[26]
A brief review of domain adaptation.Ad- vances in data science and information engineering, pages 877–894, 2021
Abolfazl Farahani, Sahar V oghoei, Khaled Rasheed, and Hamid R Arabnia. A brief review of domain adaptation.Ad- vances in data science and information engineering, pages 877–894, 2021. 1
2021
-
[28]
Model- agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model- agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning , pages 1126–1135. PMLR, 2017. 2
2017
-
[29]
Unsupervised do- main adaptation by backpropagation
Yaroslav Ganin and Victor Lempitsky. Unsupervised do- main adaptation by backpropagation. In International Con- ference on Machine Learning , pages 1180–1189. PMLR,
-
[30]
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pas- cal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky. Domain-adversarial training of neural networks. The journal of machine learn- ing research, 17(1):2096–2030, 2016. 7
2016
-
[31]
Clip-adapter: Better vision-language models with feature adapters
Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao. Clip-adapter: Better vision-language models with feature adapters. arXiv preprint arXiv:2110.04544, 2021. 2, 3, 7
2021 arXiv
-
[32]
Self-supervised in-domain representation learning for remote sensing image scene classification
Ali Ghanbarzadeh and Hossein Soleimani. Self-supervised in-domain representation learning for remote sensing image scene classification. Heliyon, 10(19):e37962, 2024. 1
2024
-
[33]
Visual-language prompt tuning with knowledge-guided context optimiza- tion
Changsheng Xu Hantao Yao, Rui Zhang. Visual-language prompt tuning with knowledge-guided context optimiza- tion. In The IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 5
2023
-
[34]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceed- ings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 3
2016
-
[35]
Machine learning for environmental monitoring
Miyuki Hino, Elinor Benami, and Nina Brooks. Machine learning for environmental monitoring. Nature Sustainabil- ity, 1(10):583–588, 2018. 1
2018
-
[36]
On the generalization ability of a global model for rapid building mapping from heteroge- neous satellite images of multiple natural disaster scenarios
Yijiang Hu and Hong Tang. On the generalization ability of a global model for rapid building mapping from heteroge- neous satellite images of multiple natural disaster scenarios. Remote Sensing, 13(5):984, 2021. 1
2021
-
[37]
Y . Hu, J. Yuan, C. Wen, X. Lu, and X. Li. Rsgpt: A re- mote sensing vision language model and benchmark. arXiv preprint arXiv:2307.15266, 2023. 3
2023 arXiv
-
[38]
Self-challenging improves cross-domain generalization
Zeyi Huang, Haohan Wang, Eric P Xing, and Dong Huang. Self-challenging improves cross-domain generalization. In European Conference on Computer Vision, pages 124–140. Springer, 2020. 2
2020
-
[39]
On the real capabilities of re- mote sensing for disaster management-feedback from real cases
Jordi Inglada and Alain Giros. On the real capabilities of re- mote sensing for disaster management-feedback from real cases. In IGARSS 2004. 2004 IEEE International Geo- science and Remote Sensing Symposium, pages 1110–1112. IEEE, 2004. 1
2004
-
[40]
Rs3lip: Consistency for remote sensing im- age classification on part embeddings using self-supervised learning and clip
Ankit Jha, Mainak Singha, Avigyan Bhattacharya, and Bi- plab Banerjee. Rs3lip: Consistency for remote sensing im- age classification on part embeddings using self-supervised learning and clip. Computer Vision and Image Understand- ing, 251:104254, 2025. 3, 7
2025
-
[41]
Few-shot scene classification of optical re- mote sensing images leveraging calibrated pretext tasks
Hong Ji, Zhi Gao, Yongjun Zhang, Yu Wan, Can Li, and Tiancan Mei. Few-shot scene classification of optical re- mote sensing images leveraging calibrated pretext tasks. IEEE Transactions on Geoscience and Remote Sensing, 60: 1–13, 2022. 2
2022
-
[42]
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In In- ternational Conference on Machine Learning, pages 4904–
-
[43]
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In Pro- ceedings of the 38th International Conference on Machine ...
2021
-
[44]
Jiang, N
X. Jiang, N. Zhou, and X. Li. Few-shot segmentation of remote sensing images using deep metric learning. IEEE Geoscience and Remote Sensing Letters, 19:1–5, 2022. 3
2022
-
[45]
Maple: Multi-modal prompt learning
Muhammad Uzair Khattak, Hanoona Rasheed, Muham- mad Maaz, Salman Khan, and Fahad Shahbaz Khan. Maple: Multi-modal prompt learning. arXiv preprint arXiv:2210.03117, 2023. Accepted at CVPR 2023. 2, 3, 7
2023 arXiv
-
[46]
Efficient prompt tuning of large vision-language model for fine-grained ship classification
Long Lan, Fengxiang Wang, Xiangtao Zheng, Zengmao Wang, and Xinwang Liu. Efficient prompt tuning of large vision-language model for fine-grained ship classification. IEEE Transactions on Geoscience and Remote Sensing ,
-
[47]
A. Li, Z. Lu, L. Wang, T. Xiang, and J.-R. Wen. Zero-shot scene classification for high spatial resolution remote sens- ing images. IEEE Transactions on Geoscience and Remote Sensing, 55(7):4157–4167, 2017. 3
2017
-
[48]
Learning to generalize: Meta-learning for do- main generalization
Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy Hospedales. Learning to generalize: Meta-learning for do- main generalization. In Proceedings of the AAAI conference on artificial intelligence, 2018. 1, 2
2018
-
[49]
Episodic training for domain generalization
Da Li, Jianshu Zhang, Yongxin Yang, Cong Liu, Yi-Zhe Song, and Timothy M Hospedales. Episodic training for domain generalization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1446– 1455, 2019. 2
2019
-
[50]
Patternnet: Visual pattern mining with deep neural network
Hongzhi Li, Joseph G Ellis, Lei Zhang, and Shih-Fu Chang. Patternnet: Visual pattern mining with deep neural network. In Proceedings of the 2018 ACM on international confer- ence on multimedia retrieval, pages 291–299, 2018. 6
2018
-
[51]
Domain generalization with adversarial feature learn- ing
Haoliang Li, Sinno Jialin Pan, Shiqi Wang, and Alex C Kot. Domain generalization with adversarial feature learn- ing. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5400–5409, 2018. 2
2018
-
[52]
Junnan Li, Dong Li, Silvio Savarese, and Steven C. Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In In- ternational conference on machine learning, pages 19730– 19742. PMLR, 2023. 3
2023
-
[53]
Pro- gressive domain expansion network for single domain gen- eralization
Lei Li, Ke Gao, Juan Cao, Ziyao Huang, Yepeng Weng, Xi- aoyue Mi, Zhengze Yu, Xiaoya Li, and Boyang Xia. Pro- gressive domain expansion network for single domain gen- eralization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 224–233,
-
[54]
Visualbert: A simple and perfor- mant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. Visualbert: A simple and perfor- mant baseline for vision and language. arXiv preprint arXiv:1908.03557, 2019. 2
1908 arXiv
-
[55]
X. Li, X. Zhang, W. Huang, and Q. Wang. Truncation cross entropy loss for remote sensing image captioning. IEEE Transactions on Geoscience and Remote Sensing , 59(6): 5246–5257, 2020. 3
2020
-
[56]
RS- CLIP: Zero-Shot Remote Sensing Scene Classification via Contrastive Vision-Language Supervision
Xiang Li, Congcong Wen, Yuan Hu, and Nan Zhou. RS- CLIP: Zero-Shot Remote Sensing Scene Classification via Contrastive Vision-Language Supervision. International Journal of Applied Earth Observation and Geoinformation, 124:103497, 2023. 2
2023
-
[57]
Vision-language models in remote sens- ing: Current progress and future trends
Xiang Li, Congcong Wen, Yuan Hu, Zhenghang Yuan, and Xiao Xiang Zhu. Vision-language models in remote sens- ing: Current progress and future trends. IEEE Geoscience and Remote Sensing Magazine, 2024. 3
2024
-
[58]
Deep do- main generalization via conditional invariant adversarial networks
Ya Li, Xinmei Tian, Mingming Gong, Yajing Liu, Tongliang Liu, Kun Zhang, and Dacheng Tao. Deep do- main generalization via conditional invariant adversarial networks. In Proceedings of the European conference on computer vision (ECCV), pages 624–639, 2018. 2
2018
-
[59]
Feature-critic networks for heterogeneous do- main generalization
Yiying Li, Yongxin Yang, Wei Zhou, and Timothy Hospedales. Feature-critic networks for heterogeneous do- main generalization. In International Conference on Ma- chine Learning, pages 3915–3924. PMLR, 2019. 2
2019
-
[60]
Y . Li, S. Fang, L. Jiao, R. Liu, and R. Shang. A multi-level attention model for remote sensing image captions. Remote Sensing, 12(6):939, 2020. 3
2020
-
[61]
Z. Li, D. Zhang, Y . Wang, D. Lin, and J. Zhang. Gen- erative adversarial networks for zero-shot remote sensing scene classification. Applied Sciences, 12(8):3760, 2022. 3
2022
-
[62]
Chen, and Bing Xu
Ziming Li, Bin Chen, Shengbiao Wu, Mo Su, Jing M. Chen, and Bing Xu. Deep learning for urban land use category classification: A review and experimental assessment. Re- mote Sensing of Environment, 311:114290, 2024. 1
2024
-
[63]
Fedrsclip: Federated learning for remote sensing scene classification using vision-language models
Hui Lin, Chao Zhang, Danfeng Hong, Kexin Dong, and Congcong Wen. Fedrsclip: Federated learning for remote sensing scene classification using vision-language models. arXiv preprint arXiv:2501.02461, 2025. 3
2025 arXiv
-
[64]
Deep few-shot learning for hyper- spectral image classification
Bing Liu, Xuchu Yu, Anzhu Yu, Pengqiang Zhang, Gang Wan, and Ruirui Wang. Deep few-shot learning for hyper- spectral image classification. IEEE Transactions on Geo- science and Remote Sensing, 57(4):2290–2304, 2018. 2
2018
-
[65]
Re- moteCLIP: A Vision-Language Foundation Model for Re- mote Sensing
Fan Liu, Delong Chen, Zhangqingyun Guan, Xiaocong Zhou, Jiale Zhu, Qiaolin Ye, Liyong Fu, and Jun Zhou. Re- moteCLIP: A Vision-Language Foundation Model for Re- mote Sensing. IEEE Transactions on Geoscience and Re- mote Sensing, 62:5622216, 2024. 2
2024
-
[66]
Lobry, D
S. Lobry, D. Marcos, J. Murray, and D. Tuia. Rsvqa: Visual question answering for remote sensing data. IEEE Trans- actions on Geoscience and Remote Sensing , 58(12):8555– 8566, 2020. 3
2020
-
[67]
Exploring models and data for remote sensing im- age caption generation
Xiaoqiang Lu, Binqiang Wang, Xiangtao Zheng, and Xue- long Li. Exploring models and data for remote sensing im- age caption generation. IEEE Transactions on Geoscience and Remote Sensing, 56(4):2183–2195, 2017. 3, 6
2017
-
[68]
X. Lu, X. Sun, W. Diao, Y . Mao, J. Li, Y . Zhang, P. Wang, and K. Fu. Few-shot object detection in aerial imagery guided by text-modal knowledge. IEEE Transactions on Geoscience and Remote Sensing, 2023. 3
2023
-
[69]
Deep learning in remote sensing applications: A meta-analysis and review.ISPRS Journal of Photogrammetry and Remote Sensing, 152:166–177, 2019
Lei Ma, Yu Liu, Xueliang Zhang, Yuanxin Ye, Gaofei Yin, and Brian Alan Johnson. Deep learning in remote sensing applications: A meta-analysis and review.ISPRS Journal of Photogrammetry and Remote Sensing, 152:166–177, 2019. 1
2019
-
[70]
Martini, V
M. Martini, V . Mazzia, A. Khaliq, and M. Chiaberge. Domain-adversarial training of self-attention-based net- works for land cover classification using multi-temporal sentinel-2 satellite imagery. Remote Sensing, 13(13):2564,
-
[71]
Domain general- ization using a mixture of multiple latent domains
Toshihiko Matsuura and Tatsuya Harada. Domain general- ization using a mixture of multiple latent domains. In Pro- ceedings of the AAAI conference on artificial intelligence , pages 11749–11756, 2020. 2
2020
-
[72]
Integrating remote sensing and demog- raphy for more efficient and effective assessment of chang- ing mountain forest distribution
Peter J Morley, Daniel NM Donoghue, Jan-Chang Chen, and Alistair S Jump. Integrating remote sensing and demog- raphy for more efficient and effective assessment of chang- ing mountain forest distribution. Ecological Informatics, 43:106–115, 2018. 1
2018
-
[73]
Spn: Stable prototypical network for few-shot learning-based hyperspectral image classification
Debabrata Pal, Valay Bundele, Biplab Banerjee, and Yo- gananda Jeppu. Spn: Stable prototypical network for few-shot learning-based hyperspectral image classification. IEEE Geoscience and Remote Sensing Letters , 19:1–5,
-
[74]
Relevant and in- variant feature selection of hyperspectral images for domain generalization
Claudio Persello and Lorenzo Bruzzone. Relevant and in- variant feature selection of hyperspectral images for domain generalization. In 2014 IEEE Geoscience and Remote Sens- ing Symposium, pages 3562–3565. IEEE, 2014. 2
2014
-
[75]
Language models as knowledge bases? arXiv preprint arXiv:1909.01066, 2019
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. Language models as knowledge bases? arXiv preprint arXiv:1909.01066, 2019. 3
1909 arXiv
-
[76]
Mlrsnet: A multi- label high spatial resolution remote sensing dataset for se- mantic scene understanding
Xiaoman Qi, Panpan Zhu, Yuebin Wang, Liqiang Zhang, Junhuan Peng, Mengfan Wu, Jialong Chen, Xudong Zhao, Ning Zang, and P Takis Mathiopoulos. Mlrsnet: A multi- label high spatial resolution remote sensing dataset for se- mantic scene understanding. ISPRS Journal of Photogram- ...
2020
-
[77]
Uncertainty-guided model generalization to unseen domains
Fengchun Qiao and Xi Peng. Uncertainty-guided model generalization to unseen domains. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6790–6800, 2021. 2
2021
-
[78]
Learning to learn single domain generalization
Fengchun Qiao, Long Zhao, and Xi Peng. Learning to learn single domain generalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12556–12565, 2020. 2
2020
-
[79]
J. Quan, C. Wu, H. Wang, and Z. Wang. Structural alignment based zero-shot classification for remote sensing scenes. In 2018 IEEE International Conference on Elec- tronics and Communication Engineering (ICECE) , pages 17–21. IEEE, 2018. 3
2018
-
[80]
Improving language understanding by gen- erative pre-training
Alec Radford. Improving language understanding by gen- erative pre-training. 2018. 3
2018
-
[81]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In International Conference on Machine Learning...
2021
-
[82]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In International Conference on Machine Learning...
2021
-
[83]
M. M. A. Rahhal, Y . Bazi, T. Abdullah, M. L. Mekhalfi, and M. Zuair. Deep unsupervised embedding for remote sens- ing image retrieval using textual cues. Applied Sciences, 10 (24):8931, 2020. 3
2020
-
[84]
M. M. A. Rahhal, Y . Bazi, S. O. Alsaleh, M. Al-Razgan, M. L. Mekhalfi, M. Al Zuair, and N. Alajlan. Open-ended remote sensing visual question answering with transform- ers. International Journal of Remote Sensing, 43(18):6809– 6823, 2022. 3
2022
-
[85]
M. M. A. Rahhal, Y . Bazi, N. A. Alsharif, L. Bashmal, N. Alajlan, and F. Melgani. Multilanguage transformer for im- proved text to remote sensing image retrieval. IEEE Jour- nal of Selected Topics in Applied Earth Observations and Remote Sensing, 15:9115–9126, 2022. 3
2022
-
[86]
M. M. A. Rahhal, M. A. Bencherif, Y . Bazi, A. Alharbi, and M. L. Mekhalfi. Contrasting dual transformer architectures for multi-modal remote sensing image retrieval. Applied Sciences, 13(1):282, 2023. 3
2023
-
[87]
Global filter networks for image classification
Yuyu Rao, Wenliang Zhao, Zhen Zhu, Jiwen Lu, and Jie Zhou. Global filter networks for image classification. In Proceedings of the Advances in Neural Information Pro- cessing Systems (NeurIPS), pages 980–993, 2021. 4
2021
-
[88]
A stochastic approx- imation method
Herbert Robbins and Sutton Monro. A stochastic approx- imation method. The annals of mathematical statistics , pages 400–407, 1951. 6
1951
-
[89]
Meta-learning for few-shot land cover classifica- tion
Marc Russwurm, Sherrie Wang, Marco Korner, and David Lobell. Meta-learning for few-shot land cover classifica- tion. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR) Workshops ,
-
[90]
Floyd F. Sabins. Remote sensing for mineral exploration. Ore Geology Reviews, 14(3-4):157–183, 1999. 1
1999
-
[91]
Multi- target domain adaptation for remote sensing classification using graph neural network
Sudipan Saha, Shan Zhao, and Xiao Xiang Zhu. Multi- target domain adaptation for remote sensing classification using graph neural network. IEEE Geoscience and Remote Sensing Letters, 19:1–5, 2022. 1
2022
-
[92]
Shi and Z
Z. Shi and Z. Zou. Can a machine generate humanlike lan- guage descriptions for a remote sensing image? IEEE Transactions on Geoscience and Remote Sensing , 55(6): 3623–3634, 2017. 3
2017
-
[93]
Autoprompt: Eliciting knowl- edge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV , Eric Wallace, and Sameer Singh. Autoprompt: Eliciting knowl- edge from language models with automatically generated prompts. arXiv preprint arXiv:2010.15980, 2020. 3
2010 arXiv
-
[94]
Test- time prompt tuning for zero-shot generalization in vision- language models
Manli Shu, Weili Nie, De-An Huang, Zhiding Yu, Tom Goldstein, Anima Anandkumar, and Chaowei Xiao. Test- time prompt tuning for zero-shot generalization in vision- language models. arXiv preprint arXiv:2209.07511, 2022. 3
2022 arXiv
-
[95]
Applenet: Visual attention parameterized prompt learning for few-shot remote sens- ing image generalization using clip
Mainak Singha, Ankit Jha, Bhupendra Solanki, Shirsha Bose, and Biplab Banerjee. Applenet: Visual attention parameterized prompt learning for few-shot remote sens- ing image generalization using clip. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...
2023
-
[96]
Improved deep metric learning with multi- class n-pair loss objective
Kihyuk Sohn. Improved deep metric learning with multi- class n-pair loss objective. In Advances in Neural Informa- tion Processing Systems, 2016. 3
2016
-
[97]
Sumbul, R
G. Sumbul, R. G. Cinbis, and S. Aksoy. Fine-grained ob- ject recognition and zero-shot learning in remote sensing imagery. IEEE Transactions on Geoscience and Remote Sensing, 56(2):770–779, 2017. 3
2017
-
[98]
Intriguing properties of neural networks
C Szegedy. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199, 2013. 2
2013 arXiv
-
[99]
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek. The information bottleneck method. arXiv preprint physics/0004057, 2000. 2
2000 arXiv
-
[100]
Do- main adaptation for the classification of remote sensing data: An overview of recent advances
Devis Tuia, Claudio Persello, and Lorenzo Bruzzone. Do- main adaptation for the classification of remote sensing data: An overview of recent advances. IEEE Geoscience and Remote Sensing Magazine, 4(2):41–57, 2016. 1
2016
-
[101]
Toward a collective agenda on ai for earth science data analysis
Devis Tuia, Ribana Roscher, Jan Dirk Wegner, Narun Ja- cobs, Xiao Xiang Zhu, and Gustau Camps-Valls. Toward a collective agenda on ai for earth science data analysis. IEEE Geoscience and Remote Sensing Magazine, 9(2):88– 104, 2021. 3
2021
-
[102]
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research , 9 (11), 2008. 6
2008
-
[103]
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research , 9 (11), 2008. 7
2008
-
[104]
Remote sensing for natural disaster man- agement
CJ Van Westen. Remote sensing for natural disaster man- agement. International archives of photogrammetry and re- mote sensing, 33(B7/4; PART 7):1609–1617, 2000. 1
2000
-
[105]
Vladimir N. Vapnik. Statistical Learning Theory . Wiley- Interscience, 1998. 7
1998
-
[106]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Ilya Polosukhin. Attention is all you need. In Ad- vances in neural information processing systems , pages 5998–6008, 2017. 2
2017
-
[107]
Show and tell: A neural image caption gen- erator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Du- mitru Erhan. Show and tell: A neural image caption gen- erator. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3156–3164, 2015. 3
2015
-
[108]
Generalizing to unseen domains via adversarial data augmentation
Riccardo V olpi, Hongseok Namkoong, Ozan Sener, John C Duchi, Vittorio Murino, and Silvio Savarese. Generalizing to unseen domains via adversarial data augmentation. Ad- vances in neural information processing systems, 31, 2018. 2
2018
-
[109]
C. Wang, G. Peng, and B. De Baets. A distance-constrained semantic autoencoder for zero-shot remote sensing scene classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 14:12545–12556,
-
[110]
Metasegnet: Metadata- collaborative vision-language representation learning for semantic segmentation of remote sensing images
Libo Wang, Sijun Dong, Ying Chen, Xiaoliang Meng, Shenghui Fang, and Songlin Fei. Metasegnet: Metadata- collaborative vision-language representation learning for semantic segmentation of remote sensing images. IEEE Transactions on Geoscience and Remote Sensing, 2024. 3
2024
-
[111]
Q. Wang, W. Huang, X. Zhang, and X. Li. Word–sentence framework for remote sensing image captioning. IEEE Transactions on Geoscience and Remote Sensing , 59(12): 10532–10543, 2020. 3
2020
-
[112]
Learning to diversify for single domain generalization
Zijian Wang, Yadan Luo, Ruihong Qiu, Zi Huang, and Mahsa Baktashmotlagh. Learning to diversify for single domain generalization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 834– 843, 2021. 2
2021
-
[113]
A fourier-based framework for domain generaliza- tion
Qinwei Xu, Ruipeng Zhang, Ya Zhang, Yanfeng Wang, and Qi Tian. A fourier-based framework for domain generaliza- tion. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 14383–14392,
-
[114]
Simde: A simple domain expan- sion approach for single-source domain generalization
Qinwei Xu, Ruipeng Zhang, Yi-Yan Wu, Ya Zhang, Ning Liu, and Yanfeng Wang. Simde: A simple domain expan- sion approach for single-source domain generalization. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 4798–4808, 2023. 2
2023
-
[115]
S3net: Spectral–spatial siamese network for few-shot hyperspec- tral image classification
Zhaohui Xue, Yiyang Zhou, and Peijun Du. S3net: Spectral–spatial siamese network for few-shot hyperspec- tral image classification. IEEE Transactions on Geoscience and Remote Sensing, 60:1–19, 2022. 2
2022
-
[116]
Domain-aware generalized meta-learning for space target recognition
Xi Yang, Dechen Kong, Dong Yang, and Miao Wang. Domain-aware generalized meta-learning for space target recognition. IEEE Transactions on Geoscience and Remote Sensing, 62:1–12, 2024. 1
2024
-
[117]
Sar-optical image matching model based on contrastive learning
Xiong Yaoting. Sar-optical image matching model based on contrastive learning. In 2023 20th International Com- puter Conference on Wavelet Active Media Technology and Information Processing (ICCWAMTIP), pages 1–5, 2023. 1
2023
-
[118]
Florence: A new foundation model for computer vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021. 2
2021 arXiv
-
[119]
Z. Yuan, W. Zhang, X. Rong, X. Li, J. Chen, H. Wang, K. Fu, and X. Sun. A lightweight multi-scale crossmodal text-image retrieval method in remote sensing.IEEE Trans- actions on Geoscience and Remote Sensing, 60:1–19, 2021. 3
2021
-
[120]
Z. Yuan, L. Mou, Q. Wang, and X. X. Zhu. From easy to hard: Learning language-guided curriculum for visual question answering on remote sensing data. IEEE Transac- tions on Geoscience and Remote Sensing , 60:1–11, 2022. 3
2022
-
[121]
Z. Yuan, W. Zhang, K. Fu, X. Li, C. Deng, H. Wang, and X. Sun. Exploring a fine-grained multiscale method for cross- modal remote sensing image retrieval. IEEE Transactions on Geoscience and Remote Sensing, 60:3078451, 2022. 3
2022
-
[122]
Z. Yuan, W. Zhang, C. Tian, X. Rong, Z. Zhang, H. Wang, K. Fu, and X. Sun. Remote sensing cross-modal text- image retrieval based on global and local information.IEEE Transactions on Geoscience and Remote Sensing, 60:1–16,
-
[123]
Lit: Zero-shot transfer with locked-image text tuning
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer. Lit: Zero-shot transfer with locked-image text tuning. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition , pages 18123–18133, 2022. 2
2022
-
[124]
Few-shot classification of aerial scene images via meta- learning
Pei Zhang, Yunpeng Bai, Dong Wang, Bendu Bai, and Ying Li. Few-shot classification of aerial scene images via meta- learning. Remote Sensing, 13(1):108, 2021. 2
2021
-
[125]
Zhang, F
S. Zhang, F. Song, X. Liu, X. Hao, Y . Liu, T. Lei, and P. Jiang. Text semantic fusion relation graph reasoning for few-shot object detection on remote sensing images. Re- mote Sensing, 15(5):1187, 2023. 3
2023
-
[126]
Segclip: Multimodal visual- language and prompt learning for high-resolution remote sensing semantic segmentation.IEEE Transactions on Geo- science and Remote Sensing, 2024
Shijie Zhang, Bin Zhang, Yuntao Wu, Huabing Zhou, Jun- jun Jiang, and Jiayi Ma. Segclip: Multimodal visual- language and prompt learning for high-resolution remote sensing semantic segmentation.IEEE Transactions on Geo- science and Remote Sensing, 2024. 3
2024
-
[127]
Zhang, X
X. Zhang, X. Li, J. An, L. Gao, B. Hou, and C. Li. Natu- ral language description of remote sensing images based on deep learning. In 2017 IEEE International Geoscience and Remote Sensing Symposium (IGARSS) , pages 4798–4801. IEEE, 2017. 3
2017
-
[128]
Zhang, X
X. Zhang, X. Wang, X. Tang, H. Zhou, and C. Li. Descrip- tion generation for remote sensing images using attribute attention mechanism. Remote Sensing, 11(6):612, 2019. 3
2019
-
[129]
Zhang, M
Y . Zhang, M. Zhang, W. Li, S. Wang, and R. Tao. Language-aware domain generalization network for cross- scene hyperspectral image classification. IEEE Transac- tions on Geoscience and Remote Sensing, 2023. 3
2023
-
[130]
Zhang, T
Z. Zhang, T. Zhao, Y . Guo, and J. Yin. Rs5m: A large scale vision-language dataset for remote sensing vision-language foundation model. arXiv preprint arXiv:2306.11300, 2023. 3
2023 arXiv
-
[131]
Maximum-entropy adversarial data augmentation for im- proved generalization and robustness
Long Zhao, Ting Liu, Xi Peng, and Dimitris Metaxas. Maximum-entropy adversarial data augmentation for im- proved generalization and robustness. Advances in Neural Information Processing Systems, 33:14435–14447, 2020. 2
2020
-
[132]
R. Zhao, Z. Shi, and Z. Zou. High-resolution remote sens- ing image captioning based on structured attention. IEEE Transactions on Geoscience and Remote Sensing, 60:1–14,
-
[133]
Multisource-domain generalization- based oil palm tree detection using very-high-resolution (vhr) satellite images
Juepeng Zheng, Wenzhao Wu, Shuai Yuan, Haohuan Fu, Weijia Li, and Le Yu. Multisource-domain generalization- based oil palm tree detection using very-high-resolution (vhr) satellite images. IEEE Geoscience and Remote Sens- ing Letters, 19:1–5, 2021. 1, 2
2021
-
[134]
Zheng, B
X. Zheng, B. Wang, X. Du, and X. Lu. Mutual attention inception network for remote sensing visual question an- swering. IEEE Transactions on Geoscience and Remote Sensing, 60:1–14, 2021. 3
2021
-
[135]
Deep domain-adversarial image generation for domain generalisation
Kaiyang Zhou, Yongxin Yang, Timothy Hospedales, and Tao Xiang. Deep domain-adversarial image generation for domain generalisation. In Proceedings of the AAAI Confer- ence on Artificial Intelligence , pages 13025–13032, 2020. 2
2020
-
[136]
Domain generalization with mixstyle
Kaiyang Zhou, Yongxin Yang, Yu Qiao, and Tao Xi- ang. Domain generalization with mixstyle. arXiv preprint arXiv:2104.02008, 2021. 2
2021 arXiv
-
[137]
Domain generalization: A survey
Kaiyang Zhou, Ziwei Liu, Yu Qiao, Tao Xiang, and Chen Change Loy. Domain generalization: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence,
-
[138]
Conditional prompt learning for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Zi- wei Liu. Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 16816– 16825, 2022. 2, 3, 5, 6, 7, 8
2022
-
[139]
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Zi- wei Liu. Learning to prompt for vision-language models. International Journal of Computer Vision , 130(9):2337– 2348, 2022. 2, 3, 6, 7
2022
-
[140]
Prompt-aligned gradient for prompt tuning
Beier Zhu, Yulei Niu, Yucheng Han, Yue Wu, and Hanwang Zhang. Prompt-aligned gradient for prompt tuning. arXiv preprint arXiv:2205.14865, 2022. 2, 3, 7
2022 arXiv
-
[141]
pytorch-opcounter, 2018
Ligeng Zhu. pytorch-opcounter, 2018. Maintained by Lyken17. Available at: pytorch-OpCounter. 8
2018
-
[142]
Deep learning in remote sensing: A comprehensive review and list of resources
Xiao Xiang Zhu, David Tuia, and Lorenzo Mou et al. Deep learning in remote sensing: A comprehensive review and list of resources. IEEE Geoscience and Remote Sensing Magazine, 5(4):8–36, 2017. 1
2017
-
[143]
U. Zia, M. M. Riaz, and A. Ghafoor. Transforming re- mote sensing images to textual descriptions. International Journal of Applied Earth Observation and Geoinformation, 108:102741, 2022. 3
2022
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.