REVIEW 3 major objections 5 minor 1 cited by
PAN-Crafter: Learning Modality-Consistent Alignment for PAN-Sharpening
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that pansharpening can be made robust to cross-modality spatial misalignment by training one U-Net to reconstruct both the fused multispectral image and its own panchromatic input, with local attention conditioned on real…
desk verdict Strong empirical pansharpening paper with a clean architecture, but the misalignment story is not supported by the ablations; still deserves a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the MARs mode switch combined with CM3A (Cross-Modality Alignment-Aware Attention). MARs duplicates each training triplet and runs the same U-Net in two modes—MS mode predicts a residual to the HRMS image, while PAN mode predicts a residual to a channel-replicated, downsampled-then-upsampled PAN image—forcing a single shared network to represent both modalities. In CM3A, the query tensor concatenates a down-sampled real image of the target modality with the current feature, the key and value tensors come from the down-sampled other-modality image concatenated with the same feature, and attention is evaluated inside a $k \times k$ local window. This makes the alignment signal data-dependent and local, so the network can compensate for shifts without estimating a global displacement. The final HRMS prediction is a residual added to the upsampled LRMS image.
What would settle it
A controlled test that applies known pixel shifts, for example 1, 3, 5, and 8 pixels, to the PAN image before fusion and measures whether the quality gap between PAN-Crafter and its stronger U-Net baseline grows with shift size would settle whether the alignment modules, rather than the baseline architecture, carry the measured gains.
Extended reading notes
Core claim
The central claim is that explicit, learned cross-modality alignment, rather than bigger reconstruction losses, is what resolves double edges and spectral distortion in pansharpening. PAN-Crafter jointly reconstructs the HRMS image in “MS mode” and a multi-channel version of the PAN image in “PAN mode,” so the same U-Net learns to move high-frequency PAN detail into the MS output while keeping spectral fidelity. In both modes, CM3A computes local attention in which the query is conditioned on a down-sampled real image of one modality and the key/value pairs on the other, letting the network adapt to whatever local displacement the sensor geometry produced. On the PanCollection benchmarks, the paper reports the best scores on nearly every metric, at 50.11× lower inference cost than the closest non-diffusion competitor and hundreds of times lower than diffusion models, and it transfers zero-shot to the unseen WorldView-2 satellite.
Load-bearing premise
The evaluation assumes that the residual misalignment present in the standard PanCollection benchmark images is representative of the cross-modality misalignment the method is designed to solve, and no controlled test with synthetic shifts or misaligned ground truth is provided.
Editorial extensions
If this is right
- If PAN-Crafter is correct, pansharpening can be done with a single feed-forward U-Net at inference, eliminating the iterative diffusion loop entirely.
- The zero-shot results on WorldView-2 suggest that a model trained on one satellite’s PAN-MS statistics can be deployed on another sensor without retraining or fine-tuning.
- Because MARs uses PAN reconstruction as auxiliary supervision, training no longer depends exclusively on having perfectly aligned HRMS ground truth; mildly misaligned pairs still provide a usable learning signal.
- Replacing fixed positional embeddings with down-sampled modality images as attention priors offers a recipe for other cross-modality fusion tasks where local geometric mismatch is the main failure mode.
Reading between the lines
- The paper does not isolate misalignment severity; a controlled test with synthetic shifts would clarify how much of the gain comes from explicit alignment handling versus the stronger U-Net baseline that the authors themselves note is already competitive.
- Applying the MARs/CM3A recipe to other modality pairs with local spatial mismatch, such as SAR-to-optical or depth-to-RGB fusion, is a natural next step the paper does not explore.
- The fixed local window ($k=3$) implies that performance may degrade for larger misalignments; making the window size adaptive per image or per scale could extend the method’s robustness.
- The paper openly notes that its full-resolution spectral metric is slightly lower on one benchmark because the method aligns the output to the LRMS image rather than the PAN image, so downstream users should check whether their application prioritizes PAN consistency or MS consistency.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes PAN-Crafter, a U-Net-based pansharpening framework whose two main components are Modality-Adaptive Reconstruction (MARs), which jointly trains the network to reconstruct both HRMS and multi-channel PAN images with a modality-conditioned modulation layer, and Cross-Modality Alignment-Aware Attention (CM3A), a local attention mechanism that replaces fixed positional embeddings with downsampled PAN/MS image features in the query/key construction. The method is evaluated on PanCollection benchmarks (WV3, GF2, QB, and a zero-shot WV2 test) using standard reduced- and full-resolution metrics, and it reports competitive or best results in most metrics with substantially lower inference time than diffusion-based baselines. The central claims are that PAN-Crafter explicitly mitigates cross-modality misalignment and that it outperforms the most recent state-of-the-art method in all metrics.
Significance. If the central claims hold, PAN-Crafter would be an attractive practical alternative to diffusion-based pansharpening: it reports strong quantitative results, large inference speedups (50.11x over CANConv), good cross-sensor zero-shot performance on WV2, and stable mean/std statistics in the supplementary. The paper follows a standard comparison protocol, uses official implementations of baselines where available, and provides unusually thorough ablations (Tables 11-17) including sensitivity to the attention kernel size and two-stage training. However, the specific attribution of the gains to explicit misalignment handling is not yet established: the ablation tables show that CM3A alone is marginal or even negative on some datasets, and the supplementary concedes that the U-Net baseline without MARs/CM3A is already competitive with prior state-of-the-art. The significance of the work would be considerably strengthened by a controlled misalignment experiment that isolates the geometric-alignment effect.
major comments (3)
- [Abstract and Section 4.3, Tables 6-8] The abstract and Section 4.3 claim that PAN-Crafter 'outperforms the most recent state-of-the-art method in all metrics', but this is contradicted by the paper's own tables. On WV3, Table 6 shows PAN-Crafter D_lambda = 0.016 +/- 0.006 versus PanDiff's 0.014 +/- 0.005; on QB, Table 8 shows PAN-Crafter D_s = 0.039 +/- 0.020 versus LAGConv's 0.035 +/- 0.009, D_lambda = 0.043 +/- 0.011 versus PanDiff's 0.028 +/- 0.011, and SAM = 4.426 +/- 0.740 versus DCPNet's 4.420 +/- 0.710. The text in Section 4.3 itself admits the WV3 D_lambda and QB D_s/SAM limitations, so the abstract-level claim should be corrected to 'most metrics' or the specific exceptions should be stated.
- [Section 4.4 and Supplementary Tables 11-13] The central attribution of the method's success to explicit cross-modality misalignment handling is not supported by the evidence. CM3A alone changes WV3 HQNR from 0.948 to 0.949 (Table 11) and decreases GF2 HQNR from 0.959 to 0.953 (Table 12); only on QB does CM3A alone give a clear gain (0.856 to 0.879, Table 13), and the actual misalignment level of that dataset is not characterized. Since Section 2.2 states that benchmark PAN-MS pairs are 'generally pre-aligned', the residual shifts in the test sets are unmeasured. The paper should include a controlled experiment with synthetic shifts (e.g., known pixel offsets between PAN and MS) or an estimated misalignment map per dataset, and show how HQNR/PSNR vary with shift magnitude, to substantiate the claim that CM3A performs geometric alignment rather than acting as a generic local attention mechanism.
- [Supplementary C.3 and main text tables] The supplementary admits that the U-Net without MARs/CM3A is already competitive with prior state-of-the-art (e.g., GF2 PSNR 43.476 versus CANConv's 43.166; WV3 HQNR 0.948 versus PanDiff's 0.952), yet this 'U-Net' baseline is not included in the main comparison tables. Because Section 4.4 and the supplementary also show that MARs is the component providing the largest gain (WV3 HQNR 0.948->0.956; GF2 0.959->0.945 for MARs alone, with the combination at 0.964), the reader cannot separate the contribution of the proposed alignment mechanism from that of the stronger architecture and auxiliary reconstruction loss. The main tables should include the U-Net+MARs variant (without CM3A) and the full baseline without either component, so that the incremental value of CM3A can be assessed directly against the stated design goal.
minor comments (5)
- [Supplementary B.1] The limitation that inter-band misalignment is not handled should be mentioned in the abstract or introduction, since the framework is described as addressing 'cross-modality' alignment in general terms; a scope clarification would prevent overgeneralization.
- [Section 4.2 and Table 15] Only the local attention kernel size k is ablated; the MARs loss weight lambda is fixed at 1.0 without sensitivity analysis, despite the text saying 'tuning lambda ensures' a balance; a small study or a robustness note would justify this choice.
- [Figure 3 caption] The PAN-mode path is described as 'Down-sampled and x4 Up-sampled Ipan' followed by 'Repeated Cms times', which is confusing because the figure shows both a down-sampled replica and an up-sampled version; please clarify the exact flow and notation in the caption or main text.
- [Section 3.3, Eqs. (9)-(12)] The notation Irep,down_pan and Ilr,down_ms is introduced without explicit definitions before first use; a one-sentence definition of these down-sampled/replicated tensors would improve reproducibility.
- [Section 4.4] There is a typo, 'evalution metric', that should be corrected.
Circularity Check
No significant circularity: the method is evaluated on external PanCollection benchmarks with official baseline implementations, and self-citations are background only.
full rationale
This paper proposes a supervised pansharpening architecture (a U-Net with CM3A local cross-attention and MARs joint HRMS/PAN reconstruction) and validates it on the external PanCollection benchmark. No step of the claimed derivation chain reduces to its own inputs by construction. The training objective (Eq. 4) is a standard L1 reconstruction loss against ground-truth HRMS and PAN targets, and the reported quality metrics are computed with the official PanCollection repository while all baselines were trained using their official codebases under the same protocol (Supp. A.2), so the comparison numbers are not manufactured by the authors. Hyperparameters (lambda = 1.0, k = 3, C = 128, seed 2025) are fixed before evaluation and are not fitted to test targets, so no fitted input is renamed as a prediction. The auxiliary PAN back-reconstruction is an internal training regularizer and is explicitly labeled as self-supervision; its target is derived from the input PAN, but it is never reported as a benchmark result or as the central contribution. Self-citations ([22] SIPSA, [21] U-Know-DiffPan, [11] C-DiffSet) appear only in related-work context and carry no load-bearing weight for the SOTA or generalization claims. The skeptic's concerns - no controlled synthetic-shift experiment, CM3A alone giving marginal or negative HQNR in Tables 11-13, and the abstract's 'all metrics' claim contradicting the paper's own tables - are evidence-attribution and consistency issues (does CM3A cause the measured gains?) rather than circularity: the reported numbers are not forced by a fit or by a self-citation chain. Those concerns belong under correctness risk, leaving the circularity score at 0.
Assumptions & free parameters
free parameters (3)
- MARs loss weight lambda =
1.0
- CM3A local attention kernel size k =
3
- Network width C =
128
assumptions (3)
- domain assumption Paired PAN and LRMS images in PanCollection share the same underlying scene and are approximately aligned up to local residual shifts.
- domain assumption The PAN back-reconstruction task is a valid auxiliary supervision signal for the HRMS reconstruction.
- domain assumption Standard stochastic training with random crops, flips, rotations, and AdamW generalizes to the test distributions of PanCollection benchmarks.
Cite this review
Pith. "Pith review of PAN-Crafter: Learning Modality-Consistent Alignment for PAN-Sharpening." pith.science (2026). https://pith.science/paper/PXF6J5IT
@misc{pith2026250523367,
author = {Pith},
title = {Pith review of: PAN-Crafter: Learning Modality-Consistent Alignment for PAN-Sharpening},
year = {2026},
howpublished = {\url{https://pith.science/paper/PXF6J5IT}},
note = {Machine review of arXiv:2505.23367}
}
abstract
PAN-sharpening aims to fuse high-resolution panchromatic (PAN) images with low-resolution multi-spectral (MS) images to generate high-resolution multi-spectral (HRMS) outputs. However, cross-modality misalignment -- caused by sensor placement, acquisition timing, and resolution disparity -- induces a fundamental challenge. Conventional deep learning methods assume perfect pixel-wise alignment and rely on per-pixel reconstruction losses, leading to spectral distortion, double edges, and blurring when misalignment is present. To address this, we propose PAN-Crafter, a modality-consistent alignment framework that explicitly mitigates the misalignment gap between PAN and MS modalities. At its core, Modality-Adaptive Reconstruction (MARs) enables a single network to jointly reconstruct HRMS and PAN images, leveraging PAN's high-frequency details as auxiliary self-supervision. Additionally, we introduce Cross-Modality Alignment-Aware Attention (CM3A), a novel mechanism that bidirectionally aligns MS texture to PAN structure and vice versa, enabling adaptive feature refinement across modalities. Extensive experiments on multiple benchmark datasets demonstrate that our PAN-Crafter outperforms the most recent state-of-the-art method in all metrics, even with 50.11$\times$ faster inference time and 0.63$\times$ the memory size. Furthermore, it demonstrates strong generalization performance on unseen satellite datasets, showing its robustness across different conditions.
Figures
Figures from the paper (9 more)
Forward citations
Cited by 1 Pith paper
-
MetaTele: Compact Refractive Metasurface Computational Telephoto Camera
Hybrid refractive-metasurface optics plus one-step diffusion fusion of structure and color measurements yields RGB telephoto imaging at telephoto ratio 0.44 with 13 mm TTL.
Reference graph
Works this paper leans on
-
[1]
Bruno Aiazzi, Luciano Alparone, Stefano Baronti, and An- drea Garzelli. Context-driven fusion of high spatial and spec- tral resolution images based on oversampled multiresolution analysis. IEEE Transactions on geoscience and remote sens- ing, 40(10):2300–2312, 2002. 2
work page 2002
-
[2]
Mtf-tailored multiscale fusion of high-resolution ms and pan imagery
Bruno Aiazzi, Luciano Alparone, Stefano Baronti, Andrea Garzelli, and Massimo Selva. Mtf-tailored multiscale fusion of high-resolution ms and pan imagery. Photogrammetric Engineering & Remote Sensing, 72(5):591–596, 2006. 2
work page 2006
-
[3]
A variational model for p+ xs image fusion
Coloma Ballester, Vicent Caselles, Laura Igual, Joan Verdera, and Bernard Roug ´e. A variational model for p+ xs image fusion. International Journal of Computer Vision, 69:43–58, 2006. 2
work page 2006
-
[4]
Hy- pertransformer: A textural and spectral feature fusion trans- former for pansharpening
Wele Gedara Chaminda Bandara and Vishal M Patel. Hy- pertransformer: A textural and spectral feature fusion trans- former for pansharpening. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 1767–1777, 2022. 2
work page 2022
-
[5]
Proximal pan- net: A model-based deep network for pansharpening
Xiangyong Cao, Yang Chen, and Wenfei Cao. Proximal pan- net: A model-based deep network for pansharpening. InPro- ceedings of the AAAI conference on artificial intelligence , pages 176–184, 2022. 2
work page 2022
-
[6]
Wjoseph Carper, Thomasm Lillesand, Ralphw Kiefer, et al. The use of intensity-hue-saturation transformations for merging spot panchromatic and multispectral image data. Photogrammetric Engineering and remote sensing , 56(4): 459–467, 1990. 2
work page 1990
-
[7]
A new adaptive component-substitution-based satellite image fusion by us- ing partial replacement
Jaewan Choi, Kiyun Yu, and Yongil Kim. A new adaptive component-substitution-based satellite image fusion by us- ing partial replacement. IEEE transactions on geoscience and remote sensing, 49(1):295–309, 2010. 2
work page 2010
-
[8]
Deformable convolutional networks
Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. In Proceedings of the IEEE international confer- ence on computer vision, pages 764–773, 2017. 3
work page 2017
Show all 65 references
-
[9]
Machine learning in pan- sharpening: A benchmark, from shallow to deep networks
Liang-Jian Deng, Gemine Vivone, Mercedes E Paoletti, Giuseppe Scarpa, Jiang He, Yongjun Zhang, Jocelyn Chanussot, and Antonio Plaza. Machine learning in pan- sharpening: A benchmark, from shallow to deep networks. IEEE Geoscience and Remote Sensing Magazine , 10(3): 279–315, 2...
2022
-
[10]
Yinyang k-means: A drop-in replace- ment of the classic k-means with consistent speedup
Yufei Ding, Yue Zhao, Xipeng Shen, Madanlal Musuvathi, and Todd Mytkowicz. Yinyang k-means: A drop-in replace- ment of the classic k-means with consistent speedup. In In- ternational conference on machine learning, pages 579–587. PMLR, 2015. 7, 9
2015
-
[11]
C-diffset: Leveraging latent diffusion for sar-to-eo image translation with confidence-guided reliable object generation
Jeonghyeok Do, Jaehyup Lee, and Munchurl Kim. C-diffset: Leveraging latent diffusion for sar-to-eo image translation with confidence-guided reliable object generation. arXiv preprint arXiv:2411.10788, 2024. 1
2024 arXiv
-
[12]
An image is worth 16x16 words: Trans- formers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. arXiv preprint a...
2010 arXiv
-
[13]
Content-adaptive non-local convolution for remote sensing pansharpening
Yule Duan, Xiao Wu, Haoyu Deng, and Liang-Jian Deng. Content-adaptive non-local convolution for remote sensing pansharpening. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 27738– 27747, 2024. 2, 3, 6, 7, 8, 9, 13, 14
2024
-
[14]
A variational pan-sharpening with local gradient constraints
Xueyang Fu, Zihuang Lin, Yue Huang, and Xinghao Ding. A variational pan-sharpening with local gradient constraints. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10265–10274, 2019. 2
2019
-
[15]
Hypercomplex quality assessment of multi/hyperspectral images
Andrea Garzelli and Filippo Nencini. Hypercomplex quality assessment of multi/hyperspectral images. IEEE Geoscience and Remote Sensing Letters, 6(4):662–665, 2009. 6
2009
-
[16]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 3
2020
-
[17]
Bidomain modeling paradigm for pan- sharpening
Junming Hou, Qi Cao, Ran Ran, Che Liu, Junling Li, and Liang-jian Deng. Bidomain modeling paradigm for pan- sharpening. In Proceedings of the 31st ACM International Conference on Multimedia, pages 347–357, 2023. 2
2023
-
[18]
Linearly-evolved transformer for pan- sharpening
Junming Hou, Zihan Cao, Naishan Zheng, Xuan Li, Xi- aoyu Chen, Xinyang Liu, Xiaofeng Cong, Danfeng Hong, and Man Zhou. Linearly-evolved transformer for pan- sharpening. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 1486–1494, 2024. 2
2024
-
[19]
Robust reference-based super-resolution via c2-matching
Yuming Jiang, Kelvin CK Chan, Xintao Wang, Chen Change Loy, and Ziwei Liu. Robust reference-based super-resolution via c2-matching. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , pages 2103–2112, 2021. 3
2021
-
[20]
Lagconv: Local-context adap- tive convolution kernels with global harmonic bias for pan- sharpening
Zi-Rong Jin, Tian-Jing Zhang, Tai-Xiang Jiang, Gemine Vivone, and Liang-Jian Deng. Lagconv: Local-context adap- tive convolution kernels with global harmonic bias for pan- sharpening. In Proceedings of the AAAI conference on arti- ficial intelligence, pages 1113–1121, 2022. 2,...
2022
-
[21]
U-know-diffpan: An uncertainty-aware knowledge dis- tillation diffusion framework with details enhancement for pan-sharpening
Sungpyo Kim, Jeonghyeok Do, Jaehyup Lee, and Munchurl Kim. U-know-diffpan: An uncertainty-aware knowledge dis- tillation diffusion framework with details enhancement for pan-sharpening. arXiv preprint arXiv:2412.06243, 2024. 3
2024 arXiv
-
[22]
Sipsa-net: Shift-invariant pan sharpening with moving object alignment for satellite imagery
Jaehyup Lee, Soomin Seo, and Munchurl Kim. Sipsa-net: Shift-invariant pan sharpening with moving object alignment for satellite imagery. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 10166–10174, 2021. 2, 3
2021
-
[23]
A review of remote sensing for envi- ronmental monitoring in china
Jun Li, Yanqiu Pei, Shaohua Zhao, Rulin Xiao, Xiao Sang, and Chengye Zhang. A review of remote sensing for envi- ronmental monitoring in china. Remote Sensing, 12(7):1130,
-
[24]
Siamtrans: zero-shot multi-frame image restoration with pre-trained siamese transformers
Lin Liu, Shanxin Yuan, Jianzhuang Liu, Xin Guo, Youliang Yan, and Qi Tian. Siamtrans: zero-shot multi-frame image restoration with pre-trained siamese transformers. In Pro- ceedings of the AAAI Conference on Artificial Intelligence , pages 1747–1755, 2022. 3, 10, 12
2022
-
[25]
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In 16 Proceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021. 3
2021
-
[26]
Hyperspectral pansharpening: A review
Laetitia Loncan, Luis B De Almeida, Jos ´e M Bioucas- Dias, Xavier Briottet, Jocelyn Chanussot, Nicolas Dobigeon, Sophie Fabre, Wenzhi Liao, Giorgio A Licciardi, Miguel Simoes, et al. Hyperspectral pansharpening: A review. IEEE Geoscience and remote sensing magazine, 3(3):27–4...
2015
-
[27]
Sgdr: Stochas- tic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter. Sgdr: Stochas- tic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016. 6
2016 arXiv
-
[28]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 6
2017 arXiv
-
[29]
Pansharpening by convolutional neural networks
Giuseppe Masi, Davide Cozzolino, Luisa Verdoliva, and Giuseppe Scarpa. Pansharpening by convolutional neural networks. Remote Sensing, 8(7):594, 2016. 2
2016
-
[30]
Pan- diff: A novel pansharpening method based on denoising diffusion probabilistic model
Qingyan Meng, Wenxu Shi, Sijia Li, and Linlin Zhang. Pan- diff: A novel pansharpening method based on denoising diffusion probabilistic model. IEEE Transactions on Geo- science and Remote Sensing, 61:1–17, 2023. 2, 3, 6, 7, 8, 9, 13, 14
2023
-
[31]
A large-scale benchmark data set for evaluating pansharpening performance: Overview and im- plementation
Xiangchao Meng, Yiming Xiong, Feng Shao, Huanfeng Shen, Weiwei Sun, Gang Yang, Qiangqiang Yuan, Randi Fu, and Hongyan Zhang. A large-scale benchmark data set for evaluating pansharpening performance: Overview and im- plementation. IEEE Geoscience and Remote Sensing Maga- zine,...
2020
-
[32]
Mapping of urban vegeta- tion with high-resolution remote sensing: A review
Robbe Neyns and Frank Canters. Mapping of urban vegeta- tion with high-resolution remote sensing: A review. Remote sensing, 14(4):1031, 2022. 1
2022
-
[33]
An introduction to convo- lutional neural networks
Keiron O’shea and Ryan Nash. An introduction to convo- lutional neural networks. arXiv preprint arXiv:1511.08458,
-
[34]
Introduction of sensor spectral response into image fusion methods
Xavier Otazu, Mar ´ıa Gonz ´alez-Aud´ıcana, Octavi Fors, and Jorge N ´u˜nez. Introduction of sensor spectral response into image fusion methods. application to wavelet-based meth- ods. IEEE Transactions on Geoscience and Remote Sensing, 43(10):2376–2385, 2005. 2
2005
-
[35]
Slide-transformer: Hierarchical vision transformer with local self-attention
Xuran Pan, Tianzhu Ye, Zhuofan Xia, Shiji Song, and Gao Huang. Slide-transformer: Hierarchical vision transformer with local self-attention. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 2082–2091, 2023. 3, 5, 10
2023
-
[36]
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Al- ban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. 2017. 6
2017
-
[37]
Fusionmamba: Efficient remote sensing im- age fusion with state space model
Siran Peng, Xiangyu Zhu, Haoyu Deng, Liang-Jian Deng, and Zhen Lei. Fusionmamba: Efficient remote sensing im- age fusion with state space model. IEEE Transactions on Geoscience and Remote Sensing, 2024. 2
2024
-
[38]
An adaptive ihs pan-sharpening method
Sheida Rahmani, Melissa Strait, Daria Merkurjev, Michael Moeller, and Todd Wittman. An adaptive ihs pan-sharpening method. IEEE Geoscience and Remote Sensing Letters, 7(4): 746–750, 2010. 2
2010
-
[39]
An efficient pan-sharpening method via a combined adaptive pca approach and contourlets
Vijay P Shah, Nicolas H Younan, and Roger L King. An efficient pan-sharpening method via a combined adaptive pca approach and contourlets. IEEE transactions on geoscience and remote sensing, 46(5):1323–1335, 2008. 2
2008
-
[40]
The discrete wavelet transform: wedding the a trous and mallat algorithms
Mark J Shensa. The discrete wavelet transform: wedding the a trous and mallat algorithms. IEEE Transactions on signal processing, 40(10):2464–2482, 1992. 2
1992
-
[41]
Pixel-adaptive convolutional neural networks
Hang Su, Varun Jampani, Deqing Sun, Orazio Gallo, Erik Learned-Miller, and Jan Kautz. Pixel-adaptive convolutional neural networks. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , pages 11166–11175, 2019. 3
2019
-
[42]
Synthesis of multispectral images to high spa- tial resolution: A critical review of fusion methods based on remote sensing physics
Claire Thomas, Thierry Ranchin, Lucien Wald, and Jocelyn Chanussot. Synthesis of multispectral images to high spa- tial resolution: A critical review of fusion methods based on remote sensing physics. IEEE Transactions on Geoscience and Remote Sensing, 46(5):1301–1312, 2008. 2, 9
2008
-
[43]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 3, 10
2017
-
[44]
A critical comparison among pansharpening algorithms
Gemine Vivone, Luciano Alparone, Jocelyn Chanussot, Mauro Dalla Mura, Andrea Garzelli, Giorgio A Licciardi, Rocco Restaino, and Lucien Wald. A critical comparison among pansharpening algorithms. IEEE Transactions on Geoscience and Remote Sensing, 53(5):2565–2586, 2014. 6, 9
2014
-
[45]
A new benchmark based on recent advances in multispectral pansharpening: Revisit- ing pansharpening with classical and emerging pansharpen- ing methods
Gemine Vivone, Mauro Dalla Mura, Andrea Garzelli, Rocco Restaino, Giuseppe Scarpa, Magnus O Ulfarsson, Luciano Alparone, and Jocelyn Chanussot. A new benchmark based on recent advances in multispectral pansharpening: Revisit- ing pansharpening with classical and emerging pansh...
2020
-
[46]
Lucien Wald. Quality of high resolution synthesised images: Is there a simple criterion? In Third conference” Fusion of Earth data: merging point measurements, raster maps and remotely sensed images”, pages 99–103. SEE/URISCA,
-
[47]
Ssconv: Explicit spectral-to-spatial convolution for pan- sharpening
Yudong Wang, Liang-Jian Deng, Tian-Jing Zhang, and Xiao Wu. Ssconv: Explicit spectral-to-spatial convolution for pan- sharpening. In Proceedings of the 29th ACM international conference on multimedia, pages 4472–4480, 2021. 2
2021
-
[48]
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Si- moncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004. 6
2004
-
[49]
Remote sensing in urban planning: Contributions towards ecologically sound policies? Landscape and urban plan- ning, 204:103921, 2020
Thilo Wellmann, Angela Lausch, Erik Andersson, Sonja Knapp, Chiara Cortinovis, Jessica Jache, Sebastian Scheuer, Peleg Kremer, Andr ´e Mascarenhas, Roland Kraemer, et al. Remote sensing in urban planning: Contributions towards ecologically sound policies? Landscape and urban p...
2020
-
[50]
Dynamic cross feature fusion for remote sensing pan- sharpening
Xiao Wu, Ting-Zhu Huang, Liang-Jian Deng, and Tian-Jing Zhang. Dynamic cross feature fusion for remote sensing pan- sharpening. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14687–14696, 2021. 2, 6, 7, 8, 13, 14
2021
-
[51]
V o+ net: An adaptive ap- 17 proach using variational optimization and deep learning for panchromatic sharpening
Zhong-Cheng Wu, Ting-Zhu Huang, Liang-Jian Deng, Jin- Fan Hu, and Gemine Vivone. V o+ net: An adaptive ap- 17 proach using variational optimization and deep learning for panchromatic sharpening. IEEE Transactions on Geoscience and Remote Sensing, 60:1–16, 2021. 2
2021
-
[52]
Empower generaliz- ability for pansharpening through text-modulated diffusion model
Yinghui Xing, Litao Qu, Shizhou Zhang, Jiapeng Feng, Xiuwei Zhang, and Yanning Zhang. Empower generaliz- ability for pansharpening through text-modulated diffusion model. IEEE Transactions on Geoscience and Remote Sens- ing, 2024. 3, 6, 7, 8, 9, 13, 14
2024
-
[53]
Ai security for geo- science and remote sensing: Challenges and future trends
Yonghao Xu, Tao Bai, Weikang Yu, Shizhen Chang, Pe- ter M Atkinson, and Pedram Ghamisi. Ai security for geo- science and remote sensing: Challenges and future trends. IEEE Geoscience and Remote Sensing Magazine, 11(2):60– 85, 2023. 1
2023
-
[54]
Learning texture transformer network for image super-resolution
Fuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu, and Bain- ing Guo. Learning texture transformer network for image super-resolution. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 5791–5800, 2020. 3, 10, 12
2020
-
[55]
Memory-augmented deep conditional unfolding network for pan-sharpening
Gang Yang, Man Zhou, Keyu Yan, Aiping Liu, Xueyang Fu, and Fan Wang. Memory-augmented deep conditional unfolding network for pan-sharpening. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1788–1797, 2022. 2
2022
-
[56]
The role of satellite remote sensing in climate change studies
Jun Yang, Peng Gong, Rong Fu, Minghua Zhang, Jing- ming Chen, Shunlin Liang, Bing Xu, Jiancheng Shi, and Robert Dickinson. The role of satellite remote sensing in climate change studies. Nature climate change, 3(10):875– 883, 2013. 1
2013
-
[57]
Pannet: A deep network architecture for pan-sharpening
Junfeng Yang, Xueyang Fu, Yuwen Hu, Yue Huang, Xinghao Ding, and John Paisley. Pannet: A deep network architecture for pan-sharpening. InProceedings of the IEEE international conference on computer vision, pages 5449–5457, 2017. 2, 6, 7, 8, 13, 14
2017
-
[58]
A multiscale and multidepth convolutional neural network for remote sensing imagery pan-sharpening
Qiangqiang Yuan, Yancong Wei, Xiangchao Meng, Huan- feng Shen, and Liangpei Zhang. A multiscale and multidepth convolutional neural network for remote sensing imagery pan-sharpening. IEEE Journal of Selected Topics in Ap- plied Earth Observations and Remote Sensing , 11(3):978...
2018
-
[59]
Deep learning in environmen- tal remote sensing: Achievements and challenges
Qiangqiang Yuan, Huanfeng Shen, Tongwen Li, Zhiwei Li, Shuwen Li, Yun Jiang, Hongzhang Xu, Weiwei Tan, Qian- qian Yang, Jiwen Wang, et al. Deep learning in environmen- tal remote sensing: Achievements and challenges. Remote sensing of Environment, 241:111716, 2020. 1
2020
-
[60]
Discrimination among semi-arid landscape endmem- bers using the spectral angle mapper (sam) algorithm
Roberta H Yuhas, Alexander FH Goetz, and Joe W Board- man. Discrimination among semi-arid landscape endmem- bers using the spectral angle mapper (sam) algorithm. In JPL, Summaries of the Third Annual JPL Airborne Geo- science Workshop. Volume 1: AVIRIS Workshop, 1992. 6
1992
-
[61]
Spatial-spectral dual back- projection network for pansharpening
Kai Zhang, Anfei Wang, Feng Zhang, Wenbo Wan, Jiande Sun, and Lorenzo Bruzzone. Spatial-spectral dual back- projection network for pansharpening. IEEE Transactions on Geoscience and Remote Sensing, 61:1–16, 2023. 2, 6, 7, 8, 13, 14
2023
-
[62]
Dcpnet: a dual-task collaborative promotion network for pansharpening
Yafei Zhang, Xuji Yang, Huafeng Li, Minghong Xie, and Zhengtao Yu. Dcpnet: a dual-task collaborative promotion network for pansharpening. IEEE Transactions on Geo- science and Remote Sensing , 62:1–16, 2024. 2, 6, 7, 8, 9, 13, 14
2024
-
[63]
Learning raw-to-srgb mappings with inaccurately aligned supervision
Zhilu Zhang, Haolin Wang, Ming Liu, Ruohao Wang, Jiawei Zhang, and Wangmeng Zuo. Learning raw-to-srgb mappings with inaccurately aligned supervision. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 4348–4358, 2021. 3
2021
-
[64]
Ssdiff: Spatial-spectral integrated diffusion model for remote sensing pansharpening
Yu Zhong, Xiao Wu, Liang-Jian Deng, Zihan Cao, and Hong-Xia Dou. Ssdiff: Spatial-spectral integrated diffusion model for remote sensing pansharpening. Advances in Neu- ral Information Processing Systems, 37:77962–77986, 2024. 3
2024
-
[65]
Probability-based global cross-modal upsam- pling for pansharpening
Zeyu Zhu, Xiangyong Cao, Man Zhou, Junhao Huang, and Deyu Meng. Probability-based global cross-modal upsam- pling for pansharpening. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 14039–14048, 2023. 2 18
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.