Pith. sign in

REVIEW 3 major objections 5 minor 71 references

Magnifier: A Multi-grained Neural Network-based Architecture for Burned Area Delineation

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read The Magnifier architecture claims that processing each satellite image twice — once at full resolution and once as a grid of local patches, then fusing the two feature maps — improves burned area segmentation by 2.65 average IoU points…

desk verdict Useful multi-granularity wrapper for burned area segmentation with real gains, but the central attribution claim is under-tested and the abstract overstates the compute savings. read the letter →

arxiv 2504.19589 v1 pith:PN6UK3EK submitted 2025-04-28 cs.CV eess.IV

classification cs.CVeess.IV
keywords burnedareadelineationsemanticsegmentationdual-encodermulti-granularitysatelliteimagerySentinel-2Landsat-8data-efficientdeeplearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that a segmentation model for burned areas can be made more accurate without collecting a single new label, by giving the same image to the network twice at different granularities. Its proposed wrapper, Magnifier, attaches to any encoder-decoder model: one encoder sees the whole image, a second sees non-overlapping patches that are encoded and recomposed into their original positions, and the two feature maps are concatenated before the decoder. Across three public satellite datasets (California, Europe, Indonesia), the authors report an average IoU gain of 2.65 percentage points over the corresponding small single-model baselines, and the magnified small models often beat their own larger versions. The claim matters because labeled wildfire imagery is scarce and expensive, and larger models tend to overfit on these datasets rather than improve.

What carries the argument

The load-bearing object is the dual-encoder fusion pipeline. A global encoder embeds the full image; a patch encoder embeds each non-overlapping $64 \times 64$ crop, and a recomposition step places each patch's embedding back into its original grid position using a stored (row, column) tuple, producing a second full-size feature map. Channel-wise concatenation of the two maps yields a tensor that the shared decoder turns into the binary burn mask. This mechanism forces the same labeled pixels to be represented at two contextual scales, which is how the paper claims to extract more information from the same data.

What would settle it

Re-run the Magnifier recipe on the same three datasets with patch sizes of $32 \times 32$, $64 \times 64$, $128 \times 128$, and $256 \times 256$ while keeping all other settings fixed; if the average IoU advantage over the single-model baseline does not persist across reasonable patch sizes, the claim that multi-grained fusion is inherently beneficial would be overturned in favor of a specific-scale effect.

Watch

Extended reading notes

Core claim

The central discovery is that multi-grained context, not raw model size, is what helps in this low-data setting. Magnifier instantiates the same small encoder twice with separate weights — one branch consuming the full image and the other consuming $64 \times 64$ patches that are encoded individually and recomposed into a full-size embedding by their positional tuples — then concatenates the two embeddings along the channel axis and feeds them to a single shared decoder. The authors report that this configuration improves F1 and IoU over single-model baselines in most tested combinations, achieves the best mean rank in four of five architecture settings, and outperforms the corresponding large single models in 14 of 15 F1 comparisons and 11 of 15 IoU comparisons, while using fewer parameters than the large versions. They attribute the improvement to the local/global distinction rather than to parameter growth.

Load-bearing premise

The whole comparison rests on a fixed patch size of $64 \times 64$ pixels applied uniformly to datasets whose ground resolutions differ (10 m, 20 m, 30 m), so the 'local' scale covers different physical areas in each dataset and no ablation over patch size is reported.

Editorial extensions

If this is right

  • Magnifier raises F1 and IoU over the corresponding small single models in the large majority of the 15 tested combinations, with the largest gains on Indonesia (e.g., +11.5 IoU points for DeepLabV3+-MobileNetV3-Small).
  • The best magnified model — DeepLabV3+ with a ResNet-18 backbone — reaches the highest F1 and IoU on all three datasets while using 22M backbone parameters versus 42M for the ResNet-101 single model.
  • Because Magnifier only requires an encoder-decoder structure, the same wrapper can be dropped onto CNN and transformer architectures alike, making the data-efficiency gain portable.
  • In the authors' comparison, magnified small models beat BurntNet on California and Indonesia with a fraction of the GFLOPs, though BurntNet wins on Europe.
  • The authors interpret the comparison against large single models as evidence that the gain comes from multi-grained inputs rather than from adding parameters.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the near-zero gains on Europe with ResNet-18 and the large gains on Indonesia suggest that the benefit may depend on base model strength, dataset size, and how much local texture matters; the +2.65 average is a mean over heterogeneous configurations, not a uniform effect.
  • Editorial inference: because the patch size is fixed while ground resolution varies across the datasets, testing dataset-specific patch sizes is the most direct untested variable; a sensitivity sweep would separate the multi-grained principle from the accident of one scale.
  • Editorial inference: the positional recomposition mechanism could be extended with learned fusion or cross-attention between the global and local branches, which the present architecture does not use.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes Magnifier, a dual-encoder segmentation architecture that processes the full input image with one encoder and non-overlapping 64x64 patches with a second encoder, recomposes the patch embeddings into a full-resolution feature map, concatenates the global and recomposed-local embeddings along the channel axis, and feeds the result to the decoder of an existing encoder-decoder model. The method is evaluated for burned area delineation on three public datasets (CaBuAr, Europe, Indonesia) with DeepLabV3+, U-Net, and SegFormer base architectures, using small and large backbones, cross-validation, and the Asymmetric Unified Focal loss. The paper reports an average +2.65% IoU gain over single-model small baselines, better mean rank in 4 of 5 configurations, and competitive or better performance than larger single models at lower GFLOPs, and it attributes this improvement to the multi-grained two-path design rather than to increased parameters.

Significance. If the central claim is supported, Magnifier is a useful and simple plug-in that improves segmentation under data scarcity while keeping models small, and the public code and cross-validated results on three real-world datasets are clear strengths. The paper is also honest in reporting the SegFormer and U-Net cases where gains are negligible or negative. However, the causal attribution to multi-granularity is not currently isolated from the doubled encoder capacity, and the headline average gain masks configuration-dependent results. The evidence supports an average improvement for several CNN configurations, but not the stated mechanism or a general claim of superiority.

major comments (3)
  1. [Section IV-F4] The claim that the improvement is 'related to its two paths and images analyzed at different levels and not to the number of parameters' is not established by the reported experiments. Magnifier doubles the encoder parameters and roughly doubles the FLOPs of the small baseline (Table IVa: DeepLabV3+ ResNet18 goes from 40.2 GFLOPs/11M to 76.9 GFLOPs/22M). The only capacity control is comparison with larger single models (ResNet101, MobileNetV3-L, MiT-B1), which differ in depth, width, and overfitting behavior; those comparisons do not control for capacity or for the effect of having two encoder branches. A single-path model with a matched parameter/computation budget, or a dual-encoder control in which both branches receive the same full image, is required to support the attribution of the gain to multi-granularity.
  2. [Section IV-F3, Table IV] The average +2.65% IoU gain is not statistically supported and is not uniform across configurations. For example, U-Net Magnifier with ResNet18 on Europe is 69.2 IoU versus 70.4 for the single model (-1.2), and SegFormer Magnifier on Indonesia is 69.9 versus 70.0 (-0.1). The paper reports only fold standard deviations, not significance tests or multiple training seeds, so the reader cannot determine whether the positive average is robust or driven by a few favorable configurations. Please add paired significance tests across folds and/or repeated-seed experiments, and report per-configuration effect sizes rather than only the average gain.
  3. [Section III-B, Section IV-E] The fixed 64x64 patch size is applied to datasets with ground sample distances of 10m, 20m, and 30m, so the local patches cover physically different areas (roughly 640m, 1280m, and 1920m per side). The 'local versus global' distinction is therefore not constant across datasets, and no patch-size sensitivity analysis is reported. Without an ablation over patch sizes on at least one dataset, the average gain could be tied to this single hand-picked value rather than to the multi-grained principle itself. Please add such an analysis or justify the choice of 64x64 per dataset.
minor comments (5)
  1. [Equation (2)] In the definition of the modified asymmetric focal loss, the first term appears to lack a summation over the rare-class pixels and the notation yi:r is undefined; please clarify the indexing and the intended sum.
  2. [Section IV-D] GFLOPs are described as 'the number of mathematical operations a system is capable of performing per second,' but FLOPs is a count of operations, not a rate; please correct the wording to avoid a units error.
  3. [Figure 1 caption] The caption says 'Average Mean IoU,' which is redundant; it should say 'average IoU' or 'mean IoU.'
  4. [Section IV-E] The phrase 'polynomial learning rate scheduler with a power of 1 for 55 iterations' is ambiguous; please specify whether this means 55 epochs, 55 training iterations, or a warmup length, and state the decay behavior.
  5. [Section IV-F5] In the discussion of the Europe results, the text refers to 'the higher difficulty due to transfer learning since the seven folds adopted in the cross-validation process are region-based'; this is not transfer learning but cross-validation, so please rephrase to avoid confusion.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central claim is empirically evaluated on held-out folds, and the cited self-authored datasets are benchmarks rather than load-bearing assumptions.

full rationale

The paper makes no formal derivational claim; its contribution is a dual-encoder architecture and an empirical comparison. The headline result (+2.65% average IoU) is computed from held-out cross-validation folds with loss hyperparameters taken from an external AUF loss paper and not fitted to the test metrics, so there is no fitted-input-called-prediction pattern. Self-citations appear only for the CaBuAr and Europe datasets used as evaluation benchmarks; the reported gains are not defined in terms of those datasets' construction, and the datasets are independently published resources. The claim that the improvement is due to multi-granular analysis rather than parameter count is under-supported because no parameter-matched single-path control is run, but that is an experimental confound or correctness risk, not circularity: nothing in the paper's equations or definitions makes the conclusion equivalent to its inputs. Hence no circular step can be exhibited, and the honest finding is a score of 0.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

No new physical entities are introduced. The only hand-chosen elements are the patch size and the loss coefficients, both fixed without ablation.

free parameters (2)
  • patch_size = 64x64 pixels
    Single crop size chosen for all datasets and architectures; controls the local/global split and is never swept. Magnifier's benefit could depend on matching patch size to sensor resolution.
  • AUF loss coefficients = lambda=0.5, delta=0.6, gamma=0.1
    Adopted unchanged from Yeung et al. [68]; not tuned here, but they shape the training objective for the imbalanced burned class.
assumptions (3)
  • domain assumption Ground truth burned-area masks are accurate for all three datasets.
    Section III-B treats CalFire, Copernicus EMS, and Indonesian expert annotations as supervision; annotation noise is not quantified.
  • domain assumption Resampling Sentinel-2 and Landsat-8 bands to common resolutions preserves the discriminative signal.
    Section III-B1/B2 up- or down-sample bands to 20m and 10m; distorted spectral or spatial mixing could bias comparisons.
  • ad hoc to paper Channel concatenation is a sufficient fusion of global and recomposed local features.
    Section III-C and Figure 4 choose concatenation without comparing to learned fusion, attention, or other fusion strategies.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Magnifier: A Multi-grained Neural Network-based Architecture for Burned Area Delineation." pith.science (2026). https://pith.science/paper/PN6UK3EK

@misc{pith2026250419589,
  author       = {Pith},
  title        = {Pith review of: Magnifier: A Multi-grained Neural Network-based Architecture for Burned Area Delineation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PN6UK3EK}},
  note         = {Machine review of arXiv:2504.19589}
}
read the original abstract

In crisis management and remote sensing, image segmentation plays a crucial role, enabling tasks like disaster response and emergency planning by analyzing visual data. Neural networks are able to analyze satellite acquisitions and determine which areas were affected by a catastrophic event. The problem in their development in this context is the data scarcity and the lack of extensive benchmark datasets, limiting the capabilities of training large neural network models. In this paper, we propose a novel methodology, namely Magnifier, to improve segmentation performance with limited data availability. The Magnifier methodology is applicable to any existing encoder-decoder architecture, as it extends a model by merging information at different contextual levels through a dual-encoder approach: a local and global encoder. Magnifier analyzes the input data twice using the dual-encoder approach. In particular, the local and global encoders extract information from the same input at different granularities. This allows Magnifier to extract more information than the other approaches given the same set of input images. Magnifier improves the quality of the results of +2.65% on average IoU while leading to a restrained increase in terms of the number of trainable parameters compared to the original model. We evaluated our proposed approach with state-of-the-art burned area segmentation models, demonstrating, on average, comparable or better performances in less than half of the GFLOPs.

Figures

Figures reproduced from arXiv: 2504.19589 by the authors.

Figure 1
Figure 1. Average Mean IoU vs. Number of parameters. The Average Mean IoU has been computed considering all datasets. Architecture Type refers to the family of base networks employed for segmentation. On average, the Magnifier backbone achieves IoU improvements compared to MobileNetV3 Small and Large, ResNet-18 and 101, and MiT-B0 and B1 without increasing the number of parameters too much. techniques, makes the Earth Observa… view at source ↗
Figure 2
Figure 2. RGB samples taken from the three datasets with the corresponding binary ground truth. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 5
Figure 5. For this procedure, the original input image [PITH_FULL_IMAGE:figures/full_fig_p005_5.png] view at source ↗
Figures from the paper (5 more)
Figure 3
Figure 3. Figure 3: Distribution of wildfires in the analyzed datasets. In (a) the areas covered by wildfires in [PITH_FULL_IMAGE:figures/full_fig_p006_3.png]
Figure 4
Figure 4. Figure 4: Magnifier architecture. In the lower branch, (i) the image is cropped in smaller patches (as shown in Figure 5), giving each patch to an encoder. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Cropping procedure. The image is cropped in patches, and each of them keeps the original position associated. embedded into a feature vector of size w0 × h0 × C0 with the Patch Encoder (bottom branch in [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 7
Figure 7. Figure 7: Base architectures of DeepLabV3+, U-Net and SegFormer. [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Example RGB images and corresponding ground truth with predictions. Images are grouped by architecture and type of backbone. [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

71 extracted references · 49 canonical work pages

  1. [1]

    Climate change impact on future wildfire danger and activity in southern europe: a review,

    J.-l. Dupuy, H. Fargeon, N. Martin-StPaul, F. Pimont, J. Ruffault, M. Guijarro, C. Hernando, J. Madrigal, and P. Fernandes, “Climate change impact on future wildfire danger and activity in southern europe: a review,” Annals of Forest Science , vol. 77, no. 2, pp. 1–24, 2020

  2. [2]

    Increased likelihood of heat-induced large wildfires in the Mediterranean Basin,

    J. Ruffault, T. Curt, V . Moron, R. M. Trigo, F. Mouillot, N. Koutsias, F. Pimont, N. Martin-StPaul, R. Barbero, J.-L. Dupuy et al., “Increased likelihood of heat-induced large wildfires in the Mediterranean Basin,” Scientific reports, vol. 10, no. 1, pp. 1–9, 2020

  3. [3]

    Changing wildfire, changing forests: the effects of climate change on fire regimes and vegetation in the Pacific Northwest, USA,

    J. E. Halofsky, D. L. Peterson, and B. J. Harvey, “Changing wildfire, changing forests: the effects of climate change on fire regimes and vegetation in the Pacific Northwest, USA,” Fire Ecology, vol. 16, no. 1, pp. 1–26, 2020. [Online]. Available: https://doi.org/10.1186/s42408-019-0062-8

  4. [4]

    Image texture analysis enhances classification of fire extent and severity using sentinel 1 and 2 satellite imagery,

    R. K. Gibson, A. Mitchell, and H.-C. Chang, “Image texture analysis enhances classification of fire extent and severity using sentinel 1 and 2 satellite imagery,” Remote Sensing, vol. 15, no. 14, p. 3512, 2023

  5. [5]

    Double-Step U-Net: A Deep Learning-Based Approach for the Estimation of Wildfire Damage Severity through Sentinel-2 Satellite Data,

    A. Farasin, L. Colomba, and P. Garza, “Double-Step U-Net: A Deep Learning-Based Approach for the Estimation of Wildfire Damage Severity through Sentinel-2 Satellite Data,” Applied Sciences , vol. 10, no. 12, 2020. [Online]. Available: https://www.mdpi.com/2076-3417/ 10/12/4332

  6. [6]

    Wild- fire detection from multisensor satellite imagery using deep semantic segmentation,

    D. Rashkovetsky, F. Mauracher, M. Langer, and M. Schmitt, “Wild- fire detection from multisensor satellite imagery using deep semantic segmentation,” IEEE Journal of Selected Topics in Applied Earth Ob- servations and Remote Sensing , vol. 14, pp. 7001–7016, 2021

  7. [7]

    Wildfire segmentation using deep vision transformers,

    R. Ghali, M. A. Akhloufi, M. Jmal, W. Souidene Mseddi, and R. Attia, “Wildfire segmentation using deep vision transformers,”Remote Sensing, vol. 13, no. 17, p. 3527, 2021

  8. [8]

    A Deep Learning Approach for Burned Area Segmentation with Sentinel-2 Data,

    L. Knopp, M. Wieland, M. R ¨attich, and S. Martinis, “A Deep Learning Approach for Burned Area Segmentation with Sentinel-2 Data,” Remote Sensing , vol. 12, no. 15, 2020. [Online]. Available: https://www.mdpi.com/2072-4292/12/15/2422

Show all 71 references
  1. [9]

    Supervised Burned Areas delineation by means of Sentinel-2 imagery and Convolu- tional Neural Networks,

    A. Farasin, L. Colomba, G. Palomba, G. Nini, and C. Rossi, “Supervised Burned Areas delineation by means of Sentinel-2 imagery and Convolu- tional Neural Networks,” in Proceedings of the 17th International Con- ference on Information Systems for Crisis Response and Management ...

  2. [10]

    Sentinel-1 flood delineation with supervised machine learning,

    G. Palomba, A. Farasin, and C. Rossi, “Sentinel-1 flood delineation with supervised machine learning,” in ISCRAM 2020 Conference Proceedings–17th International Conference on Information Systems for Crisis Response and Management , 2020, pp. 1072–1083

  3. [11]

    A Comparative Analysis for Air Quality Estimation from Traffic and Meteorological Data,

    E. Arnaudo, A. Farasin, and C. Rossi, “A Comparative Analysis for Air Quality Estimation from Traffic and Meteorological Data,” Applied Sciences , vol. 10, no. 13, 2020. [Online]. Available: https://www.mdpi.com/2076-3417/10/13/4587

  4. [12]

    Fast and accurate land-cover classification on medium-resolution remote-sensing images using segmentation models,

    W. Zhang, P. Tang, and L. Zhao, “Fast and accurate land-cover classification on medium-resolution remote-sensing images using segmentation models,” International Journal of Remote Sensing , vol. 42, no. 9, pp. 3277–3301, 2021. [Online]. Available: https: //doi.org/10.1080/0143...

  5. [13]

    Urban Land Use and Land Cover Classification Using Novel Deep Learning Models Based on High Spatial Resolution Satellite Imagery,

    P. Zhang, Y . Ke, Z. Zhang, M. Wang, P. Li, and S. Zhang, “Urban Land Use and Land Cover Classification Using Novel Deep Learning Models Based on High Spatial Resolution Satellite Imagery,” Sensors, vol. 18, no. 11, 2018. [Online]. Available: https://www.mdpi.com/1424-8220/18/11/3717

  6. [14]

    Cnn, rnn, or vit? an evaluation of different deep learning architectures for spatio-temporal representation of sentinel time series,

    L. Zhao and S. Ji, “Cnn, rnn, or vit? an evaluation of different deep learning architectures for spatio-temporal representation of sentinel time series,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , vol. 16, pp. 44–56, 2022

  7. [15]

    Scene classification in remote sensing images using dynamic kernels,

    R. Datla, V . Chalavadi, and K. M. C, “Scene classification in remote sensing images using dynamic kernels,” in 2021 International Joint Conference on Neural Networks (IJCNN) . IEEE, Jul. 2021, p. 1–8. [Online]. Available: http://dx.doi.org/10.1109/IJCNN52387.2021. 9533648

  8. [16]

    Learning scene-vectors for remote sensing image scene classification,

    R. Datla, N. Perveen, and K. M. C., “Learning scene-vectors for remote sensing image scene classification,” Neurocomputing, vol. 587, p. 127679, Jun. 2024. [Online]. Available: http://dx.doi.org/10.1016/j. neucom.2024.127679

  9. [17]

    U-Net: Convolutional Networks for Biomedical Image Segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional Networks for Biomedical Image Segmentation,” in Lecture Notes in Computer Science . Springer International Publishing, 2015, pp. 234–

  10. [18]

    Encoder- Decoder with Atrous Separable Convolution for Semantic Image Seg- mentation,

    L.-C. Chen, Y . Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder- Decoder with Atrous Separable Convolution for Semantic Image Seg- mentation,” in Computer Vision – ECCV 2018 , ser. Lecture Notes in Computer Science, V . Ferrari, M. Hebert, C. Sminchisescu, and Y . Weiss,...

  11. [19]

    SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers,

    E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers,” in Advances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y . Dauphin, P. Liang, and J. W. Vaughan, E...

  12. [20]

    Searching for MobileNetV3,

    A. Howard, M. Sandler, B. Chen, W. Wang, L.-C. Chen, M. Tan, G. Chu, V . Vasudevan, Y . Zhu, R. Pang, H. Adam, and Q. Le, “Searching for MobileNetV3,” in 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2019, pp. 1314–1324, iSSN: 2380-7504

  13. [21]

    Deep learning,

    Y . LeCun, Y . Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, p. 436–444, May 2015. [Online]. Available: http://dx.doi.org/10.1038/nature14539

  14. [22]

    Gradient-based learning applied to document recognition,

    Y . Lecun, L. Bottou, Y . Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998

  15. [23]

    Long short-term memory,

    S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comput., vol. 9, no. 8, p. 1735–1780, Nov. 1997. [Online]. Available: https://doi.org/10.1162/neco.1997.9.8.1735

  16. [24]

    Bidirectional recurrent neural net- works,

    M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural net- works,” IEEE transactions on Signal Processing , vol. 45, no. 11, pp. 2673–2681, 1997

  17. [25]

    Attention is All You Need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is All You Need,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , ser. NIPS’17. Red Hook, NY , USA: Curran Associates...

  18. [26]

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” 2020. [Online]. Available: https://arxi...

  19. [27]

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologi...

  20. [28]

    Human-level control through deep reinforcement learning,

    V . Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep...

  21. [29]

    A threshold selection method from gray-level histograms,

    N. Otsu, “A threshold selection method from gray-level histograms,” IEEE transactions on systems, man, and cybernetics , vol. 9, no. 1, pp. 62–66, 1979

  22. [30]

    Normalized cuts and image segmentation,

    J. Shi and J. Malik, “Normalized cuts and image segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 22, no. 8, pp. 888–905, 2000. 15

  23. [31]

    Mean shift: a robust approach toward feature space analysis,

    D. Comaniciu and P. Meer, “Mean shift: a robust approach toward feature space analysis,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 24, no. 5, pp. 603–619, 2002

  24. [32]

    Time-of-flight cameras in space: Pose estimation with deep learning methodologies,

    A. Koudounas, F. Giobergia, and E. Baralis, “Time-of-flight cameras in space: Pose estimation with deep learning methodologies,” in 2022 IEEE 16th International Conference on Application of Information and Communication Technologies (AICT), 2022, pp. 1–6

  25. [33]

    DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs,

    L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs,” IEEE Trans- actions on Pattern Analysis and Machine Intelligence , vol. 40, no. 4, pp. 834–848, 2018

  26. [34]

    Rethinking Atrous Convolution for Semantic Image Segmentation,

    L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam, “Rethinking Atrous Convolution for Semantic Image Segmentation,” 2017. [Online]. Available: https://arxiv.org/abs/1706.05587

  27. [35]

    A dataset for burned area delineation and severity estimation from satellite imagery,

    L. Colomba, A. Farasin, S. Monaco, S. Greco, P. Garza, D. Apiletti, E. Baralis, and T. Cerquitelli, “A dataset for burned area delineation and severity estimation from satellite imagery,” in Proceedings of the 31st ACM International Conference on Information & Knowledge Manage...

  28. [36]

    Swin Transformer: Hierarchical Vision Transformer using Shifted Win- dows,

    Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin Transformer: Hierarchical Vision Transformer using Shifted Win- dows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2021, pp. 10 012–10 022

  29. [37]

    Deep Residual Learning for Image Recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016, pp. 770–778, iSSN: 1063-6919

  30. [38]

    Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set,

    S. Ji, S. Wei, and M. Lu, “Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set,” IEEE Transactions on Geoscience and Remote Sensing , vol. 57, no. 1, pp. 574–586, 2019

  31. [39]

    reben: Refined bigearthnet dataset for remote sensing image analysis,

    K. N. Clasen, L. Hackel, T. Burgert, G. Sumbul, B. Demir, and V . Markl, “reben: Refined bigearthnet dataset for remote sensing image analysis,” arXiv preprint arXiv:2407.03653 , 2024

  32. [40]

    Scene classifica- tion of high-resolution remotely sensed image based on resnet,

    M. Wang, X. Zhang, X. Niu, F. Wang, and X. Zhang, “Scene classifica- tion of high-resolution remotely sensed image based on resnet,” Journal of Geovisualization and Spatial Analysis , vol. 3, no. 2, p. 16, 2019

  33. [41]

    Winter wheat mapping method based on pseudo-labels and u-net model for training sample shortage,

    J. Zhang, S. You, A. Liu, L. Xie, C. Huang, X. Han, P. Li, Y . Wu, and J. Deng, “Winter wheat mapping method based on pseudo-labels and u-net model for training sample shortage,” Remote Sensing , vol. 16, no. 14, 2024. [Online]. Available: https: //www.mdpi.com/2072-4292/16/14/2553

  34. [42]

    Improving crop classification accuracy with integrated sentinel-1 and sentinel-2 data: a case study of barley and wheat,

    G. R. Faqe Ibrahim, A. Rasul, and H. Abdullah, “Improving crop classification accuracy with integrated sentinel-1 and sentinel-2 data: a case study of barley and wheat,” Journal of Geovisualization and Spatial Analysis, vol. 7, no. 2, p. 22, 2023

  35. [43]

    Kan you see it? kans and sentinel for effective and explainable crop field segmentation,

    D. R. Cambrin, E. Poeta, E. Pastor, T. Cerquitelli, E. Baralis, and P. Garza, “Kan you see it? kans and sentinel for effective and explainable crop field segmentation,” 2024. [Online]. Available: https://arxiv.org/abs/2408.07040

  36. [44]

    Deepglobe 2018: A challenge to parse the earth through satellite images,

    I. Demir, K. Koperski, D. Lindenbaum, G. Pang, J. Huang, S. Basu, F. Hughes, D. Tuia, and R. Raskar, “Deepglobe 2018: A challenge to parse the earth through satellite images,” in Proceedings of the IEEE conference on computer vision and pattern recognition workshops, 2018, pp. 172–181

  37. [45]

    Ms-vacsnet: A network for multi-scale volcanic ash cloud segmentation in remote sensing images,

    G. Swetha, R. Datla, C. Vishnu, and K. M. C, “Ms-vacsnet: A network for multi-scale volcanic ash cloud segmentation in remote sensing images,” in 2023 18th International Conference on Machine Vision and Applications (MVA) . IEEE, Jul. 2023, p. 1–6. [Online]. Available: http://...

  38. [46]

    A multimodal semantic segmentation for airport runway delineation in panchromatic remote sensing images,

    R. Datla, V . Chalavadi, and K. M. Chalavadi, “A multimodal semantic segmentation for airport runway delineation in panchromatic remote sensing images,” in Fourteenth International Conference on Machine Vision (ICMV 2021) , W. Osten, D. Nikolaev, and J. Zhou, Eds. SPIE, Mar. 2...

  39. [47]

    Mutual attention inception net- work for remote sensing visual question answering,

    X. Zheng, B. Wang, X. Du, and X. Lu, “Mutual attention inception net- work for remote sensing visual question answering,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–14, 2022

  40. [48]

    Mmflood: A multimodal dataset for flood delineation from satellite imagery,

    F. Montello, E. Arnaudo, and C. Rossi, “Mmflood: A multimodal dataset for flood delineation from satellite imagery,” IEEE Access, vol. 10, pp. 96 774–96 787, 2022

  41. [49]

    Spectral signature analysis of false positive burned area detection from agricultural harvests using Sentinel-2 data,

    D. van Dijk, S. Shoaie, T. van Leeuwen, and S. Veraverbeke, “Spectral signature analysis of false positive burned area detection from agricultural harvests using Sentinel-2 data,” International Journal of Applied Earth Observation and Geoinformation , vol. 97, p. 102296, 2021....

  42. [50]

    A method for extracting burned areas from landsat tm/etm+ images by soft aggregation of multiple spectral indices and a region growing algorithm,

    D. Stroppiana, G. Bordogna, P. Carrara, M. Boschetti, L. Boschetti, and P. Brivio, “A method for extracting burned areas from landsat tm/etm+ images by soft aggregation of multiple spectral indices and a region growing algorithm,” ISPRS Journal of Photogrammetry and Remote Sen...

  43. [51]

    Burnt area index (baim) for burned area discrimination at regional scale using modis data,

    M. P. Mart ´ın, I. G ´omez, and E. Chuvieco, “Burnt area index (baim) for burned area discrimination at regional scale using modis data,” Forest Ecology and Management , no. 234, p. S221, 2006

  44. [52]

    Bais2: Burned area index for sentinel-2,

    F. Filipponi, “Bais2: Burned area index for sentinel-2,” Proceedings, vol. 2, no. 7, 2018. [Online]. Available: https://www.mdpi.com/ 2504-3900/2/7/364

  45. [53]

    Remote sensing of fire severity: as- sessing the performance of the normalized burn ratio,

    D. Roy, L. Boschetti, and S. Trigg, “Remote sensing of fire severity: as- sessing the performance of the normalized burn ratio,” IEEE Geoscience and Remote Sensing Letters , vol. 3, no. 1, pp. 112–116, 2006

  46. [54]

    Development of a Sentinel-2 burned area algorithm: Generation of a small fire database for sub-Saharan Africa,

    E. Roteta, A. Bastarrika, M. Padilla, T. Storm, and E. Chuvieco, “Development of a Sentinel-2 burned area algorithm: Generation of a small fire database for sub-Saharan Africa,” Remote Sensing of Environment , vol. 222, pp. 1–17, 2019. [Online]. Available: https://www.scienced...

  47. [55]

    Detecting Burn Severity across Mediterranean Forest Types by Coupling Medium-Spatial Resolution Satellite Imagery and Field Data,

    L. Saulino, A. Rita, A. Migliozzi, C. Maffei, E. Allevato, A. P. Garonna, and A. Saracino, “Detecting Burn Severity across Mediterranean Forest Types by Coupling Medium-Spatial Resolution Satellite Imagery and Field Data,” Remote Sensing, vol. 12, no. 4, 2020. [Online]. Availa...

  48. [56]

    Semantic Segmentation of Earth Observation Data Using Multimodal and Multi-scale Deep Networks,

    N. Audebert, B. Le Saux, and S. Lef `evre, “Semantic Segmentation of Earth Observation Data Using Multimodal and Multi-scale Deep Networks,” in Computer Vision – ACCV 2016 , S.-H. Lai, V . Lepetit, K. Nishino, and Y . Sato, Eds. Cham: Springer International Publishing, 2017, p...

  49. [57]

    Global canopy height regression and uncertainty estimation from GEDI LIDAR waveforms with deep ensembles,

    N. Lang, N. Kalischek, J. Armston, K. Schindler, R. Dubayah, and J. D. Wegner, “Global canopy height regression and uncertainty estimation from GEDI LIDAR waveforms with deep ensembles,” Remote Sensing of Environment , vol. 268, p. 112760, 2022. [Online]. Available: https://ww...

  50. [58]

    Semantic segmentation of burned areas in satellite images using a U-Net-based convolutional neural network,

    A. Brand and A. Manandhar, “Semantic segmentation of burned areas in satellite images using a U-Net-based convolutional neural network,” The International Archives of Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. 43, pp. 47–53, 2021

  51. [59]

    Uni-Temporal Multispectral Imagery for Burned Area Mapping with Deep Learning,

    X. Hu, Y . Ban, and A. Nascetti, “Uni-Temporal Multispectral Imagery for Burned Area Mapping with Deep Learning,” Remote Sensing , vol. 13, no. 8, 2021. [Online]. Available: https://www.mdpi.com/ 2072-4292/13/8/1509

  52. [60]

    Large-scale burn severity mapping in multispectral imagery using deep semantic segmentation models,

    X. Hu, P. Zhang, and Y . Ban, “Large-scale burn severity mapping in multispectral imagery using deep semantic segmentation models,” ISPRS Journal of Photogrammetry and Remote Sensing , vol. 196, pp. 228–240, 2023. [Online]. Available: https://www.sciencedirect.com/ science/art...

  53. [61]

    Attention to Fires: Multi- Channel Deep Learning Models for Wildfire Severity Prediction,

    S. Monaco, S. Greco, A. Farasin, L. Colomba, D. Apiletti, P. Garza, T. Cerquitelli, and E. Baralis, “Attention to Fires: Multi- Channel Deep Learning Models for Wildfire Severity Prediction,” Applied Sciences , vol. 11, no. 22, 2021. [Online]. Available: https://www.mdpi.com/2...

  54. [62]

    Burnt-net: Wildfire burned area mapping with single post-fire sentinel-2 data and deep learning morphological neural network,

    S. T. Seydi, M. Hasanlou, and J. Chanussot, “Burnt-net: Wildfire burned area mapping with single post-fire sentinel-2 data and deep learning morphological neural network,” Ecological Indicators , vol. 140, p. 108999, 2022. [Online]. Available: https://www.sciencedirect. com/sc...

  55. [63]

    CaBuAr: California Burned Areas dataset for delineation,

    D. Rege Cambrin, L. Colomba, and P. Garza, “CaBuAr: California Burned Areas dataset for delineation,” IEEE Geoscience and Remote Sensing Magazine, 2023

  56. [64]

    Deep learning dataset for estimating burned areas: Case study, indonesia,

    Y . Prabowo, A. D. Sakti, K. A. Pradono, Q. Amriyah, F. H. Rasyidy, I. Bengkulah, K. Ulfa, D. S. Candra, M. T. Imdad, and S. Ali, “Deep learning dataset for estimating burned areas: Case study, indonesia,” Data, vol. 7, no. 6, p. 78, 2022

  57. [65]

    Sentinel-2: ESA's optical high-resolution mission for GMES operational services,

    M. Drusch, U. D. Bello, S. Carlier, O. Colin, V . Fernandez, F. Gascon, B. Hoersch, C. Isola, P. Laberinti, P. Martimort, A. Meygret, F. Spoto, O. Sy, F. Marchese, and P. Bargellini, “Sentinel-2: ESA's optical high-resolution mission for GMES operational services,” Remote Sens...

  58. [66]

    Sentinel-2 l2a user guide,

    “Sentinel-2 l2a user guide,” 2023. [Online]. Available: https://sentinel. esa.int/web/sentinel/user-guides/sentinel-2-msi/product-types/level-2a

  59. [67]

    Landsat-8: Science and product vision for terrestrial global change research,

    D. P. Roy, M. A. Wulder, T. R. Loveland, C. E. Woodcock, R. G. Allen, M. C. Anderson, D. Helder, J. R. Irons, D. M. Johnson, R. Kennedy 16 et al. , “Landsat-8: Science and product vision for terrestrial global change research,” Remote sensing of Environment , vol. 145, pp. 154...

  60. [68]

    Unified focal loss: Generalising dice and cross entropy-based losses to handle class imbalanced medical image segmentation,

    M. Yeung, E. Sala, C.-B. Sch ¨onlieb, and L. Rundo, “Unified focal loss: Generalising dice and cross entropy-based losses to handle class imbalanced medical image segmentation,” Computerized Medical Imaging and Graphics , vol. 95, p. 102026, 2022. [Online]. Available: https://...

  61. [69]

    Statistical comparisons of classifiers over multiple data sets,

    J. Dem ˇsar, “Statistical comparisons of classifiers over multiple data sets,” J. Mach. Learn. Res. , vol. 7, p. 1–30, Dec. 2006

  62. [70]

    Refaeilzadeh, L

    P. Refaeilzadeh, L. Tang, and H. Liu, Cross-Validation. Springer US, 2009, p. 532–538. [Online]. Available: http://dx.doi.org/10.1007/ 978-0-387-39940-9 565 Daniele Rege Cambrin received a master’s de- gree in Computer Engineering from Politecnico di Torino in 2022. He is a Ph...

  63. [241]

    Available: https://doi.org/10.1007/978-3-319-24574-4 28

    [Online]. Available: https://doi.org/10.1007/978-3-319-24574-4 28

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.