{"id":"d1d58218-cb4c-4cc0-b759-d015b5273ffe","arxiv_id":"2501.10891","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"OpenEarthMap-SAR is a public benchmark of 5,033 sub-meter SAR images with aligned optical imagery and 8-class land cover labels across 35 regions in Japan, France, and the USA.","lead":"This paper releases OpenEarthMap-SAR, a public benchmark of 5,033 very-high-resolution SAR images with 8-class land cover labels across 35 regions in three countries. The dataset is meant to support all-weather land cover mapping and served as the official IEEE GRSS Data Fusion Contest Track I dataset.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Label validation is optical-centric: Table III's pseudo-vs-manual agreement does not establish that labels are valid for SAR geometry, so the benchmark's core supervision is unverified.","rationale":"The reader's weakest assumption—that optical-trained pseudo labels are accurate enough for SAR supervision—is correct and is the same area I identify, but I go further: the manual labels themselves may not be SAR-valid, and the paper provides no evidence that any label is grounded in SAR geometry. The pseudo-vs-manual agreement in Table III cannot validate SAR utility if both label types originate from optical interpretation. This is a load-bearing concern because the dataset's primary contribution is enabling SAR land cover mapping; if labels are aligned to optical but not SAR, the benchmark's training signal and evaluation protocol are compromised. I credit the authors for releasing the data on Zenodo, for transparently reporting Table III's low agreement, and for discussing temporal and sensor-mismatch limitations in the conclusion. However, the missing piece is an independent check of label correctness against SAR content. The annotation-count inconsistencies (1.5M vs 1.4708M segments; 700 vs 490 manual labels) add uncertainty but are secondary. The concrete test proposed would settle whether labels transfer to SAR geometry. Given the dataset exists and is likely useful even with noisy pseudo labels, I do not recommend changing the reader's conditional verdict; the condition should include demonstrating SAR-geometry label validity.","tokens_in":10735,"tokens_out":5737,"duration_ms":67804,"concrete_test":"Sample 100 images from the Zenodo release across regions and classes. Have two remote-sensing experts independently label the SAR amplitude image directly (with no optical reference) and compute IoU against the released pseudo labels and manual labels. Also extract acquisition dates from metadata and test whether label agreement declines with increasing SAR-optical temporal gap. If SAR-only annotation IoU is substantially below Table III values, the labels are optical-centric and the benchmark requires re-labeling or a geometry-aware alignment audit.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The dataset's central claim is that its 8-class labels support SAR land cover segmentation, yet every validation step in the paper compares labels against other labels derived from optical imagery, not against SAR-grounded truth. Section II-B states pseudo labels are generated by pre-trained OpenEarthMap models trained on optical data, and the manual annotations are described only as 'manually annotated' without stating whether annotators drew them on the optical or SAR image. The paper acknowledges manual alignment of optical and SAR pairs but reports no residual misalignment metric. At 0.15-0.5 m GSD, Umbra Spotlight SAR exhibits layover, shadow, and speckle that cause local geometric distortion even after geocoding; if labels were drawn in optical geometry, pixel-level label-SAR correspondence is suspect. Table III's 68.2% mean agreement and 56.02% mean IoU between pseudo and manual labels therefore measure consistency within the optical domain, not correctness against SAR. The low SAR-only mIoU (~35% in Table IV) could then reflect label misalignment rather than SAR's intrinsic difficulty. Additionally, the claimed '1.5 million segments' disagrees with Table II's sum of 1.4708M, and the text simultaneously reports 700 real labels and a 490-image evaluation set (14 per region), indicating the annotation pipeline's accounting is not fully reliable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces OpenEarthMap-SAR, a public benchmark dataset for sub-meter land cover segmentation from SAR imagery. It consists of 5,033 1024x1024 SAR images paired with optical imagery from 35 regions in Japan, France, and the USA, with 8-class land cover labels: pseudo labels generated by pre-trained OpenEarthMap models for all images, and manual annotations for a subset (the paper says 20 images per region, i.e., 700 images). The paper reports baseline semantic segmentation results for U-Net, SegFormer, and VMamba under Optical, SAR, and SAR+Optical modalities and five labeling scenarios (P, P+R1, P+R5, R1, R5), and positions the dataset as the official IEEE GRSS Data Fusion Contest Track I dataset.","tokens_in":11067,"tokens_out":7568,"duration_ms":74763,"significance":"The dataset addresses a real gap: there is no public sub-meter SAR land cover segmentation benchmark with aligned optical imagery and multiple geographic regions. The authors make the dataset publicly available, provide a fixed evaluation protocol, and benchmark three modern architectures. They also disclose the pseudo-label/manual agreement in Table III, which is a sign of transparency. However, the label validation is currently optical-centric and the annotation accounting is inconsistent, so the benchmark claims are not yet fully supported. If revised, this could be a useful community resource.","major_comments":[{"comment":"The label-quality validation is entirely optical-domain. Pseudo labels are generated with pre-trained OpenEarthMap models [21]-[25] trained on optical imagery, and the manual annotations are not stated to be drawn on SAR images; the paper only says experts manually aligned the paired optical and SAR datasets. Because Umbra Spotlight SAR at 0.15-0.5 m GSD is subject to layover, shadow, and speckle, residual geocoding misalignment can produce pixel-level label-SAR mismatches even after alignment. As a result, the low SAR-only mIoU values in Table IV (e.g., 35.13% for U-Net P and 34.74% for VMamba P) may reflect label misalignment rather than intrinsic SAR difficulty. Please state whether annotators labeled optical or SAR images, report a quantitative residual-alignment metric, and include at least one SAR-grounded label-consistency check (e.g., manual labels drawn directly on SAR or a visual comparison of label boundaries overlaid on SAR).","section":"Section II-B, Table III"},{"comment":"The annotation accounting is inconsistent. The text says 20 images per region were manually annotated, which gives 700 real labels, yet the evaluation set is described as 490 images containing 14 real labels per region, and the R1 and R5 training settings use 35 and 175 real labels, respectively. These numbers do not add up unless the 20 per-region annotations are further partitioned or some images are unused; the paper does not specify this partition. Please provide an exact per-region split table (manual annotations, evaluation, each training scenario) and reconcile the '700 real labels' statement with the split.","section":"Section II-B, Data split"},{"comment":"The abstract and conclusion claim '1.5 million segments,' but the segment counts in Table II sum to 1,470,800 (1.4708M), and the pixel counts sum to 5,267M, whereas 5,033 x 1,024 x 1,024 = 5,277M pixels. The differences may be due to rounding or to excluded no-data pixels, but this should be stated so the dataset statistics are reproducible.","section":"Table II and Abstract"},{"comment":"Bareland has near-zero IoU in almost all experiments (e.g., 0.00-0.16 for SAR and 0.00-1.47 for Optical in most settings) and the lowest pseudo/manual agreement in Table III (agreement 0.1063, IoU 0.0211). The paper reports this class among the eight benchmark classes without discussing why it is essentially unlearnable. This is relevant to the benchmark claim because it suggests the class definition, its rarity, or its label reliability is problematic; the authors should either provide a class-level failure analysis or revise the class set.","section":"Tables III and IV"}],"minor_comments":[{"comment":"The word 'pesudo' is a typo and should be 'pseudo'.","section":"Section II-B"},{"comment":"The class name is inconsistent: 'Agriculture land' in Table II but 'agricultural land' in the text and abstract. Please unify the terminology.","section":"Table II and text"},{"comment":"References [2] and [18] are the same article (Adriano et al., ISPRS J. Photogramm. Remote Sens., 2021) and should not be listed twice.","section":"References"},{"comment":"In the training scenario list, the phrase '4333 images, respectively' is unclear; specify whether these are the images without manual annotations and how this number relates to the 700 manually annotated images.","section":"Section II-B, Data split"},{"comment":"The phrase 'The dataset also serves the official dataset for IEEE GRSS Data Fusion Contest Track I' should be reworded to 'serves as the official dataset' for grammatical correctness.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of the journal and the dataset is potentially valuable. The main revision should focus on clarifying the annotation protocol and providing SAR-grounded label validation; the issues are fixable and do not warrant rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the resource: a public, sub-meter SAR land cover segmentation dataset with aligned optical imagery, 8 classes, 35 regions in three countries, and a fixed evaluation protocol. That has not existed at this scale, and it already has institutional weight as the official IEEE GRSS Data Fusion Contest Track I dataset. The paper also does two things well: it reports the gap between pseudo and manual labels honestly (Table III), and it admits in the conclusion that pseudo labels from optical models may introduce noise. Those are good signs.\n\nThe soft spots are real but not fatal to the resource. The stress-test note holds up on reading: every validation step compares labels against other labels derived from optical imagery, never against SAR-grounded truth. The manual annotation process is described only as \"manually annotated,\" with no statement about whether annotators drew on the optical or SAR image. Given the geocoding misalignment they had to correct, the absence of a residual misalignment metric is a genuine hole. Table III's 68% agreement and 2.1% bareland IoU measure consistency within the optical domain, not correctness against SAR geometry. So the low SAR-only mIoU could indeed be a label alignment problem rather than the intrinsic difficulty of SAR. That matters for anyone using the dataset to claim SAR is hard.\n\nThere are also accounting inconsistencies: the abstract says 1.5 million segments while Table II sums to about 1.47 million; the text says 20 manual annotations per region but the split says 14 real labels per region for evaluation; and the paper mentions 700 real labels while the split implies 490. These are the kind of things that need a correction pass before the benchmark is trustworthy.\n\nStill, the central deliverable is a downloadable dataset with a DOI, and the authors are transparent about its limitations. The \"global\" claim is overstated for three countries, but the geographic and class diversity is still a step beyond SpaceNet 6 or BRIGHT for land cover mapping. I would want a serious referee to look at this, mainly to force clarity on the annotation protocol and the numbers.\n\nRecommendation: send to peer review, conditional on the authors fixing the accounting and stating explicitly which image modality the annotators labeled. I'd cite this if I worked on multimodal land cover or SAR segmentation.","headline":"A useful public SAR land cover benchmark, but its label validation is optical-centric and the accounting has inconsistencies that need fixing before the benchmark claim is fully credible.","tokens_in":11521,"tokens_out":1224,"would_cite":true,"duration_ms":14980,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces OpenEarthMap-SAR, a public benchmark of 5,033 sub-meter SAR images with 8-class land-cover labels, designed to advance all-weather land cover mapping.","keywords":["synthetic aperture radar","land cover mapping","benchmark dataset","semantic segmentation","pseudo labels","high-resolution","all-weather mapping","missing modality"],"falsifier":"A fresh round of manual annotation of the same 700 tiles by a different set of experts, with per-pixel agreement computed against the existing manual labels, would settle whether the 'real' labels are a stable gold standard; if inter-annotator agreement is no higher than the 68.2% pseudo-label agreement, the benchmark's manual labels are too noisy to support reliable evaluation.","tokens_in":10512,"feed_emoji":"🛰️","tokens_out":7567,"duration_ms":73858,"temperature":0.7,"pith_summary":"OpenEarthMap-SAR is a new public dataset that pairs sub-meter synthetic aperture radar (SAR) imagery with optical imagery and 8-class land-cover labels. The authors aim to fill the gap left by SAR-specific semantic segmentation datasets, which are few, small, or limited to single classes. The dataset contains 5,033 1024×1024 images from 35 regions in Japan, France, and the USA, with ground sampling distances of 0.15–0.5 m. All images carry pseudo labels produced by pre-trained optical OpenEarthMap models, and 700 images were additionally manually annotated. If the dataset proves reliable, it gives the field a common testbed for all-weather, day-and-night land-cover mapping.","feed_headline":"Sub-meter radar land-cover benchmark opens all-weather mapping","feed_subtitle":"SAR sees through clouds; this dataset lets models map land cover when optical imagery fails.","key_machinery":"The load-bearing object is the dataset itself: OpenEarthMap-SAR, a collection of paired sub-meter SAR and optical images with 8-class segmentation labels and a predefined train/test split. Its design choices carry the argument: pseudo labels from pre-trained optical models give large-scale supervision without manual cost, manual labels provide a quality anchor, and the fixed evaluation protocol makes results from different methods comparable. The benchmark is what enables all-weather land-cover mapping research.","core_discovery":"The central claim is that OpenEarthMap-SAR constitutes a large-scale, geographically diverse benchmark for high-resolution SAR land-cover segmentation, with a resolution and class count that surpass previous SAR segmentation datasets. The authors support this by assembling 5,033 SAR images from Umbra's open-data catalog, pairing each with geocoded optical imagery from NAIP, IGN, and GSI, and aligning the pairs by hand. Labels use the eight OpenEarthMap classes (bareland, rangeland, developed space, road, tree, water, agricultural land, building), generated for all images from pre-trained optical models and corrected by manual annotation of 20 images per region. Baseline evaluations of U-Net, SegFormer, and VMamba across optical-only, SAR-only, and fused inputs show that the dataset is usable but challenging: SAR-only models reach about 35% mIoU, optical-only about 57–66%, and fusion generally improves water and agriculture segmentation. The authors position the dataset as the official track for the IEEE GRSS Data Fusion Contest and publish it on Zenodo.","pith_inferences":["Beyond the paper's claims, the low per-class agreement for bareland (2.1% IoU) suggests that the pseudo-label pipeline is unreliable for spectrally variable classes; a class-conditional confidence filter could improve the benchmark's usability.","We infer that the manual labels, at 20 images per region, are too few to train a robust model from scratch, so the dataset's practical value depends on semi-supervised or domain-adaptive methods that combine pseudo and real labels.","The observed gap between SAR-only (about 35% mIoU) and optical-only (about 57–66% mIoU) results implies that SAR-specific representation learning is still underdeveloped; the benchmark could be used to test self-supervised pre-training on radar imagery.","Because the SAR imagery comes from a single commercial provider (Umbra) in Spotlight mode, results may not transfer to other sensors or modes; we infer that multi-sensor SAR benchmarks remain an open need."],"forward_implications":["All-weather mapping becomes testable: models trained on this dataset can be evaluated for land-cover segmentation when optical imagery is unavailable due to clouds.","The fusion baselines give a reference point for future multi-modal SAR+optical models, with the paper reporting that fused inputs help classes like water and agricultural land.","The dataset's 35 regions across three continents support studies of geographic generalization and cross-domain adaptation.","The publicly released Zenodo repository lets researchers reproduce the baseline results and extend them with new methods.","As the IEEE GRSS Data Fusion Contest Track I official dataset, it provides a shared challenge for improving SAR segmentation."],"supporting_citations":[{"why":"Defines the 8-class label scheme and provides the pre-trained OpenEarthMap model family used to generate pseudo labels.","marker":"[21]"},{"why":"One of the OpenEarthMap model extensions used to create pseudo labels.","marker":"[22]"},{"why":"One of the OpenEarthMap model extensions used to create pseudo labels.","marker":"[23]"},{"why":"One of the OpenEarthMap model extensions used to create pseudo labels.","marker":"[24]"},{"why":"One of the OpenEarthMap model extensions used to create pseudo labels.","marker":"[25]"},{"why":"SpaceNet 6, the prior sub-meter SAR benchmark for building detection, which this dataset extends to multiclass land cover.","marker":"[19]"},{"why":"BRIGHT, a multimodal all-weather damage dataset that motivates SAR's role in disaster response and provides a comparison point.","marker":"[3]"},{"why":"FU-SAR, a meter-level SAR land cover dataset used as a comparison for task and class count.","marker":"[16]"},{"why":"GF-3 Building, a meter-level SAR building dataset used as a comparison.","marker":"[17]"},{"why":"BDD, a meter-level SAR change detection dataset used as a comparison.","marker":"[18]"}],"fun_headline_variants":["Radar benchmark maps global land cover through clouds","SAR sees through clouds to map land cover in 8 classes","OpenEarthMap-SAR: 1.5M radar segments for all-weather mapping","All-weather radar dataset for global land cover segmentation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's label quality rests on the premise that pseudo labels produced by optical-trained OpenEarthMap models are accurate enough to supervise SAR segmentation, even though the paper reports only 68.2% mean agreement with manual labels and near-zero IoU for bareland.","fun_headline_variants_meta":{"raw":{"variants":["Radar benchmark maps global land cover through clouds","SAR sees through clouds to map land cover in 8 classes","OpenEarthMap-SAR: 1.5M radar segments for all-weather mapping","All-weather radar dataset for global land cover segmentation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001759,"raw_usage":{"total_tokens":6993,"prompt_tokens":1043,"completion_tokens":5950,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":659,"completion_tokens_details":{"reasoning_tokens":5878}},"tokens_in":659,"tokens_out":5950,"duration_ms":39138,"temperature":1.0,"reasoning_tokens":5878,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:51:14.182600+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A fresh round of manual annotation of the same 700 tiles by a different set of experts, with per-pixel agreement computed against the existing manual labels, would settle whether the 'real' labels are a stable gold standard; if inter-annotator agreement is no higher than the 68.2% pseudo-label agreement, the benchmark's manual labels are too noisy to support reliable evaluation.","supporting_citations":[{"cited_title":"Generalized few-shot semantic segmentation in remote sensing: Chal- lenge and benchmark,","cited_arxiv_id":null,"evidence_quote":"One of the OpenEarthMap model extensions used to create pseudo labels."},{"cited_title":"Generating national very high-resolution land cover product of france without any labels: A comparative study,","cited_arxiv_id":null,"evidence_quote":"One of the OpenEarthMap model extensions used to create pseudo labels."},{"cited_title":"ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided Transformer,","cited_arxiv_id":null,"evidence_quote":"One of the OpenEarthMap model extensions used to create pseudo labels."},{"cited_title":"Spacenet 6: Multi-sensor all weather mapping dataset,","cited_arxiv_id":null,"evidence_quote":"SpaceNet 6, the prior sub-meter SAR benchmark for building detection, which this dataset extends to multiclass land cover."},{"cited_title":"Object-level semantic segmentation on the high-resolution gaofen-3 fuSAR-map dataset,","cited_arxiv_id":null,"evidence_quote":"FU-SAR, a meter-level SAR land cover dataset used as a comparison for task and class count."},{"cited_title":"A benchmark high-resolution gaofen-3 SAR dataset for building seman- tic segmentation,","cited_arxiv_id":null,"evidence_quote":"GF-3 Building, a meter-level SAR building dataset used as a comparison."},{"cited_title":"Learning from multimodal and multitemporal earth observation data for building damage mapping,","cited_arxiv_id":null,"evidence_quote":"BDD, a meter-level SAR change detection dataset used as a comparison."},{"cited_title":"Submeter-level land cover mapping of japan,","cited_arxiv_id":null,"evidence_quote":"One of the OpenEarthMap model extensions used to create pseudo labels."},{"cited_title":"Openearthmap: A benchmark dataset for global high-resolution land cover mapping,","cited_arxiv_id":null,"evidence_quote":"Defines the 8-class label scheme and provides the pre-trained OpenEarthMap model family used to generate pseudo labels."}],"review_version":1}