{"id":"326ddc40-1340-427e-832c-ca8d4f90c302","arxiv_id":"2501.10600","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A regression U-Net maps Amazon canopy height at 4.78 m from Planet NICFI images with a 3.68 m mean absolute error, producing a 2020-2024 Amazon height map and a ~22 m mean canopy height estimate.","lead":"The authors trained a U-Net to estimate tree canopy height across the Amazon from Planet satellite images, using aerial LiDAR as the teaching signal, and produced a 4.78 m resolution height map for 2020-2024. It reports a mean error of 3.68 m against held-out LiDAR, and finds the Amazon's average canopy height to be about 22 m.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 3.68 m validation MAE is measured on patches from the same LiDAR/NICFI tiles as training, with labels up to ±6 months older than images, so it is not independent Amazon-wide accuracy.","rationale":"","tokens_in":1143,"tokens_out":1664,"duration_ms":77405,"concrete_test":"Acquire GEDI L2B shots over the Amazon for 2019–2023 and apply standard quality filters (quality_flag=1, degrade_flag=0, sensitivity ≥0.95, slope ≤10°). Exclude any footprint within a 2 km buffer of the LiDAR flight lines used in this study. For each footprint, average the model's predicted heights over a 25 m radius matching the GEDI footprint, using the NICFI mosaic closest in time (preferably within ±1 month). Compare against GEDI RH85 or RH90, and also for the same footprints against Tolan, Lang, and Pauls. Report MAE, RMSE, and bias separately for the four Amazon regions in Figure 1. If the independent MAE is below about 5 m and regional biases within ±2 m, the generalization concern is largely resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central accuracy claim rests on a validation set that is neither spatially nor temporally independent. In Section 2.7, validation is created by selecting one 256×256 patch from each NICFI tile that overlaps the LiDAR CHM, while the remaining ~47,000 patches come from those same tiles and the same ~24 LiDAR sites. Because neighboring patches in a tile share forest type, illumination, and sensor calibration, the held-out patches are not independent of training. The validation MAE thus measures in-distribution fit, not accuracy over the full Amazon, where forest structure and imaging conditions (e.g., Guiana Shield, western Amazon, Andes slopes) are underrepresented or absent. Compounding this, the paper pairs each CHM with the nearest NICFI image and also includes images from the previous and subsequent dates, allowing a time window of up to one year (±6 months for biannual NICFI, ±1 month for monthly). For LiDAR flights from 2015–2018, canopy growth, natural disturbance, or logging within that window can corrupt the labels. The reported 3.68 m therefore estimates agreement with potentially corrupted, spatially autocorrelated labels rather than true canopy height at the image date. All comparisons to Tolan, Lang, and Pauls are made on the same validation patches, so the claimed superiority over global products is not established outside the training domain. The paper's own caveats about cloud/shade artifacts and the Andes suggest the map is less reliable in some regions, but no external validation against GEDI or independent LiDAR is provided.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains a U-Net regression model on Planet NICFI RGB-NIR imagery (4.78 m) with LiDAR-derived canopy height labels to produce a 2020-2024 Amazon canopy height map. Validation MAE is reported as 3.68 m, and the model is compared against three global canopy height products (Tolan, Pauls, Lang). The paper also demonstrates height-based detection of logging, deforestation, and regeneration, and reports an area-weighted mean Amazon canopy height of 22.09 m. The central claims are the accuracy and superiority of the model in the Amazon and the usability of the resulting map.","tokens_in":38703,"tokens_out":5394,"duration_ms":53429,"significance":"If the accuracy claim held under independent evaluation, this would be a valuable high-resolution forest structure product for a data-poor region, with potential applications to carbon monitoring and conservation. The paper uses publicly available LiDAR and NICFI data and provides a detailed description of the preprocessing and training pipeline, which are strengths. However, the validation is not spatially or temporally independent, and the Amazon-wide mean is given without uncertainty, so the headline numbers should be treated as provisional pending additional validation.","major_comments":[{"comment":"The validation is constructed by selecting one 256x256 patch from each NICFI tile that overlaps the LiDAR CHM, while the remaining ~47,000 training patches come from the same tiles and the same LiDAR programs/sites. Because neighboring patches in a tile share forest structure, illumination, and sensor calibration, the held-out patches are not spatially independent of training. The reported MAE of 3.68 m therefore measures in-distribution fit, not accuracy over the full Amazon, and the same limitation applies to the comparisons with Tolan, Pauls, and Lang in Section 3.1. Please provide a leave-one-site-out or spatially disjoint evaluation, and re-run the product comparisons on those held-out regions.","section":"Section 2.7"},{"comment":"The procedure pairs each LiDAR CHM with the closest NICFI image and also includes images from the previous and subsequent dates, allowing a time window of up to one year (±6 months for biannual NICFI, ±1 month for monthly). Height changes from growth, logging, or natural disturbance within that window corrupt the training and validation labels, and because such events are spatially correlated with forest types and disturbance history, the learned image-to-height mapping can be biased. Please quantify the sensitivity of the reported accuracy to this window, for example by restricting the validation to image-label pairs within one month and by excluding pixels with known disturbance from the analysis.","section":"Section 2.7"},{"comment":"The Amazon mean canopy height of 22.09 m is reported without any uncertainty or correction for known model biases, including underestimation above 50 m (Section 3.1), cloud/shade artifacts (Section 3.7), and the relaxed cloud masking applied in low-observation regions such as the Guiana Shield and western Amazon (Section 3.8). Please provide an error budget, confidence intervals on the mean and percentiles, or a comparison with independent regional height estimates.","section":"Section 3.9 and Table 1"},{"comment":"The claim of outperforming existing global products is established only on validation patches drawn from the training domain; this does not demonstrate superiority over the full Amazon, particularly in underrepresented regions such as the Guiana Shield, the western Amazon, and the Andes. Please report per-region accuracy (for example, the four Feldpausch regions of Figure 1) and, ideally, validation on LiDAR or GEDI data not used in training.","section":"Section 3.1"}],"minor_comments":[{"comment":"The caption mentions Tolan's model and Lang's model but the figure has four panels including Pauls's model; please correct the caption.","section":"Section 3.1, Figure 3 caption"},{"comment":"The text contains repeated phrases such as 'The the Guiana Shield' and 'in the Shield'; please edit.","section":"Section 3.9, Figure 12"},{"comment":"The heading 'T able 1' contains a stray space; please fix.","section":"Table 1"},{"comment":"Please state explicitly how many unique LiDAR tiles/sites contribute to training versus validation, and whether any training patch is taken from the same NICFI tile as a validation patch.","section":"Section 2.7"},{"comment":"Please define n, i, and the exact condition for a 'valid observation' (height above 5 m and non-cloud observation count above 5) in the text.","section":"Equation 1"},{"comment":"The decision to keep the dataset non-open is explained, but the Data Availability section should state whether the model weights, the prediction code, or the derived height map will be released, as this affects reproducibility.","section":"Section 4.2 / Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an application paper; its novelty relative to the authors' earlier California U-Net work is incremental, but the Amazon-scale product is of practical value. The main risk is that the headline accuracy is not supported by an independent validation; if the authors can provide spatially and temporally independent validation, the paper would be acceptable. The non-open data policy (Section 4.2) may conflict with journal data availability norms."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline result is real: this paper delivers the first wall-to-wall 4.78 m canopy height map for the Amazon (2020–2024) and a mean height of 22.09 m. That is a new data resource with obvious value for carbon accounting, degradation monitoring, and tall-tree discovery. The U-Net architecture is not novel—it carries over directly from the authors' California work—but the Amazon product and its public-facing descriptions of height patterns are an actual contribution.\n\nThe paper does several things well. The methods are described clearly enough to reproduce the training pipeline, the LiDAR sources are public, and the comparisons against Tolan, Lang, and Pauls on the same validation patches are honest, even if that shared evaluation frame cuts both ways. The qualitative examples of logging and regrowth detection are encouraging, and the discussion openly acknowledges the Andes/cloud limitations and the ethical choice not to open giant-tree locations. These are signs of careful, honest work.\n\nWhere the paper is soft is in the evidence for the central accuracy claim. The validation patches are drawn from the same NICFI tiles and the same ~24 LiDAR sites used for training; they are not spatially independent. The MAE of 3.68 m is an in-distribution fit, not a measurement of accuracy over the full Amazon, where Guiana Shield or western Amazon conditions are underrepresented. The temporal pairing also allows up to a ±6 month mismatch between LiDAR and image, so a logged or regrowing patch could be labelled with stale height. The paper asserts the static-height assumption but does not test it. There is no external validation against GEDI or independent LiDAR, and the Amazon-wide mean has no uncertainty attached.\n\nNone of this makes the product useless, and I do not think the paper hides these issues—they are largely implicit in the methods. But the abstract's 'outperforming existing global products' claim is only demonstrated within the training domain. A serious referee should ask for a GEDI or hold-out-region check before publishing the accuracy numbers as headline findings.\n\nWho is this for? Remote sensing scientists and forest ecologists who want a very high-resolution Amazon height layer and are willing to treat the absolute accuracy as provisional. It deserves a proper peer review, but with a request for independent validation and uncertainty estimates. I would bring it to a reading group, and I would cite it with a caveat on validation.","headline":"A useful but not yet fully validated Amazon canopy height product; the claimed 3.68 m accuracy is in-distribution, not independent Amazon-wide.","tokens_in":39345,"tokens_out":1621,"would_cite":true,"duration_ms":19587,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A U-Net regression model trained on Planet NICFI imagery with LiDAR references maps Amazon forest canopy height at ~4.78 m resolution with a 3.68 m mean absolute error, beating global height products.","keywords":["canopy height","U-Net","Planet NICFI","LiDAR","Amazon forest","deep learning regression","forest monitoring","regeneration detection"],"falsifier":"Collect fresh airborne LiDAR transects in the western Amazon and Guiana Shield within two weeks of a NICFI basemap and compare predictions; if the model's error against this independent LiDAR is systematically larger than 3.68 m, the claimed Amazon-wide low bias is falsified.","tokens_in":38221,"feed_emoji":"🌳","tokens_out":12307,"duration_ms":107593,"temperature":0.7,"pith_summary":"This paper tries to establish that the height of the Amazon forest can be mapped wall-to-wall from freely available Planet NICFI satellite mosaics at a resolution fine enough to see individual tree crowns. A U-Net regression model translates 4.78 m RGB-NIR images into canopy height, with airborne LiDAR canopy height models as training reference. On a validation sample spread across the Amazon, the model's mean absolute error is 3.68 m, and bias stays low across most of the height range, with little saturation up to 40-50 m. If correct, this yields an Amazon-wide 4.78 m height map for 2020-2024 with an average canopy height near 22 m, and makes logging, deforestation, and regrowth visible as height changes in time series.","feed_headline":"Deep learning maps Amazon tree heights with 3.68-m error","feed_subtitle":"A U-Net trained on Planet NICFI images and LiDAR beats global height maps and puts the forest's mean canopy at about 22 m.","key_machinery":"The load-bearing object is the U-Net architecture adapted for regression: an encoder-decoder convolutional network with skip connections and roughly 35 million parameters that takes 256 × 256 × 4 RGBNIR tiles at 4.78 m and outputs a single-band per-pixel height map rescaled from 0-1 to 0-100 m. It is trained with mean squared error weighted by a binary presence/absence mask of the LiDAR canopy height model, so training focuses on pixels where reference data exist. The paired reference was built by computing canopy height models from airborne LiDAR at 1 m and resampling them to 4.78 m with the median, and the same-date plus neighboring NICFI mosaics were used as inputs under a no-change assumption. The skip connections are what allow crown texture and local context to survive through the decoder, which is why the model can reproduce crown-level detail instead of only smooth height averages.","core_discovery":"The central discovery is that a locally trained regression U-Net can recover LiDAR-grade canopy height from 4.78 m Planet NICFI imagery across the whole Amazon, where global canopy height products saturate or blur. The paper reports a mean absolute error of 3.68 m on 3,436 validation tiles, close alignment of median predicted height with LiDAR medians, and reliable estimates up to 45-50 m, while a 0.5 m RGB-based global product, a 10 m Sentinel-2 product, and a 10 m Sentinel-1/2 product show stronger saturation and larger errors in this region. From these predictions, the paper derives a weighted mean Amazon canopy height of 22.09 m, a median of 22.25 m, and a 97.5th percentile of 32.10 m, and maps a ring of the tallest forests around the central Amazon and individual giant trees with crowns wider than 50 m.","pith_inferences":["I infer that the same recipe should transfer to other NICFI-covered tropical forests in Africa and Southeast Asia, because the image source is pan-tropical and the model is locally trainable; the main bottleneck will be local LiDAR reference data for calibration.","I infer that the monthly NICFI cadence could be used to build dense height time series that turn regeneration curves into carbon-recovery estimates without repeated LiDAR or commercial data, an extension the paper only sketches.","I infer that the reported 3.68 m error likely understates true error in recently disturbed areas, because the training assumption of unchanged height over up to one year corrupts labels where logging or regrowth occurred between LiDAR and satellite acquisition.","I infer that blending the optical U-Net with radar-based height estimates would improve cloud-dominated zones such as the Andes slopes, where the paper itself notes its product is weakest."],"forward_implications":["The Amazon forest has a weighted mean canopy height of about 22 m, with the tallest forests forming a roughly 1,000 km-wide arc around the central Amazon and a hotspot in the Guiana Shield.","Selective logging becomes visible as month-to-month negative height differences in Planet NICFI time series, allowing disturbance to be located even when it does not clear the canopy.","Deforestation and regrowth can be followed from height trajectories: cleared pixels fall near zero and stay low, while abandoned pasture shows gradual height increase reaching 15-20 m by 2024.","Individual giant trees, including crowns 50-70 m wide, can be identified from the height map, enabling automatic searches for the Amazon's largest trees.","On the same validation tiles, the locally calibrated U-Net outperforms three 2020 global canopy height maps, especially for heights above 30-40 m where those products saturate."],"supporting_citations":[{"why":"This reference defines the U-Net architecture that the regression model adapts for per-pixel canopy height estimation.","marker":"Ronneberger et al., 2015"},{"why":"This reference supplies the Planet NICFI basemap product and its 4.78 m resolution, monthly/biannual temporal coverage, and band specifications used as model input.","marker":"Planet Team, 2017"},{"why":"This reference provides the Sustainable Landscapes LiDAR point clouds, a main source of canopy height reference data for training.","marker":"Dos-Santos et al., 2022"},{"why":"This reference provides the EBA airborne LiDAR transects across the Brazilian Amazon used in training and validation.","marker":"Ometto et al., 2023"},{"why":"This reference provides the lidR processing chain used to turn raw LiDAR point clouds into canopy height models.","marker":"Roussel et al., 2020"},{"why":"This reference supplies the 0.5 m RGB-based global canopy height product that the validation compares against and outperforms.","marker":"Tolan et al., 2024"},{"why":"This reference supplies the 10 m Sentinel-2 global canopy height product used as a comparison baseline in validation.","marker":"Lang et al., 2023"},{"why":"This reference supplies the 10 m Sentinel-1/2 global canopy height product used as a comparison baseline in validation.","marker":"Pauls et al., 2024"},{"why":"This reference provides the deforestation/regrowth time series used to date events and interpret height changes.","marker":"Wagner et al., 2023"},{"why":"This reference provides the logging detection from NICFI images that is used to locate logging events for height-change analysis.","marker":"Dalagnol et al., 2023"}],"fun_headline_variants":["U-Net from Planet and LiDAR maps Amazon tree heights with 3.68-m error","Deep learning outperforms global models for Amazon tree height mapping","4.78-m resolution Amazon canopy height map from deep learning","Deep learning yields 3.68-m-accurate Amazon canopy height maps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that forest canopy height stayed unchanged between the LiDAR acquisition and the paired NICFI image, including images taken up to a year earlier or later.","fun_headline_variants_meta":{"raw":{"variants":["U-Net from Planet and LiDAR maps Amazon tree heights with 3.68-m error","Deep learning outperforms global models for Amazon tree height mapping","4.78-m resolution Amazon canopy height map from deep learning","Deep learning yields 3.68-m-accurate Amazon canopy height maps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00107,"raw_usage":{"total_tokens":4498,"prompt_tokens":978,"completion_tokens":3520,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":594,"completion_tokens_details":{"reasoning_tokens":3440}},"tokens_in":594,"tokens_out":3520,"duration_ms":24156,"temperature":1.0,"reasoning_tokens":3440,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:02:30.611522+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect fresh airborne LiDAR transects in the western Amazon and Guiana Shield within two weeks of a NICFI basemap and compare predictions; if the model's error against this independent LiDAR is systematically larger than 3.68 m, the claimed Amazon-wide low bias is falsified.","supporting_citations":[{"cited_title":"title Planet application program interface: In space for life on earth","cited_arxiv_id":null,"evidence_quote":"This reference supplies the Planet NICFI basemap product and its 4.78 m resolution, monthly/biannual temporal coverage, and band specifications used as model input."},{"cited_title":", author Keller, M","cited_arxiv_id":null,"evidence_quote":"This reference provides the Sustainable Landscapes LiDAR point clouds, a main source of canopy height reference data for training."},{"cited_title":", author Yang, H.I","cited_arxiv_id":null,"evidence_quote":"This reference supplies the 0.5 m RGB-based global canopy height product that the validation compares against and outperforms."},{"cited_title":", author Zimmer, M","cited_arxiv_id":null,"evidence_quote":"This reference supplies the 10 m Sentinel-1/2 global canopy height product used as a comparison baseline in validation."},{"cited_title":", author Dalagnol, R","cited_arxiv_id":null,"evidence_quote":"This reference provides the deforestation/regrowth time series used to date events and interpret height changes."},{"cited_title":", author Wagner, F.H","cited_arxiv_id":null,"evidence_quote":"This reference provides the logging detection from NICFI images that is used to locate logging events for height-change analysis."}],"review_version":1}