{"id":"f8e2e814-5422-48e4-8b74-a2fa1b15970a","arxiv_id":"2506.13649","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A 100-meter wall-to-wall map of 260 EUNIS level 3 habitats for Europe, generated by ensemble machine learning, with variable class-level accuracy.","lead":"Researchers trained machine-learning models on about 600,000 European vegetation plots to map 260 EUNIS habitat types at 100-meter resolution across Europe. The maps come with uncertainty scores and independent checks in the Netherlands, France and Austria, but many habitat classes show only modest accuracy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Dutch NLPT validation set is described as a 'hold-out' from the LVD database, which is a component of EVA used for training; the 'independent validation' claim may be circular.","rationale":"The reader's conditional verdict is appropriate. The paper provides a substantial, reproducible product with code, data, uncertainty layers, and two genuinely independent validation sets (Austrian habitat map and French forest inventory). The load-bearing weakness I identify is not the expert-system label quality per se, but the status of the Dutch NLPT dataset. The paper calls NLPT an 'independent habitat occurrence dataset' while also describing it as 'a hold-out of habitat observations from the Netherlands.' The LVD is a source database for EVA (ref. 29); thus a hold-out from LVD is not independent of the EVA training data unless explicit exclusion and independence are demonstrated. The manuscript lacks this demonstration. This matters because the strongest performance claims in the abstract rest on 'independent validation,' and the non-forest validation (grasslands, wetlands, scrub) largely relies on NLPT. If the Dutch data are in-distribution, the reported F1 scores (e.g., grasslands 0.40 on NLPT) may be optimistic. The proposed check—matching plot IDs/coordinates and auditing the code—would settle this. I therefore do not change the reader's conditional verdict, but I add a concrete condition: the authors must clarify the relationship between NLPT and EVA and, if needed, re-validate after strict exclusion.","tokens_in":21302,"tokens_out":7248,"duration_ms":73453,"concrete_test":"Check whether the NLPT (LVD) plots appear in the EVA training set by matching plot identifiers and coordinates (within a small tolerance) and by inspecting the code repository (github.com/bettasimousss/eunis-ml-mapping) for any exclusion step that removes NLPT/AT/IFN points before training. If NLPT overlaps the training data or was never excluded, rerun the entire pipeline with all Dutch plots removed from EVA and recompute Table 7 NLPT columns; if F1 scores drop substantially for grasslands and scrub, the 'independent validation' claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 'Habitat plots' states that 'Vegetation plots stored in the European Vegetation Archive (EVA) served as ground truth data for training and testing the model across Europe.' Section 'Habitat datasets for validation' then introduces 'two independent habitat occurrence datasets: a hold-out of habitat observations from the Netherlands (NLPT) and the French Forest Inventory (IFN).' The NLPT comes from the Landelijke Vegetatie Database (LVD), and reference 29 (Schaminée et al. 2012, 'The Dutch National Vegetation Database') is explicitly a source database for EVA. Calling a 'hold-out' from a component database 'independent' is contradictory. If the NLPT plots were simply withheld from the same EVA training set, this is a test split, not independent validation. Even if the 50k NLPT plots were removed from training, the remaining LVD/EVA data likely still contains thousands of Dutch plots, so spatial autocorrelation and identical sampling biases make the Dutch validation optimistic. The manuscript does not state that NLPT was excluded from EVA training, nor does it document any independent collection or labeling of NLPT. Consequently, the abstract's claim of 'independent validation' is only supported by the Austrian (AT) map and French IFN data. AT covers a single country and IFN covers forests only, leaving much of the non-forest validation (grasslands, wetlands, scrub) dependent on the possibly circular NLPT set. This directly affects the credibility of the stated F1 scores (Table 7), especially the modest values for grasslands and scrub, which would be even less reliable if the validation data share the training distribution.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a European wall-to-wall EUNIS habitat map at level 3 with 100 m resolution, produced by training formation-specific multi-class ensemble models on vegetation plots from the European Vegetation Archive (EVA) labeled by the EUNIS-ESy expert system, using climatic, topographic, edaphic, and remote-sensing predictors. The authors report spatial block cross-validation on EVA and external validation against Dutch (NLPT), Austrian (AT), and French forest (IFN) datasets, and they release probability, confidence, and top-habitat map products. The central claims are that the maps achieve strong predictive performance and that the validation includes genuinely independent datasets.","tokens_in":21595,"tokens_out":5549,"duration_ms":64958,"significance":"If the validation claims are supported, the product would be a valuable conservation and restoration planning resource: a continental, high-resolution, wall-to-wall EUNIS level 3 map with uncertainty layers, built from an unusually large training set and a transparent ensemble workflow. The paper also has concrete methodological strengths: spatial block cross-validation is used to mitigate spatial autocorrelation, multiple algorithm families are ensembled, class-imbalance corrections are compared per formation, and the code and data products are publicly available. However, the significance currently rests on two assertions that the evidence in the manuscript does not fully support: the independence of the Dutch validation set and the characterization of predictive performance as 'strong'. The class-level F1 scores reported in Table 7 are moderate for the most extensive formations, and external scores are considerably lower than the cross-validation scores.","major_comments":[{"comment":"The NLPT dataset is described as an 'independent habitat occurrence dataset' and as a 'hold-out' from the Dutch Landelijke Vegetatie Database (LVD), but reference 29 (Schaminée et al. 2012) is a national vegetation database that is used as a source for EVA, which is the training database for the models. The manuscript does not state that the NLPT plots were removed from EVA before training, nor does it document any independent field collection or labeling procedure for NLPT. Calling this set 'independent' is therefore not justified, and the abstract's claim of 'independent validation' is only supported by the Austrian map and the French IFN forest data. Since NLPT is the only external validation covering non-forest habitats in the Atlantic region, this issue directly affects the credibility of the reported grassland, wetland, and scrub F1 scores.","section":"Habitat datasets for validation"},{"comment":"The abstract and Technical Validation state that 'the habitat maps obtained strong predictive performances on the validation datasets', but the class-level F1 scores in Table 7 do not support that wording for the dominant formations. Mean forest F1 is 0.61 under EVA cross-validation but only 0.38 (NLPT), 0.33 (AT), and 0.48 (IFN); grassland F1 is 0.66 for EVA but 0.40 (NLPT) and 0.37 (AT); scrub and tundra F1 drops from 0.83 (EVA) to 0.46 (NLPT) and 0.31 (AT). These scores are modest, especially on the independent sets, and the large standard deviations (e.g., 0.32–0.39 for forests) indicate that many individual classes perform poorly. The manuscript should be revised to describe the results as moderate and to highlight the classes or formations where performance is genuinely strong, rather than claiming uniformly strong performance.","section":"Technical Validation, Table 7"},{"comment":"The regional masks used in Step 2 are computed from the same EVA vegetation plot occurrences that are also used for cross-validation, but the manuscript does not state whether the EVA cross-validation scores in Tables 7–14 are computed before or after applying these masks. If the masks are applied during validation, then each validation plot's own occurrence contributes to the mask for its ecoregion, which makes the EVA F1 scores optimistic relative to a fully forward prediction. Please clarify the validation protocol and, if masking is included, quantify the effect of the regional masks on the EVA and external validation scores.","section":"Step 2: Regional filtering rules"},{"comment":"The statement that low performance 'improved when considering the top three predictions rather than only the most likely class' is not supported by any quantitative comparison in the manuscript. No top-1 versus top-3 accuracy, precision, or F1 table is provided, nor is the improvement quantified for the classes or formations where it is claimed. Since this statement is used to mitigate the low F1 scores, the authors should add a table or figure comparing top-1 and top-3 validation results for at least the independent datasets.","section":"Technical Validation, final paragraph"}],"minor_comments":[{"comment":"The sentence 'Decision trees excel with structured tabular data and neural networks with intricate feature interactions...' appears twice, once at the end of the 'Ensemble model training' subsection and once immediately before 'Ensemble forecasting and uncertainty'; the duplicate should be removed.","section":"Methods, Ensemble model training"},{"comment":"The citation for LDAM loss (reference 47) points to Cao, Larsen, and Thorne (2001), a paper on rare species in multivariate analysis, which appears to be the wrong reference; the LDAM loss is from Cao et al. (2019), 'Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss' (NeurIPS).","section":"References, LDAM loss"},{"comment":"Table 7 contains a stray 'Strategy' row in the header and does not report results for the man-made (V) formation, although Table 1 and the mapping workflow include V; please check the table formatting and clarify whether V results are omitted because no validation data were available.","section":"Table 7"},{"comment":"The text says the domain is divided into 100 km x 100 km cells that are partitioned into five blocks, and later says '20% of the observations' are hidden per iteration; with five blocks the hidden fraction is 20%, so this is consistent, but the wording should be made explicit to avoid confusion.","section":"Methods, Spatial Block-CV partitioning"},{"comment":"Table 2 lists predictors at 1 km, 500 m, and 100 m resolutions, but the manuscript does not describe how these were resampled or harmonized to the 100 m prediction grid; please add a short paragraph on resampling, reprojection, and temporal matching of the predictor layers.","section":"Environmental predictors"},{"comment":"Table 6 lists 'P' (inland waters) among the formations for which the same MLP configuration was selected, but Table 1 and the Methods exclude inland waters; this is likely a typo for 'V' and should be corrected.","section":"Tables 1 and 6"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a serious data product, not a methodological breakthrough, and it deserves peer review. The integrated 100-m EUNIS level 3 map with 260 classes and uncertainty layers is new, and the authors have done real work assembling training data, running careful spatial block CV, and shipping code and maps. The abstract oversells the validation, and the Dutch validation set is not as independent as advertised.\n\nWhat is genuinely new is the product itself: wall-to-wall coverage of 260 habitats at 100 m, with probability cubes, top-3 maps, and uncertainty layers. Using the EUNIS-ESy expert system to label ~600k EVA plots is a transparent pipeline. The spatial block CV is sound. The Austrian and French validation sets are genuinely external, and comparing against the Austrian 10-m map is a nice independent check. Code and data are public, which is real credit.\n\nThe headline problem is the NLPT set. It is described as independent, but it is a hold-out from the Landelijke Vegetatie Database, and LVD is a component of EVA. The paper never states that the NLPT plots were excluded from training. So the Dutch F1 scores likely share spatial autocorrelation and sampling bias with the training set. That leaves grassland, scrub, and wetland validation leaning on a quasi-independent set. The abstract's \"strong predictive performances\" is also a stretch: independent F1 means are 0.38–0.48 for forests, 0.40 for grasslands, 0.46 for scrub. Those are modest, and the paper reports them. Also, there is no quantitative comparison to existing products, so the \"enhancing\" claim is not backed by numbers. The post-hoc regional masks use EVA occurrences, a mild circularity.\n\nNone of this is fatal. The map is likely the best available wall-to-wall product at this thematic detail. The fixes are straightforward: soften the abstract, document NLPT exclusion or drop it from \"independent\" claims, and add a quantitative baseline comparison.\n\nWho it is for: applied ecologists, conservation agencies, EUNIS users, and people working on EU Nature Restoration Law reporting. It is not a methods paper; the ML is standard. A serious referee should see it, but expect major revision before acceptance.\n\nMy vote: send it to peer review, with the clear expectation that the validation language gets corrected.","headline":"A genuinely useful continental habitat map product, but the 'independent' validation is partly a hold-out from the same database family and the abstract overstates accuracy.","tokens_in":22287,"tokens_out":2284,"would_cite":true,"duration_ms":24011,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Machine learning maps 260 European habitats at 100-meter resolution across Europe.","keywords":["EUNIS habitats","habitat mapping","machine learning","ensemble modelling","remote sensing","vegetation plots","spatial cross-validation","conservation planning"],"falsifier":"Score the published map against a national habitat map from a country not used in validation (for example, Spain or Denmark) at 100 m resolution; if common habitat classes there score near-zero F1 despite good cross-validation, the continental-scale claim would be falsified. A cheaper check is to count classes with F1 below 0.3 in the provided Austrian 10 m validation, since matching a finer national reference is the hardest test the product faces.","tokens_in":21104,"feed_emoji":"🗺️","tokens_out":8159,"duration_ms":85932,"temperature":0.7,"pith_summary":"The paper sets out to show that Europe can be mapped wall-to-wall for 260 EUNIS habitat types at hierarchical level 3 using machine learning, producing a product that supports conservation policy and the Nature Restoration Law. The authors train an ensemble of tree-based and neural-network classifiers on nearly 600,000 georeferenced vegetation plots whose species compositions were assigned to EUNIS level 3 classes by an expert system, using climate, topography, soil, and satellite-derived predictors at 100-meter resolution. They validate the resulting maps against independent plot datasets from the Netherlands, France, and Austria and against a spatial-block cross-validation, reporting good but uneven performance with different recall-precision trade-offs by habitat formation. If the maps hold, planners gain a spatially explicit, uncertainty-aware baseline for where habitats occur and where they do not, something that has been missing at continental scale.","feed_headline":"Machine learning maps 260 European habitats at 100 meters","feed_subtitle":"An ensemble trained on 600,000 vegetation plots adds confidence layers for conservation and restoration planning.","key_machinery":"The load-bearing mechanism is a per-formation ensemble of multi-class classifiers: for each of the terrestrial EUNIS level 1 groups (saltmarshes, coastal, wetlands, grasslands, shrublands and scrub and tundra, forests, sparsely vegetated, and man-made), a separate model predicts level 3 classes from environmental and satellite predictors, and the models from different algorithm families and spatial cross-validation folds are combined by weighted voting. Two rule-based filters make the continuous probabilities into a map: regional masks based on ecoregions and coastline distance keep predictions inside each class's known range, and a land-cover crosswalk with priority rules decides the prevailing level 1 formation and the final level 3 class at each pixel. Uncertainty is quantified at each pixel as the committee-averaging score and as model- and fold-induced disagreement.","core_discovery":"The central claim is that a wall-to-wall map of European terrestrial habitats at EUNIS level 3 is attainable by training separate multi-class machine-learning ensembles within each EUNIS level 1 formation, with roughly 597,819 vegetation plots as ground truth once their species lists are translated to level 3 by the EUNIS-ESy expert system. Instead of fitting one binary model per habitat, the authors jointly model the classes in each formation so less common habitats can borrow information from more common ones, then combine tree-based, boosting, and neural-network classifiers using weighted voting over spatial cross-validation folds. The resulting 100 m maps carry per-pixel probabilities, top-3 classes, and confidence scores, and the validation results show high F1 scores for saltmarshes, sparsely vegetated, coastal, and wetland habitats, with grasslands, shrublands, and forests more variable and with different recall-precision trade-offs on independent data from the Netherlands, France, and Austria.","pith_inferences":["A testable extension not in the paper would be to evaluate the full pipeline on pixels where a secondary formation is nearly as probable as the winner, since errors in choosing the level 1 formation would then propagate into the final level 3 label.","The three validation territories cover Atlantic, Alpine, and parts of Mediterranean and Continental Europe; a decisive next test is Scandinavia, Iberia, or Eastern Europe, where plot densities and habitat combinations differ.","Adding the missing predictors the authors flag, such as soil moisture, land-use history, and human footprint, and then re-scoring the low-performing humid-soil and abandoned-land classes would show whether the remaining errors are data gaps or model limitations."],"forward_implications":["A user can download continuous probability layers for all 260 classes and use the top-3 list to identify pixels where the first choice is fragile, making the map usable for targeting field surveys.","Conservation and restoration planners can overlay the confidence maps to distinguish well-predicted habitats from uncertain edge cases, instead of treating the single most probable class as truth.","Because the mapping workflow separates probabilities from rule-based filtering, substituting a finer land-cover layer than the one used here would yield a finer final product without retraining the models.","The per-formation ensembles make it possible to update or add habitat classes within one formation without retraining the whole continent.","For formations with few classes or strong abiotic control, per-class F1 scores above 0.9 are common, so the map is already usable for saltmarsh and sparsely vegetated habitats."],"supporting_citations":[{"why":"Supplies the nearly 597,819 vegetation plots used as training and cross-validation ground truth.","marker":"[10]"},{"why":"The EUNIS-ESy expert system that assigns each plot's species composition to a EUNIS level 3 class, defining the modeled target.","marker":"[28]"},{"why":"Defines the revised EUNIS classification and the expert-system basis that the labels and formation structure rely on.","marker":"[4]"},{"why":"Justifies the spatial block cross-validation design used to avoid inflated scores from spatial autocorrelation.","marker":"[39]"},{"why":"Provides the ensemble-forecasting rationale for combining model and data-sampling uncertainty.","marker":"[33]"},{"why":"Independent Dutch vegetation-plot dataset used for external validation.","marker":"[29]"},{"why":"Independent French forest-inventory data used for external validation.","marker":"[31]"},{"why":"Austrian 10 m habitat map used for independent regional validation.","marker":"[32]"}],"fun_headline_variants":["260 European habitats mapped at 100-m resolution with AI","AI map pinpoints 260 European habitats at fine scale","Machine learning delivers Europe-wide habitat map at 100 m","New AI map shows 260 habitats across Europe with confidence","High-res AI map of Europe's 260 habitats for conservation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the expert-system assignments used to train and validate the models are correct enough to serve as ground truth; if those labels are wrong, high validation scores would not transfer to the real world.","fun_headline_variants_meta":{"raw":{"variants":["260 European habitats mapped at 100-m resolution with AI","AI map pinpoints 260 European habitats at fine scale","Machine learning delivers Europe-wide habitat map at 100 m","New AI map shows 260 habitats across Europe with confidence","High-res AI map of Europe's 260 habitats for conservation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00045,"raw_usage":{"total_tokens":2255,"prompt_tokens":921,"completion_tokens":1334,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":1252}},"tokens_in":537,"tokens_out":1334,"duration_ms":11038,"temperature":1.0,"reasoning_tokens":1252,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:28:28.362395+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Score the published map against a national habitat map from a country not used in validation (for example, Spain or Denmark) at 100 m resolution; if common habitat classes there score near-zero F1 despite good cross-validation, the continental-scale claim would be falsified. A cheaper check is to count classes with F1 below 0.3 in the provided Austrian 10 m validation, since matching a finer national reference is the hardest test the product faces.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the nearly 597,819 vegetation plots used as training and cross-validation ground truth."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The EUNIS-ESy expert system that assigns each plot's species composition to a EUNIS level 3 class, defining the modeled target."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the revised EUNIS classification and the expert-system basis that the labels and formation structure rely on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies the spatial block cross-validation design used to avoid inflated scores from spatial autocorrelation."},{"cited_title":"& Grenouillet, G","cited_arxiv_id":null,"evidence_quote":"Provides the ensemble-forecasting rationale for combining model and data-sampling uncertainty."},{"cited_title":"& Ozinga, W","cited_arxiv_id":null,"evidence_quote":"Independent Dutch vegetation-plot dataset used for external validation."},{"cited_title":"Données brutes, Campagnes annuelles 2005 et suivantes","cited_arxiv_id":null,"evidence_quote":"Independent French forest-inventory data used for external validation."},{"cited_title":"MAES/EUNIS habitat map Austria 10m","cited_arxiv_id":null,"evidence_quote":"Austrian 10 m habitat map used for independent regional validation."}],"review_version":1}