{"id":"ad1caf1f-b8bd-48b2-8eba-30d7bce8e556","arxiv_id":"2509.04772","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"FloodVision uses GPT-4o plus a knowledge graph of object heights to estimate urban flood depth from RGB images, achieving 8.17 cm MAE on 110 crowdsourced images.","lead":"FloodVision combines GPT-4o's image understanding with a knowledge graph of real-world object heights to estimate flood water depth from ordinary photos. It reports roughly 8 cm average error on 110 crowdsourced flood images, suggesting a low-cost way to gauge flood severity without special sensors.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Evaluation rests on unvalidated citizen-reported depths; if those reports are biased or noisy, the reported 8.17 cm MAE and 20.5% improvement measure agreement with an unvalidated human guess, not physical flood depth.","rationale":"The paper's central claim is an empirical accuracy number, and the weakest link is the ground truth. The reader flagged exactly this assumption. My concern is more specific: the magnitude of the claimed improvement is 2.11 cm, which is smaller than typical rounding or precision of citizen-reported depths. Without an uncertainty model for the proxy labels, the 20.5% reduction is not distinguishable from label noise. The cross-dataset comparison to prior CNN methods is also not a controlled comparison, so the 'surpassing prior CNN-based methods' claim is not established by the reported numbers. This is an empirical validation issue rather than an internal inconsistency; no formal verification is claimed, and the knowledge graph construction is plausible. I agree with the reader's conditional verdict: the paper should not be accepted unconditionally until the proxy ground truth is validated against independent measurements. If the proposed test shows the citizen reports are reliable within a couple of centimeters, the conditional can be lifted; otherwise the headline accuracy figures would need to be substantially revised. Thus my recommendation is to keep the reader's verdict unchanged.","tokens_in":5715,"tokens_out":5175,"duration_ms":49199,"concrete_test":"Select a subset of 20-30 MyCoast NY flood reports with timestamps and coordinates; for each, obtain an independent depth reference from the nearest USGS gauge reading within a narrow time/location window, a surveyed high-water mark, or a visible staff gauge in the image. Compute bias and RMS error of the resident-reported depths against these references. Then replace the proxy labels with the independent references (or apply a measurement-error correction) and recompute FloodVision vs. GPT-4o baseline MAE. If the independent references differ from the citizen reports by more than about 2 cm, or if the recomputed improvement loses significance, the headline claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 states that resident-reported depth estimates from MyCoast New York are \"treated as proxy ground truth values,\" with no independent water-level measurements, no sensor/gauge comparison, and no screening for reporting bias. The entire quantitative evaluation in Table 1 and the abstract's headline numbers inherit this label. If citizen reports are rounded (e.g., to whole inches), systematically biased by the visual severity of the scene, or influenced by the same visual cues that GPT-4o uses, then the 2.11 cm MAE improvement over the GPT-4o baseline is within plausible label noise. The claimed superiority over prior CNN methods is additionally weakened because those results come from different datasets with different label definitions (Section 3.2). The load-bearing condition is therefore not the model architecture or knowledge graph; it is that the proxy labels faithfully represent physical water depth. The paper provides no evidence for this condition, and the 20.5% improvement is smaller than typical self-report rounding increments, making the central claim vulnerable to label noise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"FloodVision proposes a zero-shot flood depth estimation framework that combines GPT-4o's visual reasoning with a curated knowledge graph (FloodKG) of canonical object dimensions. The system identifies visible reference objects (e.g., car tires, curbs), retrieves physical heights from FloodKG, estimates submergence ratios via prompt-guided GPT-4o, and aggregates per-object depth estimates with outlier filtering. Evaluated on 110 crowdsourced images from MyCoast New York, the average estimate achieves a mean absolute error of 8.17 cm, a 20.5% improvement over a GPT-4o-only baseline, and is claimed to surpass prior CNN-based methods. The paper also discusses integration with digital twins and emergency response.","tokens_in":6047,"tokens_out":2660,"duration_ms":24135,"significance":"If the evaluation is trustworthy, FloodVision contributes a practical zero-shot alternative to task-specific CNN flood-depth estimators, addressing VLM quantitative hallucination through knowledge graph grounding. The approach is novel in combining a structural knowledge graph with a foundation VLM for physical measurement, and the authors provide a reproducible workflow (clear prompting strategy, RDF-based graph construction, public data source). The claimed 20.5% error reduction and near-real-time operation, if confirmed, would be valuable for smart city flood resilience. However, the significance is currently tempered by evaluation weaknesses: unvalidated citizen-reported ground truth, lack of uncertainty quantification, and cross-dataset comparisons that do not support the 'surpassing' claim.","major_comments":[{"comment":"The evaluation rests entirely on resident-reported depth estimates from MyCoast New York being treated as proxy ground truth, with no independent water-level measurements, sensor/gauge comparison, or screening for reporting bias. If these self-reports are rounded (e.g., to whole inches), biased by visual severity, or influenced by the same visual cues that GPT-4o uses, then the 2.11 cm MAE improvement over the baseline falls within plausible label noise. The paper must provide evidence for the reliability of these labels, such as comparison with gauge data, an analysis of report rounding, or a sensitivity study showing the conclusions are robust to label perturbations.","section":"Section 3.1"},{"comment":"Table 1 reports MAE and Pearson r values without error bars, confidence intervals, or significance tests. With n=110, the differences between FloodVision (average) at 8.17 cm and FloodVision (min) at 8.44 cm, or between FloodVision variants and the 10.28 cm baseline, may not be statistically significant. The authors should provide bootstrap confidence intervals or paired significance tests (e.g., Wilcoxon signed-rank test) to support the claim of improvement.","section":"Table 1"},{"comment":"The comparison with prior CNN-based methods (Chaudhary et al., Li et al., Alizadeh Kharazi and Behzadan, Akinboyewa et al.) is performed on different datasets with different label definitions, as the text itself concedes. The claim that FloodVision is 'surpassing prior CNN-based methods' is not supported by these incomparable numbers. A fair comparison would require re-running prior methods on the MyCoast dataset or evaluating FloodVision on the same benchmark used by those works; at minimum, the claim should be softened to 'does not require task-specific training' rather than 'surpasses'.","section":"Section 3.2"}],"minor_comments":[{"comment":"The prompt description says the model 'outputs its reasoning' but the actual output is a constrained JSON of object, height, and submergence ratio; clarify whether the JSON includes textual reasoning or only structured fields.","section":"Section 2.2"},{"comment":"The knowledge graph construction cites 'verified' dimensions from sources including Wikipedia, but the aggregation and verification procedure is not detailed; a reader cannot assess how canonical heights for objects like 'curbs' or 'trash cans' were selected from heterogeneous sources.","section":"Section 2.3"},{"comment":"The paper does not describe the image selection criteria within the MyCoast database; if the 110 images were not randomly or systematically sampled, the reported performance may not reflect typical conditions.","section":"Section 3.1"},{"comment":"The scatter plots and error distributions in Figure 3 are described qualitatively; adding numerical summaries of the residuals (e.g., median error, interquartile range) would strengthen the comparison.","section":"Section 3.2 / Figure 3"},{"comment":"There are minor typographical errors, e.g., 'anomalously depths' should be 'anomalous depths' and 'the 45°line' should be 'the 45-degree line'.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The central methodological idea is reasonable, but the evaluation's reliance on unvalidated citizen reports is a serious concern that affects the headline claims. If the authors can validate the proxy ground truth or provide a sensitivity analysis, the paper could be made publishable. The cross-dataset comparison should be reframed as contextual, not competitive."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: FloodVision is a reasonable engineering contribution—combining GPT-4o zero-shot depth estimation with a curated knowledge graph of object heights to reduce hallucination—and the writing is clear. But the central quantitative claim is not yet supported. The evaluation uses 110 citizen-submitted photos from MyCoast NY, and the \"ground truth\" depths are the residents' own estimates, with no independent water-level data. The paper itself calls them proxy ground truth in Section 3.1. If those reports are noisy or biased—rounding to whole inches, influenced by visual severity, or shaped by the same cues the model uses—the 2.11 cm MAE improvement over the GPT-4o baseline could easily be label noise. The improvement is smaller than a typical rounding increment (about 2.5 cm per inch). Also, the comparison with prior CNN methods is cross-dataset, so the word \"surpassing\" overreaches.\n\nWhat is genuinely new: the FloodKG, a structured repository of canonical dimensions for vehicles, people, and infrastructure, built from external sources, plus the override mechanism that replaces GPT-4o's provisional heights with KG values. This is cleanly designed and not circular—the heights come from outside the evaluation data. The zero-shot workflow avoids task-specific training and outputs min/max/average estimates, which is practical for emergency response.\n\nWhere it is soft:\n- No error bars, no significance tests, and only 110 images. The MAE gap (8.17 vs 10.28 cm) might be real, but the paper shows no variance.\n- The ground truth issue is load-bearing. Without any validation of citizen reports against physical measurements, the absolute MAE is uninterpretable as physical depth error.\n- The cross-dataset comparison in Section 3.2 should be clearly labeled as indicative, not a win.\n- Minor: the knowledge graph sources include a generic Wikipedia reference, so \"verified\" dimensions are only as solid as that.\n\nThe limitations section is candid about object visibility but silent on the ground truth problem.\n\nThe framework deserves a serious referee. The idea is timely and the engineering is honest. The authors need to either validate their proxy ground truth against some physical measurements or reframe the claims as agreement with citizen estimates. I would send it to review, expecting major revision rather than desk rejection. I would not cite it for the numbers, but the knowledge-graph grounding idea is worth watching.","headline":"A sensible zero-shot VLM-plus-knowledge-graph flood depth estimator whose headline 8.17 cm MAE is only as trustworthy as the unvalidated citizen-reported depths used as ground truth.","tokens_in":6420,"tokens_out":1959,"would_cite":false,"duration_ms":17531,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A zero-shot framework that fuses a vision-language model with a domain knowledge graph estimates urban flood depth to a mean error of 8.17 cm on crowdsourced images.","keywords":["urban flood depth estimation","vision-language models","knowledge graph","zero-shot learning","emergency response","flood resilience","crowdsourced imagery","submergence ratio"],"falsifier":"Collect flood photos with independent water-depth measurements from the same streets and times—staff gauges, surveyed high-water marks, or pressure-sensor logs—run FloodVision on those photos, and compare its outputs to the independent values. If the error against independent truth is substantially larger than 8.17 cm, the claimed accuracy depends on the proxy labels.","tokens_in":5538,"feed_emoji":"🌊","tokens_out":7124,"duration_ms":56916,"temperature":0.7,"pith_summary":"FloodVision sets out to show that a foundation vision-language model—a model that reads images and answers in text—can estimate how deep floodwater covers a road, if its numeric guesses are pinned to measured dimensions of everyday objects. The method asks the model to name visible reference objects like tires, curbs, and people's knees, replaces the model's guessed heights with canonical heights from a curated knowledge graph, and multiplies each height by the object's estimated submerged fraction. On 110 crowdsourced flood photos, the averaged estimate lands within 8.17 cm of the reported depth, a 20.5% improvement over the model working alone (10.28 cm) and better than earlier CNN-based approaches. The appeal is practical: depth estimates with no training data and near real-time speed could support road-closure decisions, emergency routing, and flood mapping in cities that lack dense sensor networks.","feed_headline":"FloodVision cuts flood-depth error to 8.17 cm","feed_subtitle":"A vision-language model plus physical object heights beats image-only baselines on crowdsourced flood photos.","key_machinery":"The load-bearing mechanism is FloodKG, a curated knowledge graph of canonical real-world dimensions for urban objects—vehicles, people, and infrastructure—stored as mean heights with standard deviations and linked by subClassOf and partOf relations. The pipeline uses the graph as an override: GPT-4o names candidate reference objects, matching graph entries replace the model's provisional height, and only the submergence ratio comes from the model's visual estimate. Per-object depth is height times submergence ratio, followed by statistical outlier filtering and aggregation into minimum, average, and maximum estimates. This retrieval-and-replace step is what turns qualitative scene understanding into physically anchored numbers, and its absence in the GPT-4o-only baseline is what the 20.5% error reduction measures.","core_discovery":"The central claim is that physically grounding a vision-language model's quantitative output removes enough hallucination to make zero-shot flood depth estimation viable. FloodVision prompts GPT-4o to identify up to three visually distinct reference objects, canonicalizes their names, retrieves verified height statistics from the FloodKG knowledge graph, estimates each object's submergence ratio (the fraction of its height under water), then filters outliers and aggregates into minimum, average, and maximum depth. The average variant achieves a mean absolute error of 8.17 cm and Pearson r of 0.51 on the 110-image test set, versus 10.28 cm and 0.44 for the GPT-4o-only baseline; the minimum and maximum variants also beat the baseline. The authors present this as preliminary evidence that semantic scene understanding plus external physical knowledge generalizes across diverse urban scenes where fixed object detectors and task-specific training fall short.","pith_inferences":["The same retrieve-and-override pattern could transfer to other VLM measurement tasks—snow depth against parked cars, debris height against curbs, or wrack-line surveys—wherever a stable catalog of reference sizes can be built.","Because unmatched objects fall back to the model's own estimate, graph coverage is a likely driver of accuracy; a testable prediction is that MAE shrinks monotonically as FloodKG adds reference categories.","A natural ablative test the paper does not run is swapping GPT-4o for an open-weight vision-language model; if accuracy holds, the knowledge graph rather than the specific model is carrying the improvement.","The paper leaves open how to handle images with no recognizable reference objects; future work could combine water-surface texture or reflection cues with the graph to cover those cases."],"forward_implications":["Because FloodVision needs no task-specific training, the same pipeline can be pointed at a new city or event immediately, as long as visible reference objects exist in the photos.","The 20.5% drop over the VLM-only baseline indicates that external physical knowledge is a practical cure for quantitative hallucination in vision-language models, not just for flood depth.","The minimum/average/maximum aggregation turns one image into a depth range, which supports uncertainty-aware road-closure and routing decisions during active flooding.","Near-real-time operation on geotagged citizen photos or CCTV frames makes the method a plausible perception module for urban flood digital twins and alert systems."],"supporting_citations":[{"why":"Supplies the GPT-4o vision-language model whose semantic reasoning the framework builds on and whose output serves as the baseline.","marker":"[12]"},{"why":"Supplies the 110 crowdsourced flood images and the resident-reported depth values used as proxy ground truth for evaluation.","marker":"[21]"},{"why":"Vehicle specification data used to populate dimensions for cars, SUVs, and buses in the knowledge graph.","marker":"[19]"},{"why":"National anthropometric survey data used to set human body landmark heights in the knowledge graph.","marker":"[5]"},{"why":"Reference dimensions for infrastructure objects such as curbs, hydrants, and trash cans in the knowledge graph.","marker":"[20]"},{"why":"Earlier CNN-based flood water level estimation from social media images that FloodVision is compared against.","marker":"[4]"},{"why":"Earlier CNN-based method estimating water depth from social media images that FloodVision outperforms.","marker":"[7]"},{"why":"Prior GPT-4-based flood depth estimation without knowledge graph grounding, the closest VLM predecessor.","marker":"[1]"}],"fun_headline_variants":["FloodVision: AI that sees flood depth within 8 cm","GPT-4o + knowledge graph: flood depth error down to 8 cm","Zero-shot flood depth from photos, grounded in real object sizes","FloodVision: physical grounding makes flood-depth AI reliable","From photos to flood depth: 8 cm MAE with zero-shot language model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the depth values residents typed into the crowdsourced flood-report platform are accurate enough to serve as ground truth; if those self-reports are biased or noisy, the reported 8.17 cm error does not reflect true estimation accuracy.","fun_headline_variants_meta":{"raw":{"variants":["FloodVision: AI that sees flood depth within 8 cm","GPT-4o + knowledge graph: flood depth error down to 8 cm","Zero-shot flood depth from photos, grounded in real object sizes","FloodVision: physical grounding makes flood-depth AI reliable","From photos to flood depth: 8 cm MAE with zero-shot language model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00072,"raw_usage":{"total_tokens":3235,"prompt_tokens":953,"completion_tokens":2282,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":569,"completion_tokens_details":{"reasoning_tokens":2188}},"tokens_in":569,"tokens_out":2282,"duration_ms":13317,"temperature":1.0,"reasoning_tokens":2188,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:26:42.799374+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect flood photos with independent water-depth measurements from the same streets and times—staff gauges, surveyed high-water marks, or pressure-sensor logs—run FloodVision on those photos, and compare its outputs to the independent values. If the error against independent truth is substantially larger than 8.17 cm, the claimed accuracy depends on the proxy labels.","supporting_citations":[{"cited_title":"Retrieved July 3, 2025 from https://mycoast.org/ny/flood-watch","cited_arxiv_id":null,"evidence_quote":"Supplies the 110 crowdsourced flood images and the resident-reported depth values used as proxy ground truth for evaluation."},{"cited_title":"AutomobileDimension.com","cited_arxiv_id":null,"evidence_quote":"Vehicle specification data used to populate dimensions for cars, SUVs, and buses in the knowledge graph."},{"cited_title":"Fryar, Margaret D","cited_arxiv_id":null,"evidence_quote":"National anthropometric survey data used to set human body landmark heights in the knowledge graph."},{"cited_title":"Retrieved July 3, 2025 from https://www.wikipedia.org/","cited_arxiv_id":null,"evidence_quote":"Reference dimensions for infrastructure objects such as curbs, hydrants, and trash cans in the knowledge graph."},{"cited_title":"Chaudhary, S","cited_arxiv_id":null,"evidence_quote":"Earlier CNN-based flood water level estimation from social media images that FloodVision is compared against."},{"cited_title":"Naser Lessani, and Zhenlong Li","cited_arxiv_id":null,"evidence_quote":"Prior GPT-4-based flood depth estimation without knowledge graph grounding, the closest VLM predecessor."}],"review_version":2}