{"id":"a038a7f0-65e7-4bae-8078-b6bf2a26fedd","arxiv_id":"2605.25784","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"VertiCue-Bench shows MLLMs can read raw CHM height cues but largely fail to integrate them into reliable semantic reasoning, often underperforming RGB-only baselines.","lead":"The paper introduces VertiCue-Bench, a diagnostic benchmark of 1,534 instances across 17 tasks that tests whether multimodal LLMs can use canopy height models to resolve 2D visual ambiguities in remote sensing natural scenes. A smart generalist might read it to learn that current models show a clear gap between perceiving height data and actually using it for semantic reasoning.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Benchmark curation may not isolate height perception from semantic reasoning without explicit validation of unique disambiguation cases.","rationale":"The reader's weakest_assumption directly identifies the curation step as the load-bearing point, and the abstract provides no quantitative evidence (e.g., inter-annotator checks or cue-uniqueness statistics) that would refute it. Because the full text is referenced but not reproduced here, no stronger internal inconsistency can be diagnosed; the concern therefore aligns exactly with the reader's and leaves the UNVERDICTED status unchanged.","tokens_in":1808,"tokens_out":347,"duration_ms":19302,"concrete_test":"For a random sample of 100 instances, independently verify whether removing the CHM channel (while keeping RGB) changes the ground-truth label or makes the question unsolvable; if >15% of cases remain solvable via RGB alone or other cues, recompute the joint-input accuracy gap and check whether the dissociation statistic holds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of a perception-reasoning dissociation depends on the 1,534 instances across 17 tasks being constructed such that low-level CHM reading is separable from ambiguity-aware semantic use, and that joint RGB+CHM inputs require the height cue for correct reasoning. The abstract states the instances were 'carefully curated to explicitly disentangle' these, but if curation selected scenes where RGB texture already correlates with height or where alternative non-height cues suffice, then underperformance on joint inputs could reflect fusion or prompting artifacts rather than a true geometry-to-semantics gap. Counterfactual modality testing would only support the dissociation if the tasks were pre-validated to have no such confounds.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces VertiCue-Bench, a diagnostic benchmark with 1,534 instances across 17 tasks that explicitly disentangles low-level height perception (via Canopy Height Models) from ambiguity-aware semantic reasoning in remote sensing natural scenes. Evaluations of 14 general and remote-sensing-specialized MLLMs, combined with counterfactual modality testing, claim to reveal a perception-reasoning dissociation: models show emerging competence in reading raw CHM cues but largely fail to translate this into reliable semantic reasoning and often underperform RGB-only baselines when joint constraints are required.","tokens_in":1955,"tokens_out":494,"duration_ms":21342,"significance":"If the dissociation holds after proper validation of the benchmark tasks, the work would be significant for exposing a geometry-to-semantics gap in current MLLMs for geospatial reasoning. It provides a new diagnostic benchmark that could guide improvements in integrating 3D structural data with semantic understanding, particularly in natural scenes with spectral confusion.","major_comments":[{"comment":"The central claim of a perception-reasoning dissociation rests on the 1,534 instances across 17 tasks being constructed such that low-level CHM reading is separable from ambiguity-aware semantic use, and that joint RGB+CHM inputs require the height cue for correct reasoning. The abstract states the instances were 'carefully curated to explicitly disentangle' these capabilities, but provides no details on validation of the disentanglement, confirmation of unique disambiguation cases, or checks against confounds such as RGB texture already correlating with height or alternative non-height cues sufficing. Without such validation, underperformance on joint inputs could reflect fusion or prompting artifacts rather than a true gap (see also the weakest assumption in the provided reader's take).","section":"Abstract / VertiCue-Bench construction"}],"minor_comments":[{"comment":"The abstract reports evaluations on 14 models but does not mention error bars, statistical significance testing, or variance across runs, which are needed to assess whether the reported dissociation is robust.","section":"Abstract"},{"comment":"Details on task construction, data selection criteria, and how the 17 tasks were designed to isolate height perception from semantic reasoning are absent from the abstract and would be required for reproducibility.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the thoughtful review and for identifying the need for greater transparency on benchmark validation. We address the major comment below and will incorporate clarifications in the revision.","responses":[{"response":"We agree that the abstract is necessarily concise and does not itself contain the validation details. Section 3.2 and Appendix A of the manuscript describe the curation pipeline: candidate scenes were first filtered for spectral confusion where RGB-only performance is near chance; expert annotators then verified that CHM supplies the sole disambiguating cue (with inter-annotator agreement reported); and each task includes explicit RGB-only, CHM-only, and joint conditions to isolate the contribution of height. We acknowledge, however, that quantitative checks for residual RGB-height correlations and exhaustive tests ruling out alternative non-height cues (e.g., shadow or texture gradients) are only summarized rather than tabulated. We will add a dedicated subsection and supplementary table reporting these confound analyses in the revised manuscript.","revision_made":"partial","referee_comment":"[Abstract / VertiCue-Bench construction] The central claim of a perception-reasoning dissociation rests on the 1,534 instances across 17 tasks being constructed such that low-level CHM reading is separable from ambiguity-aware semantic use, and that joint RGB+CHM inputs require the height cue for correct reasoning. The abstract states the instances were 'carefully curated to explicitly disentangle' these capabilities, but provides no details on validation of the disentanglement, confirmation of unique disambiguation cases, or checks against confounds such as RGB texture already correlating with height or alternative non-height cues sufficing. Without such validation, underperformance on joint inputs could reflect fusion or prompting artifacts rather than a true gap (see also the weakest assumption in the provided reader's take)."}],"tokens_in":1436,"tokens_out":388,"duration_ms":28199,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that VertiCue-Bench tries to isolate whether models can read canopy height from CHM inputs and then use that geometry to disambiguate scenes where RGB textures overlap. They run 14 models across 17 tasks and 1,534 instances with modality swaps, and the pattern they describe is that raw height reading shows up but joint RGB+CHM reasoning often drops below RGB-only performance.\n\nWhat is new is the explicit split between low-level height perception tasks and the ambiguity-aware semantic tasks, plus the counterfactual tests. That framing is not in the earlier remote-sensing MLLM benchmarks referenced in the abstract, and it targets a real practical issue in natural scenes where vertical structure should matter.\n\nThe paper does a reasonable job naming the geometry-to-semantics gap and showing that current models are not closing it. The idea of holding out the height cue and measuring whether it helps is straightforward and worth testing.\n\nThe soft spot is that the abstract gives almost no information on how the 1,534 instances were selected or validated to guarantee that height is the only reliable disambiguator. If RGB texture already correlates with height in the chosen scenes, or if other cues suffice, then the reported underperformance on joint inputs could come from fusion issues rather than a clean failure to use geometry. The stress-test concern about curation therefore lands until the full paper shows explicit checks for unique disambiguation cases. No error bars or significance numbers appear in the abstract either, so the size of the dissociation is hard to judge.\n\nThis is for people working on multimodal models for geospatial or remote-sensing applications. A reader who needs a diagnostic for height cue usage would find the task split useful even if the numbers need more scrutiny.\n\nIt should go to peer review. The benchmark direction is worth referee time, but the construction details and statistical support will need to be checked in the full manuscript.","headline":"The paper builds a benchmark to check if MLLMs actually apply CHM height data to resolve spectral ambiguities in remote sensing scenes, and reports that perception works better than the downstream reasoning step.","tokens_in":2444,"tokens_out":469,"would_cite":false,"duration_ms":19893,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Current MLLMs read canopy height data but rarely apply it to resolve spectral ambiguities in natural remote sensing scenes.","keywords":["multimodal large language models","remote sensing","canopy height models","geospatial reasoning","benchmark","3D perception","semantic disambiguation","height cues"],"falsifier":"Run the same 14 models on a fresh collection of scenes where height differences are the only reliable cue for correct labels; if accuracy stays near or below the RGB-only baseline, the dissociation claim holds.","tokens_in":2712,"feed_emoji":"","tokens_out":710,"duration_ms":15885,"temperature":0.7,"pith_summary":"The paper presents VertiCue-Bench, a diagnostic set of 1,534 instances across 17 tasks, to test whether multimodal large language models can use explicit vertical structure from canopy height models to disambiguate remote sensing scenes where different land covers share similar textures. Evaluations of 14 general and remote-sensing MLLMs show they achieve some competence at extracting height information yet largely fail to integrate it with image data for correct semantic judgments, and sometimes perform worse than RGB-only versions. This dissociation indicates that models treat height as an isolated signal rather than a constraint that narrows possible interpretations of the scene. A sympathetic reader would care because many geospatial tasks require exactly this step from geometry to meaning, such as distinguishing forest types or urban features that look alike from above.","feed_headline":"Benchmark finds MLLMs read height maps but ignore them for scene labels","feed_subtitle":"1,534 cases across 17 tasks show models often perform worse with both RGB and CHM than with RGB alone on natural scenes.","key_machinery":"VertiCue-Bench, a benchmark of 1,534 curated instances across 17 tasks that separate low-level height perception from ambiguity-aware semantic reasoning.","core_discovery":"VertiCue-Bench reveals a perception-reasoning dissociation: models exhibit emerging competence in reading raw CHM height cues, yet largely fail to translate geometric perception into reliable semantic reasoning and often underperform RGB-only baselines when joint constraints are required.","pith_inferences":["Future benchmarks could add controlled noise to the height maps to measure how sensitive reasoning remains when the geometric cue is imperfect.","The observed dissociation suggests that training objectives focused only on captioning or question answering may not encourage the model to treat height as evidence that rules out certain interpretations.","Architectures that maintain separate streams for appearance and structure until a late fusion stage might close the gap observed here."],"forward_implications":["Models that pass isolated height-reading tests will still produce unreliable answers on tasks requiring joint image-and-height constraints.","Adding canopy height data can degrade performance relative to image-only input when the model lacks a mechanism to treat height as a disambiguating constraint.","Progress on geospatial MLLMs will require explicit training or architectures that enforce consistency between geometric measurements and semantic categories.","Existing remote-sensing benchmarks that omit vertical structure will continue to overestimate model capability on natural scenes."],"fun_headline_variants":["MLLMs read CHM height but fail to apply it in reasoning","Height perception exists yet semantic use lags in MLLMs","VertiCue shows models ignore CHM when labeling scenes","Benchmark reveals perception-reasoning dissociation in MLLMs","Tests find MLLMs underperform with RGB plus CHM input"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The 1,534 instances and 17 tasks were designed and selected so that success or failure can be attributed specifically to the ability to combine height with appearance rather than to other factors in data curation or prompting.","fun_headline_variants_meta":{"raw":{"variants":["MLLMs read CHM height but fail to apply it in reasoning","Height perception exists yet semantic use lags in MLLMs","VertiCue shows models ignore CHM when labeling scenes","Benchmark reveals perception-reasoning dissociation in MLLMs","Tests find MLLMs underperform with RGB plus CHM input"]},"model":"grok-4.3","cost_usd":0.003897,"raw_usage":{"total_tokens":1933,"prompt_tokens":695,"num_sources_used":0,"completion_tokens":84,"cost_in_usd_ticks":38965500,"prompt_tokens_details":{"text_tokens":695,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1154,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":695,"tokens_out":84,"duration_ms":13191,"temperature":1.0,"reasoning_tokens":1154,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T22:29:23.252454+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Run the same 14 models on a fresh collection of scenes where height differences are the only reliable cue for correct labels; if accuracy stays near or below the RGB-only baseline, the dissociation claim holds.","supporting_citations":[],"review_version":1}