{"id":"625e6957-bd6d-44a9-b00d-cfb96b7645e2","arxiv_id":"2512.24470","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Semantic Lookout applies constrained vision-language models to select cautious fallback maneuvers in maritime scenes, outperforming geometry baselines on hazard relief and aligning with human judgments in 40 harbor tests plus a field run.","lead":"The paper introduces Semantic Lookout, a camera-only system that uses vision-language models to detect semantic hazards such as fires or diver flags and select short-horizon, human-overridable fallback maneuvers for autonomous boats. This targets regulatory needs for handling situations where meaning matters more than geometry alone.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Evaluation limited to 40 harbor scenes plus one field run leaves open whether VLM semantic understanding generalizes without missing hazards or introducing unsafe maneuvers.","rationale":"The reader's weakest assumption directly identifies the alignment-and-safety gap; the concrete limitation is the narrow evaluation scale and missing robustness checks, which prevents moving from UNVERDICTED to a stronger verdict without additional data.","tokens_in":1835,"tokens_out":306,"duration_ms":20713,"concrete_test":"Re-run the 40-scene protocol plus the fire-hazard standoff metric on an expanded set of 150 scenes that include night, fog, and crossing-traffic cases; compute false-negative hazard detection rate against expert annotations. If the rate exceeds 8% or standoff improvement disappears, the generalization claim for real maritime conditions is unsupported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that VLM outputs on camera images produce water-valid trajectories that align with human consensus, relieve risk, and stay safe under the IMO MASS Code constraints. The reported evidence consists of per-call metrics on 40 scenes, majority-of-three human alignment, standoff improvement on fire scenes, and a single on-water handover demonstration. No quantitative false-negative rates for critical hazards, no ablation on weather/lighting/traffic variability, and no comparison against expert mariner ground truth beyond the small consensus panel are described. This leaves the safety and compatibility argument dependent on an untested extrapolation from the narrow test distribution.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces Semantic Lookout, a camera-only VLM-based fallback maneuver selector for maritime autonomy that chooses cautious, water-valid actions (or station-keeping) from world-anchored trajectories under continuous human authority. It argues that VLMs supply the semantic awareness needed for out-of-distribution hazards (e.g., diver-down flags, nearby fires) that classical geometry-based stacks miss, and that a fast-slow anomaly pipeline makes this practical within the IMO MASS Code alert-to-takeover window. The central empirical support consists of per-call metrics on 40 harbor scenes (scene understanding, latency, majority-of-three human alignment, standoff improvement on fire scenes) plus a single on-water handover demonstration; sub-10 s models are claimed to retain most awareness while outperforming geometry baselines.","tokens_in":1973,"tokens_out":564,"duration_ms":35962,"significance":"If the reported alignment and risk-relief results prove robust, the work supplies a concrete, regulator-compatible mechanism for injecting foundation-model semantics into short-horizon maritime fallback systems. It explicitly credits the field demonstration and the latency-aware comparison of fast versus slow VLMs, and it motivates hybrid architectures that pair VLM semantics with multi-sensor BEV perception.","major_comments":[{"comment":"Evaluation on 40 harbor scenes: the abstract and results summary report quantitative improvements in human alignment and fire-hazard standoff distance, yet provide no exact metric definitions, statistical tests, error bars, or data-exclusion criteria. Because these numbers are the primary evidence for the claim that VLMs produce water-valid trajectories compatible with the draft IMO MASS Code, the absence of these details is load-bearing for the central safety argument.","section":"Results on 40 harbor scenes"},{"comment":"Field-run demonstration: the single on-water alert-to-fallback-to-handover trial is presented as verification of end-to-end operation, but no quantitative false-negative rates for critical hazards, no ablation across weather/lighting/traffic conditions, and no comparison against expert mariner ground truth beyond the three-person consensus panel are reported. This leaves the generalization claim dependent on an untested extrapolation from the narrow test distribution.","section":"Field demonstration"}],"minor_comments":[{"comment":"The term 'Semantic Lookout' is introduced without a concise formal definition or pseudocode; a short algorithmic box would clarify the candidate-constrained selection step.","section":"Method"},{"comment":"Figure captions for the harbor scenes and trajectory overlays should explicitly state the VLM prompt template and the exact human-voting protocol used for the majority baseline.","section":"Figures"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive and detailed feedback. We address each major comment point by point below, indicating the revisions we will make to strengthen the manuscript's clarity and empirical rigor while preserving the original contributions.","responses":[{"response":"We agree that the current presentation lacks sufficient detail on the evaluation protocol. In the revised manuscript we will expand the Results and Methods sections to provide: exact definitions of all metrics (scene-understanding accuracy, majority-of-three human alignment procedure with inter-rater agreement, standoff-distance computation in meters, and latency); statistical tests (e.g., paired comparisons with p-values against geometry baselines); error bars or confidence intervals on all reported figures and tables; and explicit data-exclusion criteria (e.g., scenes discarded for camera failure, extreme glare, or insufficient visibility). These additions will be placed before the quantitative claims to make the safety argument fully transparent.","revision_made":"yes","referee_comment":"Evaluation on 40 harbor scenes: the abstract and results summary report quantitative improvements in human alignment and fire-hazard standoff distance, yet provide no exact metric definitions, statistical tests, error bars, or data-exclusion criteria. Because these numbers are the primary evidence for the claim that VLMs produce water-valid trajectories compatible with the draft IMO MASS Code, the absence of these details is load-bearing for the central safety argument."},{"response":"We acknowledge the limitations of a single field trial. The demonstration was intended only to confirm that the alert-to-fallback-to-handover sequence can execute within the IMO MASS Code time window under real conditions; it was never presented as a statistical evaluation. In revision we will (i) explicitly state the scope and constraints of the field run, (ii) add a dedicated limitations paragraph discussing the absence of false-negative rates and multi-condition ablations due to regulatory and safety requirements for on-water testing, and (iii) clarify that the three-person consensus serves as an initial human-alignment benchmark while noting the need for broader expert validation in future work. The primary quantitative evidence remains the 40 harbor scenes.","revision_made":"partial","referee_comment":"Field-run demonstration: the single on-water alert-to-fallback-to-handover trial is presented as verification of end-to-end operation, but no quantitative false-negative rates for critical hazards, no ablation across weather/lighting/traffic conditions, and no comparison against expert mariner ground truth beyond the three-person consensus panel are reported. This leaves the generalization claim dependent on an untested extrapolation from the narrow test distribution."}],"tokens_in":1578,"tokens_out":544,"duration_ms":48816,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that this work puts vision-language models into a constrained role as a semantic selector for short-horizon safety maneuvers when a maritime system encounters meaning-dependent situations like hazards or flags. They build Semantic Lookout around candidate water-valid trajectories, keep continuous human authority, and tie it to the draft IMO MASS Code needs for fallback and takeover windows. Sub-10-second models keep most of the scene awareness while the selector beats geometry-only baselines on alignment and standoff distance in the reported tests. The field run shows the full alert-to-maneuver-to-handover loop can run on water. That combination of design choices and initial numbers is the concrete contribution here. The evaluation covers 40 harbor scenes for latency, human consensus match via majority vote, and risk relief on fire cases, plus the one real-world check. This gives a usable template for anyone trying to add semantic awareness without replacing the rest of the stack. The soft spots sit in the test scope and reporting. Forty scenes plus one field run leave open how the VLM handles weather shifts, lighting changes, or denser traffic, and there are no error bars, false-negative rates for missed hazards, or ablations on those variables. The human panel is small, so it is not yet clear how well it tracks expert mariner judgment beyond the test set. The abstract summarizes outcomes without the full statistical detail that would let a reader judge robustness directly. This paper is for robotics and autonomy groups working on regulated maritime systems or hybrid foundation-model fallbacks. Readers who need a worked example of candidate-constrained VLM use in a safety loop will get value from the pipeline and the regulatory framing. It deserves a serious referee because the applied setup addresses a real gap with measurable early results, even if the current evidence is preliminary and would benefit from broader testing requests during review. I would send it out with that in mind.","headline":"This paper shows a practical way to use VLMs for picking human-overridable fallback maneuvers in maritime autonomy on harbor scenes, with some gains over geometry baselines, but the evidence stays narrow.","tokens_in":2513,"tokens_out":458,"would_cite":false,"duration_ms":32445,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/RealityFromDistinction.lean","rs_theorem":"reality_from_one_distinction","paper_passage":"We introduce Semantic Lookout, a camera-only, candidate-constrained VLM fallback maneuver selector that selects one cautious action (or station-keeping) from water-valid, world-anchored trajectories under continuous human authority."},{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"On 40 harbor scenes we measure per-call scene understanding and latency, alignment with human consensus (model majority-of-three voting), short-horizon risk-relief on fire hazard scenes"}],"headline":"VLM-based maritime fallback selector for IMO MASS Code compliance; no structural overlap with RS forcing chain","alignment":"orthogonal","rationale":"The paper's central machinery is a candidate-constrained VLM selector (FB-1/FB-n majority voting over water-gated straight-line primitives) that produces short-horizon, human-overridable actions for semantic OOD hazards (diver flags, fire, MOB). This is a practical robotics application of foundation-model reasoning under regulatory constraints. RS derives spacetime, c=1, ℏ, G, φ, J-cost, and 8-tick periodicity from a single distinction via machine-checked theorems (reality_from_one_distinction, Jcost functional uniqueness, AlexanderDuality D=3 forcing). The paper contains none of these elements, no cost-function reasoning, no ratio symmetry, no parameter-free constant derivations, and operates in a domain (camera-only maritime anomaly handling) on which RS is silent.","tokens_in":59103,"confidence":"high","tokens_out":396,"duration_ms":12355,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Vision-language models select safe fallback maneuvers for autonomous ships by interpreting semantic hazards.","keywords":["vision-language models","maritime autonomy","semantic hazard detection","fallback maneuvers","IMO MASS Code","autonomous vessels","safety systems","hazard response"],"falsifier":"A test case in which the VLM fails to identify a hazard such as a diver in the water or chooses a trajectory that reduces rather than increases safety distance to a fire.","tokens_in":2730,"feed_emoji":"🚢","tokens_out":656,"duration_ms":70786,"temperature":0.7,"pith_summary":"The paper establishes that vision-language models can supply the semantic understanding needed for maritime autonomous vessels to handle unexpected hazards where meaning matters, such as recognizing a diver-down flag or a nearby fire. It introduces Semantic Lookout as a practical system that uses camera images to choose cautious short-horizon actions from possible water-based trajectories while keeping human oversight. Evaluations across 40 harbor scenes confirm alignment with human consensus, better performance than geometry-based methods, risk reduction on fire scenes, and operation within latency limits suitable for regulatory handover periods. A field demonstration shows the full alert to maneuver to operator takeover sequence works in practice.","feed_headline":"VLMs choose safe fallback maneuvers for autonomous ships","feed_subtitle":"Camera-based semantic awareness supports quick human-overridable actions to comply with maritime autonomy regulations","key_machinery":"Semantic Lookout, the candidate-constrained vision-language model fallback maneuver selector operating on camera images to pick cautious water-valid actions under human authority.","core_discovery":"Semantic Lookout is a camera-only VLM system that selects one cautious fallback maneuver or station-keeping from water-valid trajectories. It provides semantic awareness for out-of-distribution situations, shows alignment with human majority voting on scene understanding, outperforms geometry-only baselines by increasing standoff on fire hazards, and runs in sub-10 seconds to fit the alert-to-takeover window of the draft IMO MASS Code. End-to-end functionality is confirmed in a field run.","pith_inferences":["This semantic approach could integrate with existing multi-sensor perception systems to handle both meaning and precise geometry.","Adapting the models to maritime-specific training data might improve accuracy for local conditions and hazards.","The method opens possibilities for applying similar VLM-based fallbacks in other autonomous domains with regulatory handover needs.","Future work could test if these selections remain safe over longer horizons or in more varied sea states."],"forward_implications":["Sub-10s VLM models preserve most scene awareness of slower models.","The system increases standoff distance to fire hazards compared to geometry-only baselines.","Full pipeline from alert through fallback maneuver to operator handover is verified on water.","VLMs act as semantic fallback selectors that fit the draft IMO MASS Code requirements within practical latency limits."],"fun_headline_variants":["VLMs select semantic fallbacks for ships","Camera VLMs detect maritime semantic hazards","Semantic Lookout picks cautious autonomy actions","VLMs support semantic ship safety maneuvers"],"cache_read_input_tokens":64,"weakest_assumption_plain":"VLM scene understanding from camera images will match human agreement and produce safe choices when limited to water-valid paths, without overlooking important hazards or adding risks under real maritime conditions.","fun_headline_variants_meta":{"raw":{"variants":["VLMs select semantic fallbacks for ships","Camera VLMs detect maritime semantic hazards","Semantic Lookout picks cautious autonomy actions","VLMs support semantic ship safety maneuvers"]},"model":"grok-4.3","cost_usd":0.007152,"raw_usage":{"total_tokens":3281,"prompt_tokens":787,"num_sources_used":0,"completion_tokens":50,"cost_in_usd_ticks":71515500,"prompt_tokens_details":{"text_tokens":787,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2444,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":787,"tokens_out":50,"duration_ms":91859,"temperature":1.0,"reasoning_tokens":2444,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-21T15:34:43.918927+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A test case in which the VLM fails to identify a hazard such as a diver in the water or chooses a trajectory that reduces rather than increases safety distance to a fire.","supporting_citations":[],"review_version":1}