{"id":"8a4264ee-44df-4c14-b932-6cda7652679e","arxiv_id":"2505.07084","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Fine-tuning multimodal language models on a new SOTIF-focused driving dataset improves question answering and captioning, but the open-ended gains are measured by an LLM judge with no independent human scoring.","lead":"The paper introduces DriveSOTIF, a dataset of driving images with captions and safety question-answer pairs, and fine-tunes multimodal language models on it to identify perception-related driving risks. The authors report better question-answering and captioning after tuning, but their open-ended scores are awarded by the same kind of AI that wrote the dataset.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table VII contradicts 'fine-tuning improves across all evaluation metrics': several models score worse on relevance and overall after fine-tuning, and the abstract's 11.8%/12.0% gains are selected best rows, not aggregate results.","rationale":"The paper is a useful engineering contribution: it releases a SOTIF-oriented VQA/captioning dataset built from PeSOTIF images, fine-tunes eight MLLMs with LoRA, and reports deployment latency. Close-ended accuracy gains are directionally consistent across all eight rows, which suggests the dataset has some measurable effect. The central problem is the mismatch between the text's universal claim and the reported data. Sec. IV.C says 'fine-tuning improves performance across all evaluation metrics,' but Table VII shows substantial declines in relevance, trustworthiness, and overall scores for three of the eight models. The abstract selects the best single-model gains rather than an aggregate, so the headline numbers overstate the evidence. This is not a disagreement with field consensus; it is an internal inconsistency. The reader's weakest assumption about LLM-generated ground truth is also valid, but it is an external-validity concern; the internal contradiction is more decisive because even granting the GPT-4.1 labels and judge scores, the table does not support the headline. The paper's own acknowledgements that the model missed the hidden child in the box scenario and that inference-time hallucination persists further support a cautious framing. I therefore retain the reader's conditional verdict, with the added condition that aggregate statistics replace cherry-picked maxima and that 'across all evaluation metrics' be removed unless verified.","tokens_in":19018,"tokens_out":10811,"duration_ms":103981,"concrete_test":"Recompute from Table VII all paired FT-minus-Base differences per model and per metric, then run a two-sided sign test or paired bootstrap across the eight model rows on the direction of changes for accuracy, relevance, trustworthiness, clarity, coherence, and overall score. Also recompute the abstract's headline improvements as the mean and median relative gain over all rows rather than the maximum row. If multiple models show decreased relevance or overall scores, or if the aggregate open-ended gain is not clearly positive, replace 'improves across all evaluation metrics' and the 11.8%/12.0% claims with honest aggregate statistics.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Sec. IV.C claims that fine-tuning improves performance across all evaluation metrics, but Table VII reports multiple counterexamples: Qwen2.5-VL-7B relevance drops from 4.57 to 4.13 and overall from 4.76 to 4.61; InternVL3-2B relevance drops from 4.22 to 3.92 and overall from 4.48 to 4.47; InternVL3-8B relevance drops from 4.56 to 4.12, trustworthiness from 4.70 to 4.53, and overall from 4.79 to 4.62. The abstract's headline numbers are not aggregate improvements: 11.8% close-ended is the relative gain of InternVL3-2B specifically (58.01 to 64.84), and 12.0% open-ended is InternVL3-1B's overall score change (3.84 to 4.30). Across the eight evaluated models, fine-tuning effects on open-ended and reasoning metrics are mixed, with several larger models flat or worse. The central claim that fine-tuning on DriveSOTIF broadly improves perception-related SOTIF reasoning therefore fails on the paper's own reported measurements, independent of the separate concern that GPT-4.1 serves as both annotation validator and open-ended judge. Close-ended accuracy does improve in all eight rows, so the dataset is not inert; but the universal and open-ended claims need to be corrected or substantiated with aggregate statistics.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DriveSOTIF, a visual question answering (VQA) and image captioning dataset for perception-related Safety of the Intended Functionality (SOTIF) in autonomous driving. The dataset is generated by a multi-agent pipeline using GPT-4v, GPT-4o, and Claude 3 Opus, with GPT-4.1-based validation, and contains 1,114 images, 1,114 captions, and 5,570 question-answer pairs. The authors benchmark several open-source multimodal LLMs before and after LoRA-based supervised fine-tuning, reporting gains in close-ended VQA accuracy and in GPT-4.1-judged open-ended VQA scores, along with deployment latency measurements on RTX 3090 and Jetson Orin platforms and qualitative real-world case studies from Canada and China. The paper claims to be the first application of domain-specific MLLM fine-tuning to the SOTIF domain.","tokens_in":19320,"tokens_out":4363,"duration_ms":40395,"significance":"If the reported results are taken at face value, the paper offers a useful new resource: DriveSOTIF is, to my knowledge, the first VQA and captioning dataset focused on perception-related SOTIF, and the authors make the dataset and code publicly available. The multi-agent generation pipeline with validation and the careful deployment study are pragmatic contributions, and the close-ended accuracy improvements are consistent across all eight evaluated models. The paper also contains an honest discussion of hallucination and of the hidden-child failure case. However, the central open-ended evaluation is circular because the annotation, validation, and judging all come from the same LLM family, and the paper's universal-improvement claim is contradicted by its own Table VII. These issues affect the main quantitative claims and need to be addressed before the paper can be accepted.","major_comments":[{"comment":"The claim that \"Fine-tuning improves performance across all evaluation metrics\" is directly contradicted by Table VII. For example, Qwen2.5-VL-7B relevance drops from 4.57 to 4.13 and overall from 4.76 to 4.61; InternVL3-2B relevance drops from 4.22 to 3.92 and overall from 4.48 to 4.47; InternVL3-8B relevance drops from 4.56 to 4.12, trustworthiness from 4.70 to 4.53, and overall from 4.79 to 4.62. The abstract's headline 11.8% and 12.0% figures are also selected best-model gains (InternVL3-2B for close-ended accuracy and InternVL3-1B for overall open-ended score), not aggregate results across the eight models. Please report aggregate statistics with per-model effect sizes, and either substantiate the universal claim or replace it with a more precise statement about which metrics and model sizes improve.","section":"Section IV.C and Table VII"},{"comment":"The open-ended evaluation is circular. The ground-truth answers are generated by GPT-4o and Claude 3 Opus and validated by GPT-4.1 (Section III.B), the open-ended responses are scored by GPT-4.1 as an LLM judge (Section III.D), and the fine-tuned models are trained to reproduce GPT-style answers. The reported 12.0% open-ended gain may therefore reflect alignment with the judge's stylistic preferences rather than improved safety-relevant reasoning. The 3-4% human-review error rate on 595 samples does not bound errors across the full training set. Please add a human evaluation of a held-out sample of open-ended responses, report judge-human agreement, and/or use an independent judge model from a different model family than the one used for annotation.","section":"Sections III.B, III.D, and IV.C"},{"comment":"The abstract states that fine-tuned models maintain \"real-time performance with a 0.59-second average inference time per image,\" but Table VIII shows that 0.59 s is specifically the InternVL3-1B result on an RTX 3090. Other configurations are substantially slower, e.g., Qwen2-VL-2B ranges from 0.74 to 0.87 s on the RTX 3090 and 2.93 to 3.70 s on Jetson Orin, and Qwen2.5-VL-3B takes 1.01 to 1.39 s on RTX 3090 and 5.63 to 6.86 s on Orin. Please report latency per model, size, quantization, and platform, and avoid implying that 0.59 s is a general average across the evaluated systems. In addition, the continuous-inference experiments at 30 Hz show many requests queued or dropped, so the claim of real-time suitability should be qualified relative to the actual frame rate and timeout policy.","section":"Section V.A and Abstract"},{"comment":"The real-world case-study evaluation is qualitative and lacks a systematic comparison: there are no baseline model outputs, no quantitative risk-detection or false-alarm rates, and no scoring rubric for the four scenarios. Moreover, the hidden-child scenario is reported as a failure (the model \"did not accurately capture and interpret the partially or fully obscured child\"), which is inconsistent with the abstract's claim that fine-tuned models \"correctly identify safety risks that challenge even experienced human drivers.\" Please either provide quantitative evaluation on a labeled real-world set with baseline comparison, or substantially soften the abstract and conclusion in line with the mixed results.","section":"Sections V.B, V.C, and Abstract"},{"comment":"No confidence intervals, error bars, or significance tests are reported for any of the quantitative comparisons. The test set is small (555 questions, 111 images), and several reported differences are only one or two points on the LLM-judge scales, so the improvements may not be statistically reliable. Please report confidence intervals or bootstrap estimates, and paired significance tests where appropriate, especially for the close-ended accuracy and captioning metrics.","section":"Tables VI and VII"}],"minor_comments":[{"comment":"The sentence \"Results in Table III show that proprietary LLM models can provide accurate and contextually relevant answers\" appears to cite the wrong table, since Table III shows sample dataset annotations rather than benchmark results; please re-check the reference.","section":"Section VI.B"},{"comment":"The paragraph beginning \"The multi-agent system employed varying temperature settings\" ends with \"format standardization.\" and is followed by a fragment beginning \"with close-ended and open-ended questions automatically categorized\". Please complete the sentence or merge it with the preceding one.","section":"Section III.C"},{"comment":"The conclusion repeats \"average inference time of 0.59 seconds per image\" without noting that this is the best-case InternVL3-1B result on a specific GPU; please qualify this statement as in the major comment above.","section":"Section V.A"},{"comment":"The hyperparameter details are deferred to the supplemental material; given that LoRA rank and fine-tuning hyperparameters directly affect the reported gains, please state at least the rank and learning rate in the main text.","section":"Section IV.B"}],"recommendation":"major_revision","confidential_remarks":"The dataset contribution is real and likely useful to the community, and the close-ended accuracy gains are consistent across models. The main risks are the circular open-ended evaluation and the overbroad claims that contradict the paper's own Table VII. I think a major revision is appropriate: the authors should re-analyze and re-report the open-ended results with human evaluation or an independent judge, correct the universal-improvement and real-time claims, and add statistical rigor. If the corrected open-ended results still show consistent gains, the paper would be publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is worth a serious referee, but not in its current form. What is actually new: DriveSOTIF is the first VQA/captioning dataset aimed at perception SOTIF, built from PeSOTIF images with a multi-agent pipeline and human review on 595 samples. That is a real resource, and the closed-ended accuracy improvements are credible: all eight fine-tuned rows in Table VII go up, and captioning metrics improve by large margins. The real-world case studies are also useful, including the honest failure on the child-hidden-in-a-box case, which shows the authors are not just selling the method.\n\nThe soft spots are load-bearing but fixable. First, the open-ended evaluation is circular: the ground-truth answers were generated by GPT-4o/Claude 3 Opus and validated by GPT-4.1, and the same GPT-4.1 is used as the LLM judge for open-ended scores. The 12.0% open-ended gain therefore measures alignment with one family of proprietary LLM opinions more than it measures safety-relevant reasoning. Second, the abstract and Sec. IV.C overstate the results. The 11.8%/12.0% numbers come from specific best-performing rows, not aggregate results, and Table VII itself contains multiple models whose relevance or overall scores get worse after fine-tuning. That directly contradicts the claim that fine-tuning improves performance across all evaluation metrics. Third, no confidence intervals or significance tests are reported, and the 0.59-second real-time claim is specifically InternVL3-1B on an RTX 3090, not a general result.\n\nThe reader's circularity concern lands, and the stress-test note shows the overclaim is visible in the paper's own table. The fix is not hard: report aggregate statistics across all models, separate close-ended from open-ended, and ground the open-ended evaluation in human-annotated subsets rather than GPT-4.1 alone. The dataset itself is reusable and the paper is a reasonable first step for SOTIF researchers who want a VQA benchmark. I would send it to review, with a request for major revision.","headline":"A useful new SOTIF VQA dataset and a reasonable fine-tuning study, but the headline gains are cherry-picked and the open-ended evaluation is GPT judging GPT.","tokens_in":19904,"tokens_out":1696,"would_cite":true,"duration_ms":16675,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuning a small multimodal language model on a SOTIF-specific VQA dataset lets it detect, explain, and recommend responses to perception-related driving hazards in near real time, with a measured 11.8% gain on close-ended and 12.0% on…","keywords":["SOTIF","Safety of the Intended Functionality","multimodal large language models","visual question answering","autonomous driving","LoRA fine-tuning","LLM-as-judge","long-tail traffic scenarios"],"falsifier":"Have a panel of human driving-safety experts label several hundred unseen SOTIF images for risk presence, cause, and recommended action, then compare the fine-tuned model's answers to those human labels rather than to LLM-generated references; if agreement is near chance on cases where experts agree, the claim collapses. The paper's own child-hidden-in-a-box case is a partial falsifier already, since the model misses the hazard.","tokens_in":18745,"feed_emoji":"🚗","tokens_out":10747,"duration_ms":93282,"temperature":0.7,"pith_summary":"This paper claims that a general-purpose multimodal large language model, fine-tuned on a purpose-built visual question-answering and captioning dataset, can take on the perception-related Safety of the Intended Functionality (SOTIF) reasoning that autonomous vehicles currently lack. The authors introduce DriveSOTIF, the first SOTIF-specific VQA dataset, and report that fine-tuning lifts close-ended accuracy by 11.8% and open-ended rubric scores by 12.0% over baselines, at 0.59 seconds per image on a consumer GPU. Real-world tests in Canada and China show the model identifying snowy-road, merging-truck, and headlight-glare hazards, though one extremely subtle case (a child hidden in a box) is missed. If the result holds, a small onboard model could act as a near-real-time, explainable SOTIF risk-assessment layer for autonomous driving.","feed_headline":"Small fine-tuned AI model flags road hazards in 0.59 seconds","feed_subtitle":"Domain-specific fine-tuning lifts SOTIF risk-question accuracy by ~12% and runs on a 2 GB onboard GPU.","key_machinery":"The central object is DriveSOTIF, a dataset of 1,114 images, 1,114 captions, and 5,570 question-answer pairs built from a long-tail perception-SOTIF image collection. The generative machinery is a multi-agent pipeline in which vision-language models alternate roles for captioning, question generation, answer generation, and validation, with failures regenerated and a second model checking image relevance, question suitability, and answer correctness. The training machinery is LoRA-based supervised fine-tuning of open-source multimodal models, and the evaluation machinery for open-ended answers is an LLM-as-judge scoring relevance, trustworthiness, clarity, and coherence. Together these parts turn a generic visual question-answering model into a domain-specific SOTIF risk assessor.","core_discovery":"The paper's central claim is that domain-specific supervised fine-tuning gives multimodal LLMs the spatial and causal intelligence that human drivers use to judge safety: open-world generalization to unseen hazards, causal reasoning about combinations of conditions such as rain, night, and glare, and contextual understanding of scene actors. Fine-tuned models produce longer, situation-aware captions and more detailed, scenario-specific visual question-answering responses than their baselines. The largest gains appear in small models, with the 1-billion-parameter model's accuracy rising from 55.7% to 63.5% and its overall judge score from 3.84 to 4.30, making the capability compatible with onboard, resource-limited deployment. The paper also reports a limit: the model misses a child hidden in a box, which it attributes to the visual encoder's capacity.","pith_inferences":["Beyond the paper: because both training answers and evaluation scores come from large language models, part of the measured gain may be alignment with those models' prior opinions rather than with safety truth; a deployment-grade claim needs human-expert or physical ground truth.","Beyond the paper: the same generation-and-fine-tuning pipeline should transfer to LiDAR, radar, and bird's-eye-view inputs; a multi-sensor SOTIF risk dataset is a natural next test and would show whether the approach scales beyond camera images.","Beyond the paper: evaluating a DriveSOTIF-fine-tuned model on an independent corner-case benchmark for risk localization would separate genuine SOTIF reasoning from dataset-specific phrasing.","Beyond the paper: the reported trade-off between response time and concurrency suggests a fast-slow architecture in practice, with lightweight perception running continuously and the MLLM triggered for semantic risk assessment rather than invoked on every frame at high speed."],"forward_implications":["A model with roughly 1 billion parameters and about 2 GB of GPU memory can serve as an onboard, near-real-time SOTIF risk assessment module instead of a cloud call.","Fine-tuning lifts both captioning and visual question answering across model sizes, with the largest gains in small models, so resource-constrained vehicles benefit most.","Learned SOTIF reasoning transfers across countries and weather conditions, suggesting a single fine-tuned model can cover diverse operational design domains.","Connecting the model to a decision-making layer would let an autonomous vehicle factor perception-risk explanations into real-time planning, with continuous improvement via human-in-the-loop data.","The approach does not yet handle extremely subtle hazards, as the hidden-child case shows, so a safety case would need to bound the miss rate on such scenarios."],"supporting_citations":[{"why":"Supplies the long-tail perception-SOTIF images from which DriveSOTIF is built and defines the problem domain.","marker":"[4]"},{"why":"Representative road-corner-case dataset that motivates the need for long-tail perception data.","marker":"[7]"},{"why":"Prior VQA dataset for driving corner cases that does not directly target SOTIF risk, the gap DriveSOTIF fills.","marker":"[9]"},{"why":"Existing driving VQA benchmark using template-based questions, motivating the open-ended, natural question design.","marker":"[24]"},{"why":"LoRA is the parameter-efficient fine-tuning technique that makes small-model adaptation feasible.","marker":"[42]"},{"why":"Establishes the LLM-as-judge evaluation approach used for open-ended answer scoring.","marker":"[34]"},{"why":"Supplies the multi-criteria scoring rubric (relevance, trustworthiness, clarity, coherence) used by the judge.","marker":"[35]"},{"why":"Real-world adverse-weather driving data used in the Canadian validation case.","marker":"[49]"}],"fun_headline_variants":["Fine-tuned MLLM spots road risks in 0.59s","Small AI model boosts hazard detection by 12%","On-device MLLM improves SOTIF perception 12%","Domain tuning lifts AV hazard accuracy 12%","Tiny MLLM catches hazards humans miss in 0.59s"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model's notion of a correct SOTIF risk assessment is inherited from the large language models that wrote the answers and from the language model that scores them, so the reported gains could measure agreement with those models' opinions rather than with safety truth.","fun_headline_variants_meta":{"raw":{"variants":["Fine-tuned MLLM spots road risks in 0.59s","Small AI model boosts hazard detection by 12%","On-device MLLM improves SOTIF perception 12%","Domain tuning lifts AV hazard accuracy 12%","Tiny MLLM catches hazards humans miss in 0.59s"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000188,"raw_usage":{"total_tokens":1321,"prompt_tokens":926,"completion_tokens":395,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":306}},"tokens_in":542,"tokens_out":395,"duration_ms":4288,"temperature":1.0,"reasoning_tokens":306,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:24:56.161016+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a panel of human driving-safety experts label several hundred unseen SOTIF images for risk presence, cause, and recommended action, then compare the fine-tuned model's answers to those human labels rather than to LLM-generated references; if agreement is near chance on cases where experts agree, the claim collapses. The paper's own child-hidden-in-a-box case is a partial falsifier already, since the model misses the hazard.","supporting_citations":[{"cited_title":"Pesotif: A challenging visual dataset for perception sotif problems in long-tail traffic scenarios,","cited_arxiv_id":null,"evidence_quote":"Supplies the long-tail perception-SOTIF images from which DriveSOTIF is built and defines the problem domain."},{"cited_title":"Canadian adverse driving conditions dataset,","cited_arxiv_id":null,"evidence_quote":"Real-world adverse-weather driving data used in the Canadian validation case."}],"review_version":1}