{"id":"3ad399a2-3ca9-4b3a-af91-a63f06cf0711","arxiv_id":"2506.19683","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Ultrasound images are converted into scene graphs that, together with a large language model, generate plain-language explanations and probe movement guidance for non-expert users.","lead":"The paper converts ultrasound images into small semantic scene graphs that describe anatomical structures and their relationships, then uses a large language model to explain the image and to suggest how to move the probe. It targets non-expert and point-of-care users, but the validation is limited to 27 test images from a small volunteer cohort.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Scanning guidance is reported without a closed-loop validation: probe motion rests on an unspecified 'lateral movement' mechanism, and Task II only scores LLM text, leaving the guidance half of the central claim unsupported.","rationale":"I read the paper as a proof-of-concept that an SG extracted by RelTR can feed an LLM to explain carotid ultrasound images and to suggest probe movement that brings missing anatomies into view. The authors are appropriately cautious in the Discussion about carotid-only validation and small data; they also report standard SG metrics and compare many LLMs, which is useful evidence. My concern is not with the SG metrics or the explanation task. It is that Task II is evaluated only at the level of generated text. Section 2's lateral-movement mechanism is a single sentence and is never operationalized; Section 3.3 measures Acc, METEOR, and ROUGEL against GPT-4o reference texts, not any scanning outcome. Thus the second half of the central claim is currently an unvalidated pipeline stage. This is a correctness-risk concern: the proposed guidance loop could fail in a way that no amount of SG accuracy or LLM fluency would reveal. A closed-loop scanning experiment with a clear success criterion, namely whether the target anatomy appears, is the minimal check that would settle it. I therefore keep the reader's CONDITIONAL verdict rather than moving it, but the condition should explicitly include validating the guidance loop, not only improving data diversity.","tokens_in":7484,"tokens_out":6910,"duration_ms":74121,"concrete_test":"Run a closed-loop scanning experiment with a fresh volunteer and a portable probe: start from a view missing one of the five entities, have a non-expert or scripted holder follow the LLM's guidance for a fixed time, and measure whether the target anatomy appears in the field of view. Compare against a no-guidance or random-sweep baseline across at least 10 start views, and log lateral-movement estimates from consecutive frames to verify the signal when a visible anatomy is tracked.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2 states that probe lateral movement is obtained \"by comparing two consecutive detections of the target anatomy,\" but Section 3 never describes or evaluates that procedure. No consecutive-frame dataset is mentioned, and the 27 test images are described as a static image set. Task II in Table 2 therefore measures whether an LLM can turn a supplied SG plus a supplied, unvalidated lateral-movement label into text that matches GPT-4o-derived reference text, as judged by experts. It does not test whether a user can actually move a probe to reveal a missing anatomy, whether the motion estimate is correct, or whether the generated instructions are actionable in a real scan. This is load-bearing because the paper's central claim includes scanning guidance: if the lateral-movement signal is unreliable or the text-to-motion mapping is ambiguous, the guidance half fails even under perfect scene-graph prediction. The problem is internal validity, not merely out-of-distribution generalization; the concern would remain if the SG model were retrained on a larger dataset. The Discussion's limitation about carotid-only validation is real but does not address this missing guidance-loop evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a scene-graph-based framework for ultrasound image understanding and scanning guidance. A RelTR transformer predicts triplets of the form <entity-predicate-entity> for five carotid anatomies, and the resulting triplets, along with lateral side and lateral movement information, are fed to an LLM to generate image explanations (Task I) and probe motion instructions (Task II). Experiments on a 289-image dataset evaluate scene graph prediction with standard detection and recall metrics, and LLM outputs are scored with subjective accuracy, METEOR, and ROUGE-L against GPT-4o-derived references.","tokens_in":7704,"tokens_out":3529,"duration_ms":36308,"significance":"If the framework works as claimed, it could make ultrasound interpretation and basic probe navigation accessible to non-expert users in point-of-care settings. The paper is among the first to apply semantic scene graphs to ultrasound, and using a single-stage transformer avoids an explicit object detection step. The scene graph predictor is evaluated with standard metrics and an ablation over encoder-decoder layer counts, and the LLM comparison spans several open and closed models. However, the experiments do not currently demonstrate that the scene graph adds value beyond the LLM itself, and the scanning-guidance half of the central claim is not validated in a closed loop.","major_comments":[{"comment":"The lateral movement signal is load-bearing for the central scanning-guidance claim, but it is never defined or evaluated. The paper states only that lateral movement is obtained 'by comparing two consecutive detections of the target anatomy,' and no consecutive-frame dataset, algorithm description, or validation appears in Section 3. The 27 test images are described as a static image set, and Table 2 scores LLM-generated text rather than probe motion. As a result, the experiments do not show that a user can actually move the probe to reveal a missing anatomy, even if the scene graph predictions were perfect.","section":"Section 2 (US Scanning Guidance) and Section 3.3 (Task II)"},{"comment":"No baseline or ablation isolates the contribution of the scene graph. All LLM conditions appear to receive the same SG-derived grounding prompt, and there is no condition without the scene graph (for example, raw image captioning, object-list-only prompts, or a direct vision-language model). Without such a comparison, the claim that the SG representation is what enables explanation and guidance is not supported.","section":"Section 3.3, Table 2"},{"comment":"The subjective 'Acc' score is reported to varying precision but without any definition of the rubric, the number of third-party experts, their clinical or ultrasound background, or inter-rater agreement (for example, Cohen's kappa). The reference texts were generated using GPT-4o and 'followed by manual verification,' but the verification protocol is not described. These omissions make the headline accuracies, including the 1.000 scores for Gemma 2 and Grok 3 in Task I, difficult to interpret.","section":"Section 3.3, Evaluation Metrics and Table 2"},{"comment":"The test set consists of 27 images from a single ultrasound machine and probe, and the Discussion explicitly states that the method is only validated on carotid images. This is an honest limitation, but the abstract and introduction frame the method as a general step toward 'democratizing ultrasound' beyond the carotid region. The paper should either temper the general claim or provide evidence across additional anatomies and devices; as written, the framing overstates the external validity.","section":"Section 3.1 and Discussion"}],"minor_comments":[{"comment":"The phrase 'to explain image content to ordinary and provide guidance' should read 'to ordinary users' or 'to non-experts'; similar grammatical issues appear elsewhere, for example 'for ordinaries' in the abstract.","section":"Abstract and Introduction"},{"comment":"There are several typos and formatting issues, including 'tirplets' and 'thetirplets' in Section 2, and 'T able' in the Table 1 caption; these should be corrected.","section":"Section 2 and Section 3"},{"comment":"The grounding prompt structure is only described verbally; providing the actual prompt template or a precise pseudo-code description would improve reproducibility, since Task II depends on the exact wording and ordering of the grounding information.","section":"Section 2"},{"comment":"The identification of the left or right lateral side is mentioned as an output of the system but is never evaluated; the paper should explicitly state that lateral-side accuracy is not measured or include an evaluation of this component.","section":"Section 3.2"},{"comment":"The precision of the Acc values is inconsistent across entries, ranging from 0.265 to 1.000; the authors should specify the intended number of decimal places or the number of trials on which each accuracy is based.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The scene graph prediction experiments are competently executed with standard metrics, but the scanning-guidance claim is under-validated, and the added value of the scene graph over a plain LLM is not demonstrated. A revision that adds a non-SG baseline and an evaluation of the lateral-motion estimation would address the most serious gaps. I do not see a novelty disclosure concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it really is the first to build a semantic scene graph for ultrasound images, with a sensible entity/predicate taxonomy for carotid and thyroid anatomy, and it shows a clean way to feed that graph to an LLM for plain-language explanation. Second, the paper's title and abstract promise \"scanning guidance,\" but that half of the system is never validated. Section 2 says probe lateral movement is obtained by comparing two consecutive detections; Section 3 never describes how that works or evaluates it on any consecutive-frame data. Task II only measures whether an LLM can turn a supplied scene graph plus a supplied, unvalidated lateral-movement label into text matching GPT-4o-generated references. The stress-test note is right: this is an internal validity gap, not just an out-of-distribution generalization issue. If the lateral-movement estimate is wrong, the guidance collapses even with perfect scene graphs.\n\nWhat the paper does well: the motivation is clear, the SG prediction evaluation uses standard metrics (mAP, Recall@K, mR@K), the choice of RelTR as a single-stage method is reasonable, and the authors are honest that the dataset is small and carotid-only. The lightweight-versus-large LLM comparison is useful and gives a practical read on what might run on a portable device. The custom annotation tool and the explicit discussion of why data augmentation is hard for SG prediction are also good signs of careful thinking.\n\nWhere it is soft: the 27-image test set is very small, there is no baseline isolating whether the SG representation adds value over, say, just feeding object detections to the LLM, and the subjective Accuracy score has no inter-rater reliability reported. None of these are fatal for a proof-of-concept, but they matter because the headline claims are about democratizing ultrasound for non-experts, and no non-expert user study was run. The paper's own discussion acknowledges the carotid-only limitation, which I take as good faith, but it does not address the missing guidance-loop validation.\n\nThe citation pattern is fine—the authors lean on their own prior robotic ultrasound work, but that is standard in a methods continuity sense and not a problem here. There is no fitted-constant-as-prediction circularity; the SG predictions are checked against manually annotated ground truth and the LLM outputs against verified references.\n\nWho is this for? Researchers working on point-of-care ultrasound, image-guided self-learning, or scene-graph applications in medical imaging. They will get a clear, honest proof of concept and a concrete taxonomy they can build on. It deserves a serious referee, but the referee should push hard on the guidance claim: either remove it from the scope or run a real probe-motion experiment. I would accept it for peer review with a clear expectation of major revision, and I would probably cite the taxonomy even while being explicit that the guidance part is unvalidated.","headline":"A genuine first application of scene graphs to ultrasound with an honest small-data story, but the scanning-guidance half of the central claim is not actually tested—only an LLM text-generation task is evaluated.","tokens_in":8201,"tokens_out":1331,"would_cite":true,"duration_ms":16215,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A predicted scene graph of anatomical triplets is enough for an LLM to explain an ultrasound image and to say which way the probe should move to find missing anatomy.","keywords":["ultrasound image analysis","scene graph","large language models","image explanation","scanning guidance","point-of-care ultrasound","relation transformer","carotid artery"],"falsifier":"Run the trained system on, say, fifty carotid scans from new volunteers acquired with a different ultrasound machine and probe, and measure relation recall against manually annotated triplets. If recall on new acquisitions drops to near chance level, or if the LLM's missing-anatomy guidance names the wrong absent structure in more than a small fraction of cases, the scene-graph pipeline fails at its load-bearing step and the explanation and guidance claims fall with it.","tokens_in":1617,"feed_emoji":"🩺","tokens_out":1769,"duration_ms":74718,"temperature":0.7,"pith_summary":"The paper proposes that a single ultrasound image can be reduced to a small semantic scene graph: a set of relation triplets such as \"thyroid is contiguous with carotid.\" It then argues that this graph is sufficient for an LLM to perform two useful jobs for a non-expert: explain what the image shows in plain language, and tell the user which way to move the probe so an absent anatomy comes into view. The system is built for carotid and thyroid neck scans, using a single-stage transformer trained on 262 annotated images and tested on 27 images from different volunteers. The motivation is democratizing ultrasound: a portable device could coach a lay user through a basic scan without a clinician on site and without training a large ultrasound-specific vision-language model.","feed_headline":"Scene graphs turn ultrasound images into plain-language scan coaching","feed_subtitle":"A compact graph of neck anatomy plus an LLM explains scans and tells novices which way to move the probe.","key_machinery":"The load-bearing object is the ultrasound scene graph: a set of relation triplets $<\\text{entity}_1,\\ \\text{predicate},\\ \\text{entity}_2>$ over five anatomies and three predicates, generated by a transformer-based one-stage relation detector that predicts subjects, objects, and predicates in a single pass. This graph carries the argument because everything downstream, from neck-side inference and lateral-movement estimation to the LLM explanation and missing-anatomy guidance, is computed from the triplets rather than from raw pixels or a full clinical report. The graph is deliberately compact and concept-level, which is what lets a small model predict it and lets an LLM reason over it in a resource-constrained portable setting.","core_discovery":"The central claim is that a semantic scene graph predicted directly from a 2D ultrasound image, using five anatomical entities and three spatial predicates, provides a sufficient intermediate representation for patient-friendly explanation and probe guidance. The graph is predicted end-to-end by a one-stage relation transformer, so no separate object detector and no large ultrasound-specific vision-language model is required. The same triplets are fed to an LLM as a grounding prompt, together with the inferred side of the neck and the probe's lateral movement direction, letting the LLM answer a user's question about a focus anatomy and identify an anatomy missing from the current view with a recommended movement direction. The authors validate the pipeline on carotid-region images and report that larger-capacity LLMs follow the graph more accurately and produce better summaries and guidance than small quantized models; they also state explicitly that the method is currently validated only on carotid images.","pith_inferences":["Editor's inference: if the scene graph is the reason the pipeline works, the same graph could serve as a state representation for robotic ultrasound, directly steering a motorized probe until the missing anatomy triplet appears rather than only telling a human which way to move.","Editor's inference: the paper measures text quality and instruction-following, not whether a non-expert actually completes a more standardized scan; a testable extension would give novices the system and measure scan completion time, number of views, or expert-rated completeness against unassisted scanning.","Editor's inference: the lateral-movement signal derived from consecutive detections suggests a cheap, label-free probe-motion estimate; a natural extension is to feed a full video stream through the same graph predictor to build a temporal scan trajectory and catch drift.","Editor's inference: since only carotid images were validated, the strongest next test is an out-of-domain transfer experiment with a different probe frequency or a different neck region; if the predicate vocabulary transfers poorly, a per-anatomy vocabulary may be necessary."],"forward_implications":["If the scene graph is reliably predicted, a portable ultrasound device can explain a carotid scan to a lay user in plain language without streaming images to a remote expert.","The same predicted graph can drive scanning guidance: the LLM can name an anatomy missing from the current view and say which way to move the probe, supporting more complete self-scans.","Because the representation is a compact graph rather than a full report, the approach avoids training a large ultrasound-specific vision-language model, which matters when annotated ultrasound data are scarce.","The dependence on only a handful of entities and predicates suggests the framework could be adapted to other anatomies by redefining the entity and predicate vocabulary and retraining on a new small dataset.","Even lightweight local LLMs can perform the explanation task with reasonable accuracy, pointing toward on-device deployment in portable ultrasound hardware."],"supporting_citations":[{"why":"Supplies the single-stage transformer that predicts subjects, predicates, and objects in one pass, the core of the scene-graph pipeline.","marker":"[7]"},{"why":"Introduces scene graphs for medical images (CT), the precedent the paper extends to ultrasound.","marker":"[26]"},{"why":"Generates the reference texts used to score the LLM's summaries and guidance after manual verification.","marker":"[1]"},{"why":"Defines the COCO AP@50 and AP@[50:95] protocols used to evaluate object detection.","marker":"[19]"},{"why":"Defines Recall@K, the relation-detection metric used in the experiments.","marker":"[20]"},{"why":"Defines mean Recall@K, used alongside Recall@K for relation detection.","marker":"[6]"},{"why":"Provides the METEOR metric used to measure LLM output similarity to reference texts.","marker":"[4]"},{"why":"Provides the ROUGE-L metric used to measure LLM output similarity to reference texts.","marker":"[18]"}],"fun_headline_variants":["Scene graphs make ultrasound images explainable and guide scanning","Ultrasound scene graphs power plain-language explanations and probe tips","LLM plus scene graph decodes ultrasound for non-experts and guides scans","Carotid-tested scene graphs from ultrasound: explain and coach movement","Semantic scene graph from ultrasound: explanation and scanning guidance"],"cache_read_input_tokens":10368,"weakest_assumption_plain":"The whole pipeline collapses if the relation transformer, trained on 262 carotid images from five volunteers on one ultrasound machine, does not predict correct scene graphs for new users, different probes, or other anatomies, and the paper only tests 27 images from different volunteers while stating that only carotid images were validated.","fun_headline_variants_meta":{"raw":{"variants":["Scene graphs make ultrasound images explainable and guide scanning","Ultrasound scene graphs power plain-language explanations and probe tips","LLM plus scene graph decodes ultrasound for non-experts and guides scans","Carotid-tested scene graphs from ultrasound: explain and coach movement","Semantic scene graph from ultrasound: explanation and scanning guidance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1447,"prompt_tokens":947,"completion_tokens":500,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":413}},"tokens_in":563,"tokens_out":500,"duration_ms":5975,"temperature":1.0,"reasoning_tokens":413,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:27:25.874925+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained system on, say, fifty carotid scans from new volunteers acquired with a different ultrasound machine and probe, and measure relation recall against manually annotated triplets. If recall on new acquisitions drops to near chance level, or if the LLM's missing-anatomy guidance names the wrong absent structure in more than a small fraction of cases, the scene-graph pipeline fails at its load-bearing step and the explanation and guidance claims fall with it.","supporting_citations":[{"cited_title":"IEEE Transactions on Pattern Analysis and Machine Intelli- gence 45(9), 11169–11183 (2023)","cited_arxiv_id":null,"evidence_quote":"Supplies the single-stage transformer that predicts subjects, predicates, and objects in one pass, the core of the scene-graph pipeline."},{"cited_title":"In: MICCAI","cited_arxiv_id":null,"evidence_quote":"Introduces scene graphs for medical images (CT), the precedent the paper extends to ultrasound."},{"cited_title":"In: Computer vision– ECCV 2014: 13th European conference, zurich, Switzerland, September 6-12, 2014, proceedings, part v 13","cited_arxiv_id":null,"evidence_quote":"Defines the COCO AP@50 and AP@[50:95] protocols used to evaluate object detection."},{"cited_title":"In: Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14","cited_arxiv_id":null,"evidence_quote":"Defines Recall@K, the relation-detection metric used in the experiments."},{"cited_title":"In: CVPR","cited_arxiv_id":null,"evidence_quote":"Defines mean Recall@K, used alongside Recall@K for relation detection."},{"cited_title":"In: Text sum- marization branches out","cited_arxiv_id":null,"evidence_quote":"Provides the ROUGE-L metric used to measure LLM output similarity to reference texts."}],"review_version":2}