{"id":"ddd1ca99-5148-4860-9838-a736264437b8","arxiv_id":"2506.13956","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Synthetic dialogues and images generated by LLMs and diffusion models improve LLaVA's zero-shot action selection on the Do-I-Demand robotic assistance benchmark.","lead":"The paper creates realistic fake conversations and images of everyday rooms to train a robot to understand indirect requests. The authors report that this synthetic data improves a vision-language model's accuracy on a real robotic assistance benchmark.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Action-based augmentation conditions BLIP-Diffusion on an undisclosed 'reference image from our real-world data collection'; if that image is a Do-I-Demand evaluation sample, the reported gains are leakage rather than synthetic-data generalization.","rationale":"The provenance of the reference image is the most load-bearing issue because it determines whether the central claim is about synthetic-data generalization or about leaking a target-domain image into the generation process. The reader's weakest assumption identifies exactly this point, and I agree. The other weaknesses noted by the reader—missing external SOTA comparisons, no error bars, no code/data release—are important but secondary: they are reporting deficiencies that could be repaired without altering the experimental design. A leakage of the reference image, by contrast, would invalidate the main conclusion. The paper is internally consistent and the ablations are coherent, and nothing here suggests intentional misconduct; the concern is an undeclared, testable provenance issue. The appropriate verdict remains CONDITIONAL, with release/disclosure of the reference image and a disjoint-source rerun as the conditions.","tokens_in":10324,"tokens_out":7526,"duration_ms":87603,"concrete_test":"Ask the authors to disclose the reference image(s) used in Section 3.2 and verify that none is a Do-I-Demand sample, ideally by releasing the generation script. Then independently rerun the best configuration (both augmentations) with a reference image taken from a disjoint room-scene source that is not part of Do-I-Demand, and compare Table 1. If accuracy drops substantially (e.g., more than 5 points) or the action-based augmentation gain disappears, the reported result is contaminated; if accuracy holds within noise, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that synthetic dialogues and images alone produce the reported zero-shot gains on Do-I-Demand—requires that the action-based augmentation pipeline does not touch target-domain images. Section 3.2 says BLIP-Diffusion generation 'incorporate[s] a reference image from our real-world data collection,' but the paper never identifies this image, how many reference images are used, or whether the image comes from the same 400-sample Do-I-Demand evaluation set. The authors include co-authors of the Do-I-Demand dataset, and no train/eval split is described; it is therefore plausible that the reference image is one of the evaluation images. Because BLIP-Diffusion is designed to preserve the subject and visual style of the reference, the generated action-based training images would inherit room layout, objects, and lighting from a target-domain image. In that case, the Table 1 gains (e.g., 31.5 vs 20.3 for LLaVA-13B utterance-only) and the Table 3 ablation (6-point drop when BLIP is replaced by SDXL) would partly measure target-domain leakage rather than the value of synthetic augmentation. This is the assumption that, if violated, breaks the zero-shot generalization interpretation of the main result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes ASMR, a data augmentation framework for robotic life-support action classification. It uses GPT-3.5 to generate human-robot dialogues in two modes (place-based and action-based) and diffusion models (SDXL and BLIP-Diffusion) to generate corresponding environmental images. LLaVA-7B/13B is fine-tuned with LoRA on the synthetic data and evaluated zero-shot on the Do-I-Demand benchmark (400 samples). Tables 1 through 3 report consistent accuracy gains from augmentation, with a best accuracy of 48.8%, and the bucket analysis in Section 4.5 shows improvements on labels that initially had zero accuracy. The paper claims that the approach achieves state-of-the-art performance.","tokens_in":10558,"tokens_out":4747,"duration_ms":48514,"significance":"If the reported gains are genuine, the framework would substantially reduce the cost of collecting real human-robot interaction data for intent-to-action mapping, and the within-paper evidence is encouraging: gains appear across two base model sizes, two response encoders, and two augmentation pathways, with ablations isolating the contribution of the BLIP-Diffusion image generator and a prompt-variation study. However, the central claim is not yet fully supported because the evaluation lacks statistical grounding, external comparisons, and a clear guarantee that the image-generation pipeline does not touch the evaluation partition. The paper also provides no code or data release, so the synthetic-data pipeline is not independently reproducible as reported.","major_comments":[{"comment":"The reference image used to condition BLIP-Diffusion is described only as 'a reference image from our real-world data collection,' with no identification of the image, the number of reference images, or the partition of Do-I-Demand from which it is drawn. Because BLIP-Diffusion preserves the subject and visual style of the reference, and because the paper never describes a train/evaluation split of the 400-sample benchmark, the gains in Table 1 (for example, 31.5 vs. 20.3 for LLaVA-13B utterance-only with SBERT) and the 6-point drop in Table 3 when BLIP is replaced by SDXL could partly reflect leakage of evaluation-domain images into the training set rather than the value of synthetic augmentation. Please disclose the reference-image provenance, ensure it is disjoint from the evaluation set, and report results with the reference image held out from both generation and evaluation.","section":"Section 3.2"},{"comment":"The paper is inconsistent about whether any real Do-I-Demand training samples are used. Section 1 says training is 'supplemented with a small, real-world dataset,' while Section 4.1 states 'We fine-tune the base models using our augmentation dataset' and evaluates zero-shot accuracy on the evaluation dataset. Specify exactly which real data, if any, are used for fine-tuning, and describe how the 400 samples are divided into any training and evaluation subsets; otherwise the 'zero-shot' interpretation of the results cannot be assessed.","section":"Section 4.1"},{"comment":"The claims of 'significant improvement' and 'state-of-the-art performance' are not supported by any statistical evidence or external comparison. The table reports single accuracies with no error bars, no number of seeds, no significance tests, and no comparison against previously published results on Do-I-Demand. Please add repeated-seed experiments with variance or confidence intervals and a comparison table that includes prior published results on this benchmark before claiming state-of-the-art performance.","section":"Section 4.2 and Table 1"},{"comment":"No human or automatic quality check of the generated dialogues or images is reported, and the diverse-prompt experiment in Table 2 shows that some prompt variants decrease accuracy (for example, LLaVA-7B utterance-only drops from 30.3 to 23.3). A quality-filtering step or at least a small human evaluation of the generated data would strengthen the claim that the improvements come from realistic synthetic scenarios rather than from accidental distribution matching or from the specific prompt wording.","section":"Section 3.1 and Section 4.3"}],"minor_comments":[{"comment":"The manuscript contains inconsistent typography, including 'LLaV A' with a space instead of 'LLaVA' and 'DO-I-DEMAND' versus 'Do-I-Demand'; please unify these.","section":"Throughout"},{"comment":"The 'Description + Utterance' setting is described as combining the human request with 'a description of the environment, inferred from an image'; please clarify whether the actual image is also provided to LLaVA or only the text description, since this affects the interpretation of the multimodal claim.","section":"Section 4.1"},{"comment":"Please define the bucket construction precisely, including the number of labels per bucket and whether the buckets are computed from baseline or augmented-model predictions, and add y-axis labels to Figures 3 and 4.","section":"Section 4.5 and Figures 3-4"},{"comment":"The phrase 'large-scale dataset' overstates the synthetic data size (10 places and 43 actions, each with ten dialogues, yields roughly 530 examples); consider a more neutral description.","section":"Abstract and Section 3"},{"comment":"The prompt templates for dialogue generation are given, but the image-generation prompts for both SDXL and BLIP-Diffusion are not; please include the exact image prompts for reproducibility.","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The author team includes co-authors of the Do-I-Demand dataset (Ref. [33]), and the paper does not describe any train/eval split for that benchmark. Combined with the undisclosed BLIP-Diffusion reference image, this makes the data-provenance question the single most important issue to resolve before the results can be trusted. The editor may wish to insist on a clear statement of which specific image(s) are used as references and a verification that they are disjoint from the evaluation set. The absence of external comparisons also makes the 'state-of-the-art' claim difficult to evaluate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a straightforward, sensible extension: use GPT-3.5 to write ambiguous-request dialogues conditioned on either places or action labels, render the described rooms with SDXL or BLIP-Diffusion, and fine-tune LLaVA on the synthetic set alone. Within-paper, the augmentation helps consistently, and the bucket analysis shows the gains are concentrated on labels that were at zero. That is a real, useful finding for anyone facing small robot-interaction datasets.\n\nSecond, the stress-test concern lands. The action-based pipeline feeds BLIP-Diffusion a \"reference image from our real-world data collection\" and never says which image, how many, or whether it comes from the Do-I-Demand evaluation set. BLIP-Diffusion preserves the subject and visual style of that reference, so if the reference is an evaluation image, the generated training images inherit the same rooms, objects, and lighting as the test set. The authors include Do-I-Demand co-authors and no train/eval split is described, so this is not a paranoid reading. The paper's zero-shot generalization claim depends on the reference image being outside the evaluation partition. As written, that is uncheckable. This is the load-bearing soft spot.\n\nThe rest of the weaknesses are more ordinary. No error bars or significance tests on 400 samples. No comparison to prior published numbers on Do-I-Demand, so \"state-of-the-art\" is unsupported. No code or data release. The ablation swapping BLIP for SDXL is useful, but it also cannot separate the value of subject-preserving generation from the leakage concern. If the provenance is cleaned up, the method looks like a legitimate contribution; if not, the main gains could largely be test-set leakage.\n\nI would send this to peer review, because the question is important and the fix is straightforward: disclose and verify the reference image partition, add confidence intervals, and compare against the published Do-I-Demand baselines. But I would not cite it as evidence for synthetic-data generalization until that is resolved. For a reading group, it is a good case study in how easy it is to leave a data-provenance hole in an augmentation pipeline.","headline":"Useful augmentation study with an unresolved leakage risk: the action-based pipeline's BLIP-Diffusion reference image is undisclosed, making the zero-shot claim uncheckable as written.","tokens_in":11070,"tokens_out":2176,"would_cite":false,"duration_ms":23708,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that synthetic scenario data—LLM-written dialogues plus diffusion-drawn room images—can replace real human-robot interaction samples when fine-tuning a vision-language model for robotic action selection, raising…","keywords":["robotic action selection","multimodal classification","data augmentation","large language models","diffusion models","visual instruction tuning","ambiguous requests","Do-I-Demand benchmark"],"falsifier":"Check which image from the real-world data collection was used as the BLIP-Diffusion reference and which partition it belongs to. If that image is in the Do-I-Demand test set, retraining without it should be run, and if accuracy drops to baseline, the augmentation's reported gains are not evidence of generalization. A second check is to generate action-based images without any reference image and compare accuracy, since the paper's ablation only swaps BLIP-Diffusion for SDXL rather than removing the conditioning.","tokens_in":10142,"feed_emoji":"🤖","tokens_out":4763,"duration_ms":42245,"temperature":0.7,"pith_summary":"The paper tries to establish that a robot's ability to infer the right life-support action from an ambiguous human request plus a camera view can be improved without collecting more real human-robot interaction data. It generates synthetic training material: a large language model writes plausible dialogues between human and robot, and diffusion models draw the room from the robot's perspective. Fine-tuning LLaVA on this synthetic data, with no real target-domain training samples, raises zero-shot accuracy on the Do-I-Demand benchmark from 34.8% to 47.8% (13B, utterance-plus-description setting). The claim matters because real interaction data is expensive and slow to collect, and this pipeline offers a scalable substitute.","feed_headline":"Synthetic data lifts robot action accuracy to 48.8%","feed_subtitle":"LLM-written dialogues and diffusion-drawn images alone beat unmodified LLaVA on the 43-action Do-I-Demand benchmark.","key_machinery":"The machinery is a two-path data generation pipeline. Place-based augmentation prompts gpt-3.5 for ten ambiguous-request dialogues per location, while action-based augmentation prompts for dialogues that lead to one of 43 predefined robot actions, then BLIP-Diffusion generates an image from the environment description using a reference image from the real data collection, with 'room' as the constant subject. The resulting image-text pairs fine-tune LLaVA via LoRA, and at inference a sentence encoder matches LLaVA's free-form response to the action label.","core_discovery":"The central discovery is that action-tied synthetic dialogues, combined with diffusion-generated images conditioned on a reference room, transfer to real-world action selection. The paper reports that action-based augmentation generally beats place-based augmentation, and the two combined give the best results, with both LLaVA-7B and LLaVA-13B improving substantially over unmodified baselines. It also reports that replacing BLIP-Diffusion with SDXL hurts accuracy, indicating that the reference-image conditioning is doing real work.","pith_inferences":["Beyond the paper, the reported 'zero-shot' accuracy is zero-shot only with respect to the target dataset; the generation pipeline still saw a reference image from that real-world collection, so a stricter test would withhold any real image from the generation pipeline entirely.","Beyond the paper, the benchmark has only 400 samples and 43 actions, so gains of roughly ten to fifteen points may shrink or vary on larger, more diverse datasets.","Beyond the paper, a testable extension is to run the same augmentation structure with an open-weight language model in place of gpt-3.5, which would show whether the benefit comes from the LLM's scale or from the augmentation structure itself.","Beyond the paper, the sentence-encoder matching step is a potential bottleneck; using a fixed output head over the 43 actions instead of cosine-similarity matching might change the measured benefit of augmentation."],"forward_implications":["If the pipeline works as reported, robot assistants can be adapted to new homes or action sets by generating scenario data instead of running thousands of human-robot interactions.","The combination of place-based and action-based augmentation outperforms either alone, so both diversity of locations and coverage of the action set matter.","The method's gains concentrate on labels that were at zero accuracy before augmentation, meaning the synthetic data helps the hardest action categories most.","Action-based augmentation is more effective than place-based augmentation, suggesting that action-specific data is the stronger training signal.","The ablation shows that reference-image conditioning contributes to the gain, so image realism and consistency matter beyond the text alone."],"supporting_citations":[{"why":"Supplies the Do-I-Demand dataset, the real-world benchmark with 400 samples and 43 actions that defines the task and evaluation.","marker":"[33]"},{"why":"Supplies LLaVA, the base multimodal model that is fine-tuned on the augmented data and evaluated.","marker":"[17]"},{"why":"Supplies BLIP-Diffusion, the model used to generate action-based images conditioned on a reference image, whose removal in the ablation lowers accuracy.","marker":"[16]"},{"why":"Supplies SDXL, the diffusion model used for place-based image generation and as the substitute in the ablation that shows BLIP-Diffusion's importance.","marker":"[25]"},{"why":"Supplies LoRA, the efficient fine-tuning technique used to adapt LLaVA to the augmented scenario data.","marker":"[10]"},{"why":"Supplies Sentence-BERT, one of the encoders used to match LLaVA's response text to the action labels during evaluation.","marker":"[30]"},{"why":"Supplies GPT-3, the alternative sentence encoder used for the same response-to-action matching.","marker":"[3]"}],"fun_headline_variants":["Synthetic data from LLMs and diffusion lifts robot action accuracy to 48.8%","Robot action accuracy hits 48.8% thanks to synthetic dialogues and images","LLM-diffusion pairing crafts training data that boosts robot action accuracy to 48.8%","AI-generated scenes and dialogues teach robots to pick actions with 48.8% accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes the reference image used to condition BLIP-Diffusion is not taken from the evaluation portion of Do-I-Demand; if it is, the reported gains could be test-set leakage rather than generalization.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic data from LLMs and diffusion lifts robot action accuracy to 48.8%","Robot action accuracy hits 48.8% thanks to synthetic dialogues and images","LLM-diffusion pairing crafts training data that boosts robot action accuracy to 48.8%","AI-generated scenes and dialogues teach robots to pick actions with 48.8% accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000933,"raw_usage":{"total_tokens":3924,"prompt_tokens":808,"completion_tokens":3116,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":424,"completion_tokens_details":{"reasoning_tokens":3023}},"tokens_in":424,"tokens_out":3116,"duration_ms":24714,"temperature":1.0,"reasoning_tokens":3023,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:24:43.094859+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check which image from the real-world data collection was used as the BLIP-Diffusion reference and which partition it belongs to. If that image is in the Do-I-Demand test set, retraining without it should be run, and if accuracy drops to baseline, the augmentation's reported gains are not evidence of generalization. A second check is to generate action-based images without any reference image and compare accuracy, since the paper's ablation only swaps BLIP-Diffusion for SDXL rather than removing the conditioning.","supporting_citations":[{"cited_title":"IEEE Access12, 11,774–11,784 (2024)","cited_arxiv_id":null,"evidence_quote":"Supplies the Do-I-Demand dataset, the real-world benchmark with 400 samples and 43 actions that defines the task and evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies LLaVA, the base multimodal model that is fine-tuned on the augmented data and evaluated."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies BLIP-Diffusion, the model used to generate action-based images conditioned on a reference image, whose removal in the ablation lowers accuracy."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies SDXL, the diffusion model used for place-based image generation and as the substitute in the ablation that shows BLIP-Diffusion's importance."},{"cited_title":"In: International Conference on Learning Repre- sentations (2021)","cited_arxiv_id":null,"evidence_quote":"Supplies LoRA, the efficient fine-tuning technique used to adapt LLaVA to the augmented scenario data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Sentence-BERT, one of the encoders used to match LLaVA's response text to the action labels during evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies GPT-3, the alternative sentence encoder used for the same response-to-action matching."}],"review_version":1}