{"id":"2da9273a-5d09-4918-925a-6a159bdf5353","arxiv_id":"2508.05366","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Self-feedback from LLMs improved, hurt, or left unchanged retrieval-augmented answers depending on model and question type in early BioASQ 2025 experiments.","lead":"This paper tested whether large language models improve their own answers in biomedical question answering by critiquing and revising their outputs. It reports that self-feedback works unevenly, depending on the model and the task.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's 'varied performance' may not actually compare self-feedback to a no-feedback baseline; without that comparison, the central claim is uninterpretable.","rationale":"The paper is abstract-only, so the reader's UNVERDICTED rating is appropriate. My stress-test identifies a more fundamental evidential gap than the reader's same-model-self-bias concern: the abstract's empirical statement may not even contain a comparison with a no-feedback condition. The sentence 'varied performance for the self-feedback strategy across models and tasks' is logically compatible with merely observing that different models/tasks have different absolute scores. If that is all the paper reports, then the central claim—that self-feedback has no universal beneficial effect—is not addressed at all. This is not an accusation of error; it is a precise statement of what would have to be present in the full text for the claim to be supported. Because no full text is available, the concern cannot be resolved here, and the verdict stays UNVERDICTED. I partially agree with the reader's weakest_assumption because both point to the need for a proper comparison, but the missing no-feedback control is the more operationally load-bearing issue.","tokens_in":798,"tokens_out":3195,"duration_ms":35902,"concrete_test":"Locate the BioASQ results table for the four models. For each model and answer type, compute Δ = score(self-feedback) − score(no-feedback control) for the same test questions, with a bootstrap or sign-test error estimate. If no control column or no paired no-feedback result is reported, the abstract's 'varied performance' does not test the effect of self-feedback. If Δs are available, check whether their signs and magnitudes actually vary across model-task pairs rather than being within noise; only then does 'varied performance' support the central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, as the reader articulates it, is that self-feedback has no universal beneficial effect across models and tasks. But the only evidence sentence in the abstract—'Preliminary results indicate varied performance for the self-feedback strategy across models and tasks'—is ambiguous. It may simply report that absolute performance varies across the four models and four answer types, which would be trivially true and says nothing about the effect of self-feedback. For the claim to be load-bearing, the paper must compare, per model and answer type, the self-feedback run against a matched run with no self-feedback (and ideally against an independent-critic or human-feedback baseline). Without that contrast, any 'variation' could be driven by model/task difficulty, answer-type scoring, prompt order, or random seed, not by whether the critique was useful. The reader's concern about same-model self-agreement is secondary; the first question is whether a no-feedback control exists and what the deltas are. This is not an internal inconsistency—the abstract is underspecified—but it is the point on which the empirical claim hinges. If the full text includes such a control, this concern dissolves.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates self-feedback in agentic retrieval-augmented generation for biomedical question answering, using the BioASQ CLEF 2025 benchmark. The authors compare four LLMs (Gemini-Flash 2.0, o3-mini, o4-mini, DeepSeek-R1) with and without a self-feedback loop in which the model generates an output, evaluates it, and refines it, for query expansion and four answer types. The abstract reports 'varied performance' of self-feedback across models and tasks and concludes with insights into LLM self-correction and the need to compare LLM-generated feedback with human expert input.","tokens_in":1082,"tokens_out":1506,"duration_ms":16460,"significance":"If substantiated, the claim that self-feedback has inconsistent or no universal benefit across modern reasoning and non-reasoning models on a professional biomedical search task is valuable: it would temper expectations for a widely used technique and support the case for human-in-the-loop validation. The multi-model, multi-answer-type design is well matched to the BioASQ setting and the study addresses a timely question. However, at the abstract level the evidence is qualitative only: there are no effect sizes, significance tests, error bars, dataset splits, or explicit comparison against a no-feedback baseline. The paper's contribution cannot be assessed from the provided text, and the central empirical claim currently rests on an underspecified result sentence.","major_comments":[{"comment":"The sentence 'Preliminary results indicate varied performance for the self-feedback strategy across models and tasks' is ambiguous: it could mean that absolute performance varies across model/task combinations, which would be trivially true and would not measure the effect of self-feedback. The central claim requires a per-model, per-answer-type contrast between a self-feedback run and a matched no-feedback run. Please report the deltas, the direction of variation, and the number of test questions per condition. Without this control, the results sentence does not support the stated investigation of whether iterative self-correction improves performance.","section":"Abstract, results sentence"},{"comment":"The self-feedback loop uses the same model for generation, evaluation, and refinement. Because the critic and generator share the same biases and training objective, observed changes may reflect the model's tendency to agree or disagree with its own initial output rather than a general property of critique. To interpret 'varied performance' as evidence about self-correction, the paper should include a baseline with an independent critic (e.g., a different LLM or human feedback) or at least report the content of the critiques and the rate at which the model accepted or rejected its own suggestions. Without such a contrast, the mechanism behind the variation is unidentified.","section":"Abstract, methodology sentence"},{"comment":"The abstract reports no quantitative results: no accuracy, F1, or other BioASQ score is given for any model or condition. Even for a short paper, a compact results table or effect-size summary is needed to support the empirical claim. As written, the reader cannot determine whether the 'varied performance' is statistically or practically meaningful, nor whether the paper's conclusions follow from the data.","section":"Abstract, overall"}],"minor_comments":[{"comment":"'Agentic Retrieval Augmented Generation (RAG) and deep research systems' is a somewhat loose characterization; consider specifying the system types and citing representative works to frame the novelty of self-feedback.","section":"Abstract, first sentence"},{"comment":"The claim that automated systems 'may reduce user involvement and misalign with expert information needs' would benefit from a concrete reference or definitional clarification, since the paper does not appear to measure user involvement directly.","section":"Abstract, motivation"},{"comment":"The phrase 'informs future work' is vague; please indicate what specific design or measurement insight the preliminary results provide, even at a qualitative level.","section":"Abstract, final sentence"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an abstract-only submission in this review, so the formal assessment is based on the abstract's claims. The missing no-feedback baseline and quantitative reporting are addressable in revision if the full paper contains the underlying data. If the full text also lacks such comparisons, the verdict should be reconsidered toward rejection, because the central claim would then be unsupported. The paper fits the venue's interest in LLM evaluation for biomedical search, but currently reads more like an extended abstract than a completed study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Candid take: the abstract is too thin to tell whether the central claim is actually supported, and the one sentence that carries the paper—'varied performance for the self-feedback strategy'—is ambiguous. It might mean self-feedback helped on some models/tasks and hurt on others, or it might just mean absolute performance varied across models and tasks, which is trivially true and says nothing about self-feedback. If the full paper doesn't compare each self-feedback run to a matched no-feedback baseline, the result is uninterpretable. If it does, the concern goes away.\n\nWhat's genuinely new here: the specific combination of current reasoning and non-reasoning LLMs (Gemini-Flash 2.0, o3-mini, o4-mini, DeepSeek-R1) with an iterative self-feedback loop on BioASQ 2025 answer types. That combination hasn't been measured before, and the field genuinely doesn't know whether self-correction helps professional search. Mixed or negative evidence is a useful data point. The design—generate, critique, refine—is standard but appropriate.\n\nThe soft spots are also clear. First, the missing no-feedback control is the load-bearing issue. Second, same-model self-critique is inherently circular; the critique isn't independent, so any effect could be the model agreeing or disagreeing with itself rather than a general property of self-critique. The abstract acknowledges the human-feedback comparison only as future work, which is fine but should be flagged as a limitation. Third, there are no effect sizes, significance tests, or error bars reported, so even the 'varied performance' cannot be judged for size or reliability. These are criticisms of the abstract; the full text may well address them.\n\nThis paper is for people working on LLM self-correction, agentic search, and BioASQ. It's not a breakthrough, but it's a legitimate empirical question. Send it to peer review—if the full paper includes a baseline comparison, it deserves a serious look. If it doesn't, the authors need a major revision before it can stand.","headline":"The abstract is too thin to know whether the self-feedback result is real, but the research question is legitimate and worth a peer-review look.","tokens_in":1511,"tokens_out":2860,"would_cite":false,"duration_ms":24404,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Self-feedback in LLMs helps some BioASQ tasks, hurts others.","keywords":["retrieval augmented generation","self-feedback","self-correction","biomedical question answering","BioASQ","reasoning models","query expansion","large language models"],"falsifier":"Run one model on a fixed set of BioASQ questions under two conditions: the real self-feedback loop, and a control loop where the critique step is replaced by a generic instruction of the same length (e.g., 'now revise your answer'). If the control produces the same distribution of score changes, the critique content is not driving the effect. Alternatively, compare self-feedback against human expert revisions; if self-feedback never beats the human baseline, the claim that models can critique themselves into better biomedical answers would fail.","tokens_in":737,"feed_emoji":"🔁","tokens_out":3653,"duration_ms":34299,"temperature":0.7,"pith_summary":"This paper asks whether an LLM can improve its own retrieval-augmented answers by generating, critiquing, and revising them. Testing four models on biomedical questions from the BioASQ CLEF 2025 challenge, it applies this self-feedback loop to query expansion and to yes/no, factoid, list, and ideal answer types. The paper's central claim, given the preliminary results, is that self-feedback has no consistent beneficial effect: performance varies across models and tasks. It also claims that reasoning models are not reliably better at producing useful feedback than non-reasoning models. A sympathetic reader would take away that self-correction is not a safe default in professional search and needs per-task validation.","feed_headline":"LLM self-feedback helps some tasks, hurts others","feed_subtitle":"Four models critiquing their own BioASQ answers show no consistent gain, so self-correction can't be a default.","key_machinery":"The central mechanism is the self-feedback loop: a single LLM acts as generator, evaluator, and refiner of its own output. The loop is applied at two stages—query expansion and final answer generation—and across four answer formats (yes/no, factoid, list, ideal). The mechanism carries the comparison because all models receive the same loop, so any difference in outcome is attributed to the model and the task type.","core_discovery":"On its own terms, the paper reports a comparative study of agentic retrieval augmented generation with a self-feedback mechanism. Each model—Gemini-Flash 2.0, o3-mini, o4-mini, and DeepSeek-R1—generated an initial output, evaluated that output, and then refined it, both for expanding queries and for producing four answer types in the BioASQ CLEF 2025 biomedical question-answering challenge. The discovery is negative in shape: preliminary results show varied performance, meaning the same self-feedback loop improves some model-task combinations and degrades others. The paper frames this as evidence that self-correction cannot be assumed to help, and that reasoning models do not consistently ge","pith_inferences":["If the preliminary pattern holds, part of the 'varied performance' could be an artifact of self-agreement bias: a model tends to rate its own first answer highly, so the refinement step may simply confirm the initial output unless the critique is strong enough to override it.","A stronger test would compare self-feedback with cross-model feedback—one model critiquing another's answer—and with human expert edits. If cross-model feedback outperforms self-feedback, the results would point to critique quality rather than self-correction as the active ingredient.","In a professional search setting, the paper's implied design could be extended into an interactive loop where model self-feedback is offered as an optional revision rather than applied automatically, preserving user expertise and transparency."],"forward_implications":["Self-feedback should not be treated as a free improvement in RAG pipelines; deployments need task-level evaluation before enabling it.","Reasoning-oriented LLMs should not be assumed to give better self-critique; feedback quality has to be checked separately from answer quality.","For professional search, expert-formulated benchmarks like BioASQ can expose where automated refinement conflicts with expert information needs.","If self-feedback is applied, its effect on transparency and user involvement should be reported, not just its effect on accuracy."],"supporting_citations":[],"fun_headline_variants":["Self-feedback in LLM RAG: no consistent gain across tasks","Why LLM self-critique can backfire in biomedical search","Self-correction: unreliable for LLM retrieval in bio QA","LLM self-feedback: hit or miss on BioASQ tasks","Reasoning models don't guarantee better self-feedback"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that a model's critique of its own output carries a meaningful correction signal; if the critique mostly echoes the model's initial biases or adds prompt noise, the varied performance says little about self-correction.","fun_headline_variants_meta":{"raw":{"variants":["Self-feedback in LLM RAG: no consistent gain across tasks","Why LLM self-critique can backfire in biomedical search","Self-correction: unreliable for LLM retrieval in bio QA","LLM self-feedback: hit or miss on BioASQ tasks","Reasoning models don't guarantee better self-feedback"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000207,"raw_usage":{"total_tokens":1258,"prompt_tokens":784,"completion_tokens":474,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":385}},"tokens_in":528,"tokens_out":474,"duration_ms":4280,"temperature":1.0,"reasoning_tokens":385,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:22:13.828587+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run one model on a fixed set of BioASQ questions under two conditions: the real self-feedback loop, and a control loop where the critique step is replaced by a generic instruction of the same length (e.g., 'now revise your answer'). If the control produces the same distribution of score changes, the critique content is not driving the effect. Alternatively, compare self-feedback against human expert revisions; if self-feedback never beats the human baseline, the claim that models can critique themselves into better biomedical answers would fail.","supporting_citations":[],"review_version":1}