{"id":"d599de22-17d1-48c5-9da1-4245bd65ceed","arxiv_id":"2508.15283","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":7.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper claims a few-shot prompting attack on neural ranking models, but the body is a mismatch: it contains no methods or experiments for that claim.","lead":"The abstract describes FSAP, a black-box attack that uses few-shot prompting of LLMs to generate adversarial documents that outrank credible results in neural ranking models. The manuscript body, however, is an unrelated astronomy paper on galaxy background subtraction, so the claimed experiments and framework are not present in the submitted text.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's FSAP claims are unsupported by the manuscript body, which is an unrelated astronomy paper (HostSub GP, arXiv:2508.15278v2); no methods or results for FSAP appear in the submission.","rationale":"The reader's verdict identifies the abstract/full-text mismatch as the central problem, and my independent read of the submission agrees: the body is an astronomy paper with no connection to the FSAP abstract. The most load-bearing concern is therefore not a subtle flaw in the attack method—there is no method present to scrutinize—but the complete absence of any supporting content for the empirical claims. This concern directly undermines the strongest claim in the abstract, since 'consistently outrank' and 'low detectability' are empirical assertions requiring experiments that do not appear anywhere in the manuscript. The reader's weakest_assumption focuses on the support-set dependency, which would be relevant if the FSAP methodology were actually described; that is a secondary issue. I therefore mark agreement as 'partial': the reader's rationale flags the same mismatch, but their stated weakest assumption is downstream of the more fundamental evidentiary gap. No ad hominem is intended; the issue is the text itself, not the authors. A REJECT verdict is appropriate because the submission cannot be evaluated as a coherent research paper on its stated topic, and the central claim is unsupported by any derivable evidence here.","tokens_in":12377,"tokens_out":2360,"duration_ms":26634,"concrete_test":"Download the full LaTeX/PDF source of arXiv:2508.15283 and search the body for 'FSAP', 'Few-Shot Adversarial Prompting', 'TREC', 'neural ranking', or 'misinformation'. If no such terms appear outside the abstract, the central claim has no in-text support. Additionally, compare the body's title, abstract, and arXiv header with arXiv:2508.15278v2; if they match, the submission is the astronomy paper under a different arXiv ID, confirming the mismatch.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim is that FSAP-generated documents consistently outrank credible documents on TREC 2020/2021 Health Misinformation Tracks across four neural ranking models, with strong stance alignment and low detectability. For that claim to be supported, the submission must contain the FSAP method, the experimental setup, the support-set construction, and the ranking results. The full text contains none of these: it is a self-contained astronomy paper titled 'HostSub GP: Precise Galaxy Background Subtraction in Transient Long-slit Spectroscopy with Gaussian Processes', whose arXiv header identifies it as 2508.15278v2 rather than the submitted 2508.15283. There is no section defining few-shot prompting, no description of the support set, no experimental protocol, no tables of TREC results, and no analysis of stance alignment or detectability. Thus the abstract's empirical assertions are without derivation or verification in this submission. This is not a matter of contested methodology; it is a complete absence of the evidentiary body required to assess the claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission is presented under the title 'Adversarial Attacks against Neural Ranking Models via In-Context Learning' and its abstract claims that a Few-Shot Adversarial Prompting (FSAP) framework generates documents that 'consistently outrank credible, factually accurate documents' on the TREC 2020 and 2021 Health Misinformation Tracks across four neural ranking models, with strong stance alignment and low detectability. However, the full text of the submission is an entirely unrelated astronomy paper, 'HostSub GP: Precise Galaxy Background Subtraction in Transient Long-slit Spectroscopy with Gaussian Processes'. None of the sections, equations, figures, tables, or references in the body concerns FSAP, neural ranking models, TREC, LLMs, or adversarial attacks. No experimental setup, support-set construction, model list, evaluation protocol, or result tables supporting the abstract's claims appear anywhere in the manuscript.","tokens_in":12586,"tokens_out":2804,"duration_ms":32893,"significance":"If the abstract's claims were supported, the proposed FSAP attack would be a significant contribution to adversarial IR, as it would demonstrate a scalable, black-box, in-context-learning-based threat to neural ranking systems, with implications for misinformation and retrieval security. However, the submitted manuscript contains none of the evidence needed to assess these claims. There is no method section, no dataset description, no experimental protocol, no code release, and no falsifiable result. The astronomy content in the body, while possibly of value in its own field, is irrelevant to the advertised topic. The paper cannot be evaluated as a research contribution because its stated subject and its actual content are disjoint.","major_comments":[{"comment":"The full text is a self-contained astronomy paper ('HostSub GP') with no mention of FSAP, neural ranking models, TREC, LLMs, or adversarial attacks. The abstract's central empirical claim—that FSAP-generated documents consistently outrank credible documents across four ranking models—is therefore completely unsupported. This is not a local gap in methodology or a missing robustness check; the evidentiary body for the claimed contribution is absent.","section":"Entire manuscript (Sections 1–6)"},{"comment":"The submitted title and abstract describe an adversarial-ranking paper, but the full-text header identifies it as arXiv:2508.15278v2, titled 'HostSub GP: Precise Galaxy Background Subtraction...', with an entirely different author list and subject. This identity mismatch means the text cannot be verified as the paper described by the abstract. At minimum, the submission must be accompanied by the correct, matching full text before any substantive review can occur.","section":"arXiv header and title"},{"comment":"Even if one attempted to treat the abstract as a standalone claim, there is no description of the support set, its size or construction, no list of the four ranking models, no definition of the TREC evaluation measures, no baseline comparisons, and no analysis of stance alignment or detectability. The assertions of 'consistently outrank' and 'low detectability' are not accompanied by any data, tables, or statistical tests.","section":"Experimental protocol (missing)"},{"comment":"The stress-test concern that FSAP may depend on a representative support set remains unaddressed, but this is secondary: the manuscript contains no FSAP method at all. The only limitations discussed (Section 5 of the astronomy text) concern Gaussian-process host subtraction, not the claimed adversarial attack. A limitations discussion for FSAP, including support-set requirements and topic-coverage constraints, is absent.","section":"Limitations and failure modes"}],"minor_comments":[{"comment":"The title in the PDF body does not match the submission title. The body's header lists 'HostSub GP' and an astronomy abstract, while the submission metadata lists FSAP. The author list also appears different. This is not a simple typo and should be corrected at the submission level.","section":"Title and abstract"},{"comment":"The reference list contains only astronomy-related citations. No prior work on adversarial IR, in-context learning, or TREC misinformation tracks is cited, further confirming that the body is not the paper described in the abstract.","section":"References"}],"recommendation":"reject","confidential_remarks":"This manuscript appears to be a submission error: the abstract and the full text are two entirely different papers. The editor may wish to verify the uploaded files and submission metadata. If this is an accidental upload, a corrected version should be submitted as a new manuscript with the correct full text; as submitted, the paper cannot be reviewed or accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this one is two papers stapled together. The abstract promises a few-shot adversarial attack on neural ranking models (FSAP) with experiments on TREC health misinformation tracks. The full-text body is an astronomy paper on Gaussian-process host galaxy subtraction (arXiv:2508.15278v2). None of the FSAP method, the support-set construction, the experimental setup, or the ranking results appears anywhere in the submission. The mismatch is mechanical but it is decisive: there is no evidentiary support for the abstract's claims in the submitted text.\n\nWhat is new: the idea of using in-context learning with a small support set of harmful examples to generate whole adversarial documents, rather than token-level perturbations or manual rewrites, is a plausible and potentially novel attack vector. The two instantiations (same-query and cross-query transfer) and the reported generalization across proprietary and open-source LLMs are worth reading about. But this submission does not contain that work. The HostSub GP paper may be a legitimate contribution to transient spectroscopy—it has a clear method, synthetic and real-data tests, and released software—but it is not the paper the abstract describes.\n\nSoft spots are not subtle: the central claim has no derivation or verification in the body. The reader's low confidence is appropriate; we cannot assess FSAP's soundness because none of the apparatus is present. The astronomy paper might be internally sound, but that does not rescue a submission whose abstract and body are different works. If this is an upload error, the authors should resubmit the correct PDF; if it is not, a referee cannot review what is here.\n\nFor peer review: desk reject. There is no coherent manuscript to evaluate. I would not bring it to reading group and would not cite it. The abstract's idea is interesting, but an abstract alone is not a paper.","headline":"The submission is two papers stuck together: the abstract promises a few-shot adversarial attack on neural rankers, while the body is a GP-based galaxy subtraction paper; neither supports the other.","tokens_in":13065,"tokens_out":2505,"would_cite":false,"duration_ms":22757,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A pure prompting attack makes generated documents outrank credible health answers","keywords":["adversarial attack","neural ranking models","in-context learning","large language models","few-shot prompting","health misinformation","black-box attack","information retrieval"],"falsifier":"Run FSAP against the same TREC 2020/2021 queries with support sets drawn from unrelated topics or with factual rather than harmful examples; if the generated documents no longer outrank credible answers, the effect depends on topic-matched harmful content rather than on the prompting mechanism itself. A second check would measure whether detected-document flagging rises when the generated set is clustered by style.","tokens_in":12284,"feed_emoji":"⚔️","tokens_out":5041,"duration_ms":59201,"temperature":0.7,"pith_summary":"This paper tries to establish that neural ranking models can be defeated without any access to their internals: an attacker simply prompts a large language model with a small set of previously observed harmful documents and asks it to write a new document for a target query. The resulting text is fluent and topically plausible, and on the TREC 2020 and 2021 Health Misinformation Tracks it outranks credible, factually accurate documents across four ranking models. The attack works in two modes, one that reuses harmful examples from the same query and one that transfers patterns across unrelated queries. If true, this matters because it turns document ranking—a core component of search and question answering—into a target that can be manipulated through ordinary text generation, with no gradients or model instrumentation.","feed_headline":"Few-shot prompts make fake health documents outrank real answers","feed_subtitle":"Black-box LLM attack conditions on harmful examples to beat credible health info in TREC tests.","key_machinery":"The central object is the Few-Shot Adversarial Prompting (FSAP) framework and its two instantiations, FSAP-IntraQ and FSAP-InterQ. The load-bearing mechanism is in-context learning: the support set of harmful examples shapes the LLM's generation so that fluency and topical coherence are preserved while misleading content is embedded. This replaces token-level gradient attacks and manual rewriting with a pure prompt-level attack that does not require any gradient access or internal model instrumentation.","core_discovery":"FSAP treats adversarial generation as an in-context learning problem rather than a search over token perturbations. Given a query and a support set of harmful documents, the LLM produces a grammatically fluent, topically coherent document that embeds false or misleading claims. FSAP-IntraQ uses harmful examples from the same query to maximize topical fidelity; FSAP-InterQ transfers adversarial patterns from unrelated queries to broaden coverage. On the TREC 2020 and 2021 Health Misinformation Tracks, documents generated this way consistently rank above credible documents for four neural ranking models, show strong stance alignment with the misinformation topic, and are not easily detected as","pith_inferences":["The paper evaluates on health misinformation; the same prompting mechanism could plausibly transfer to other domains where ranking decides what is seen, such as news, product reviews, or code, though the abstract does not test this.","A natural next experiment would vary support-set size, topical distance, and example ordering to map where the rank advantage appears, since those degrees of freedom are not analyzed in the abstract.","If an LLM can write documents that outrank credible ones, downstream systems that use top-ranked results as training labels could inherit a systematic misinformation bias."],"forward_implications":["Any deployed neural ranker that admits LLM-generated documents into its candidate pool is exposed to a black-box attack an ordinary API user could run.","FSAP-InterQ's transfer across unrelated queries implies the attack is not confined to a few memorized queries; a small corpus of harmful examples may seed a much wider set of attacks.","If generated documents outrank credible ones on health misinformation topics, rankers built on transformer encoders are not robust to in-context adversarial text, contradicting the assumption that fluency and topicality alone indicate trustworthiness.","The reported low detectability means simple filter-based defenses are unlikely to stop the attack without more sophisticated content-verification signals.","Because the method requires no gradient access, it also applies to proprietary rankers whose internal parameters are hidden."],"supporting_citations":[],"fun_headline_variants":["Few-shot prompts let LLMs craft fake health docs that beat real ones","Black-box attack uses few-shot examples to outrank truthful health content","In-context learning powers a prompt-only attack on neural rankers","LLMs mislead search by generating fluent fake docs from harmful examples","No gradients needed: prompt-based attack fools neural ranking models"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The attack's effectiveness hinges on having a support set of harmful examples that are sufficiently representative of the target query for the LLM to imitate; the abstract does not specify how large or how closely matched that set must be for the reported outranking to occur.","fun_headline_variants_meta":{"raw":{"variants":["Few-shot prompts let LLMs craft fake health docs that beat real ones","Black-box attack uses few-shot examples to outrank truthful health content","In-context learning powers a prompt-only attack on neural rankers","LLMs mislead search by generating fluent fake docs from harmful examples","No gradients needed: prompt-based attack fools neural ranking models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001119,"raw_usage":{"total_tokens":4507,"prompt_tokens":767,"completion_tokens":3740,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":3650}},"tokens_in":511,"tokens_out":3740,"duration_ms":24230,"temperature":1.0,"reasoning_tokens":3650,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:59:04.726082+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run FSAP against the same TREC 2020/2021 queries with support sets drawn from unrelated topics or with factual rather than harmful examples; if the generated documents no longer outrank credible answers, the effect depends on topic-matched harmful content rather than on the prompting mechanism itself. A second check would measure whether detected-document flagging rises when the generated set is clustered by style.","supporting_citations":[],"review_version":1}