{"id":"92f411f0-45aa-4957-bb42-d8d914d79bd5","arxiv_id":"2508.16261","paper_version":1,"verdict":"UNVERDICTED","confidence":"UNKNOWN","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The abstract promises a survey and taxonomy of federated LLM post-training methods, but the full text is an unrelated ANN filtering benchmark paper.","lead":"This submission's abstract describes a survey of federated post-training methods for large language models, organized by whether the model is white-box, gray-box, or black-box to the trainer. The full text supplied is a different paper, an experimental benchmark on attribute-filtered nearest-neighbor search, so the survey described cannot be reviewed.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Manuscript body is a different paper ('Attribute Filtering in Approximate Nearest Neighbor Search', arXiv:2508.16263v2); the claimed FedLLM survey/taxonomy has no supporting text in this submission, so the central claim is unverifiable.","rationale":"The reader's strongest_claim is the abstract's survey claim, and the reader's weakest_assumption rightly identified that the submitted manuscript may not contain the described survey. My stress-test confirms this directly: the full text is a different paper with a different title, author list, arXiv identifier, and research topic. No section or equation of the FedLLM survey exists in the supplied body, so none of the survey's assertions about taxonomy coverage, representative methods, or open challenges can be checked. This is not a disagreement over interpretation or completeness within a bounded survey; it is a total absence of the artifact that would need to be reviewed. Because the central claim cannot be read, let alone audited, the correct disposition is to leave the reader's UNVERDICTED verdict unchanged. I also note that this observed inconsistency could arise from an upload or metadata mix-up rather than any intent to mislead; the concern is about the submission's verifiability, not about the authors' integrity.","tokens_in":4739,"tokens_out":2586,"duration_ms":28153,"concrete_test":"Fetch the arXiv record for 2508.16261 directly from arXiv (e.g., https://arxiv.org/abs/2508.16261) and its PDF. Verify that title, author list, and abstract match the claimed FedLLM survey. Then run full-text searches in the PDF body for 'federated', 'post-training', 'white-box', 'gray-box', 'black-box', and 'taxonomy'. If any of these terms are absent, or if the body is the ANN-filtering paper, the submission cannot support the survey's claims. If a corrected PDF is supplied and does contain the survey, proceed to audit its search methodology and completeness.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract promises a comprehensive survey of federated post-training of LLMs with a white/gray/black-box taxonomy. But the supplied full text is entirely the paper 'Attribute Filtering in Approximate Nearest Neighbor Search: An In-depth Experimental Study' by Li, Yan, Lu, Zhang, Cheng, and Ma, which carries arXiv:2508.16263v2 [cs.DB]. Title, authors, and subject matter differ from the abstract; there is no body section defining the FedLLM taxonomy, no method families, no search methodology, and no completeness analysis. Since no survey content is present, the load-bearing conditions for the central claim—that the taxonomy accurately organizes existing FedLLM studies and that the black-box inference-only paradigm is adequately covered—cannot be checked. This is a missing-support condition, not a disagreement with a scientific consensus. It is not resolved by plausibility of the abstract; the submitted artifact itself is internally inconsistent.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission claims to be a comprehensive survey of federated post-training of large language models (FedLLM), proposing a taxonomy based on model access (white-box, gray-box, black-box) and parameter-efficiency optimization. However, the supplied full text is a different paper: 'Attribute Filtering in Approximate Nearest Neighbor Search: An In-depth Experimental Study' by Mocheng Li, Xiao Yan, Baotong Lu, Yue Zhang, James Cheng, and Chenhao Ma (arXiv:2508.16263v2, cs.DB). The body contains no FedLLM taxonomy, no representative FedLLM methods, no survey tables, and no discussion of the black-box inference-only paradigm. The central claim of the abstract is therefore unverifiable from the submitted artifact.","tokens_in":4832,"tokens_out":2222,"duration_ms":28051,"significance":"If the intended survey existed as described, it would be a potentially useful organizing contribution: a validated white/gray/black-box taxonomy for FedLLM could provide a common vocabulary for a rapidly growing literature and give visibility to inference-only approaches. However, in this submission there is nothing to evaluate. The actual body text is an unrelated ANN filtering survey with a different title, author list, and subject matter. No credit can be given for content that is absent, and the claimed contribution is not merely flawed but missing.","major_comments":[{"comment":"The abstract promises a comprehensive survey on federated tuning for LLMs and a white-box/gray-box/black-box taxonomy. The full text is a different paper: 'Attribute Filtering in Approximate Nearest Neighbor Search: An In-depth Experimental Study' by Li et al., arXiv:2508.16263v2 [cs.DB]. The body contains no FedLLM content whatsoever, so the load-bearing premise of the submission is false.","section":"Abstract vs. Full Text"},{"comment":"No section of the body defines or even mentions the FedLLM taxonomy claimed in the abstract. The only taxonomy presented concerns Filtering ANN algorithms based on attribute types and filtering strategies. The claimed two-axis classification of FedLLM studies cannot be checked for correctness, completeness, or internal consistency because it does not exist in the manuscript.","section":"Body text (all sections)"},{"comment":"Even taking the abstract at face value, no search strategy, inclusion criteria, or coverage analysis is provided for the claimed FedLLM survey. Comprehensiveness is a defining requirement for a survey, and the abstract alone cannot establish it. This is an additional missing-support issue, secondary to the full-text mismatch but relevant if the authors resubmit the intended paper.","section":"Survey methodology (claimed but absent)"}],"minor_comments":[{"comment":"The full text carries an ACM Reference Format block with a placeholder DOI (https://doi.org/XXXXXXX.XXXXXXX) and a ©2018 copyright line, which is inconsistent with the 2025 submission date. Such metadata must be corrected in any resubmission.","section":"Metadata and formatting"}],"recommendation":"reject","confidential_remarks":"This is not a standard content rejection: the submitted manuscript body is an entirely different paper. I recommend that the editor desk-reject rather than send for full review, since no substantive evaluation of the claimed FedLLM survey is possible. The authors should be asked to resubmit the correct manuscript if the intended survey exists."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the punchline: the submission is not the paper it claims to be. The abstract promises a comprehensive survey of federated post-tuning of LLMs, with a two-axis taxonomy and white/gray/black-box classes; the full text is a different paper, \"Attribute Filtering in Approximate Nearest Neighbor Search,\" by a different author team, carrying a different arXiv ID. The body contains no section defining the FedLLM taxonomy, no representative method families, no search methodology, and no completeness analysis. The central claim has no supporting text.\n\nI can't review the survey described in the abstract because it doesn't exist in this artifact. The abstract's framing—model-access and parameter-efficiency axes—is a reasonable way to organize the FedLLM literature, and a good survey on those lines would be genuinely useful. But an abstract alone can't support judgments about coverage, accuracy, or novelty. The reader's UNVERDICTED is the correct verdict.\n\nCredit where it's due: the body text, read on its own, is a real experimental study of filtered ANN search. It proposes a unified interface and taxonomy, benchmarks 10 algorithms across 4 datasets, and ships code in a public repository. That is reproducible work. But it belongs to a different submission under its own title; as evidence for the FedLLM survey, it's irrelevant.\n\nThe mismatch is not a subtle argumentative flaw. It's a desk-reject-level internal inconsistency: the abstract and body are different papers. This is a missing-support condition, not a disagreement with the authors' conclusions. No referee can evaluate the claimed taxonomy from this material. I would not send this to peer review. The authors should resubmit with the correct full text (or submit the ANN paper properly). If the FedLLM survey actually exists and is as comprehensive as advertised, it would deserve a serious referee—but that is a conditional, not a property of this submission.\n\nBottom line: desk reject with an invitation to resubmit a corrected version. The abstract's taxonomy is worth reading if it ever materializes; this artifact is not.","headline":"The submission is not the paper it claims to be: the abstract describes a FedLLM survey, but the full text is a different paper on filtered ANN search, so there is nothing to review.","tokens_in":5436,"tokens_out":4128,"would_cite":false,"duration_ms":42802,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The full text delivered with this submission is not the federated-LLM survey announced in the abstract; it is an experimental study that proposes a unified interface and taxonomy for attribute-filtered approximate nearest-neighbor search.","keywords":["federated learning","large language models","approximate nearest neighbor search","attribute filtering","vector database","benchmark","taxonomy"],"falsifier":"Compare the full text's title and content with the abstract: the mismatch is already visible without any experiment. For the benchmark itself, a concrete falsifier would be adding a filtering-ANN method whose filtering strategy falls outside the proposed taxonomy and showing that it wins on the same four datasets, or showing that the relative ranking of the ten algorithms changes when a different 10-million-item dataset is used.","tokens_in":4515,"feed_emoji":"🔍","tokens_out":9805,"duration_ms":102492,"temperature":0.7,"pith_summary":"The submitted document consists of two mismatched parts. The abstract announces a taxonomy of federated post-training methods for large language models, grouped by whether they use white-box, gray-box, or black-box access to the model. The full text is a different paper: it builds a unified interface for attribute-filtered approximate nearest-neighbor search, classifies algorithms by attribute type and filtering strategy, and benchmarks 10 algorithms and 12 methods on four datasets holding up to 10 million items. The contribution that can actually be checked is therefore the filtering-ANN study, not the FedLLM survey. That study matters because vector search with structured attribute constraints underpins retrieval-augmented generation, recommendation, and vector databases.","feed_headline":"The promised FedLLM survey is not in the submission","feed_subtitle":"Full text instead compares 10 filtering-ANN algorithms and 12 variants on 10-million-item datasets.","key_machinery":"The central objects are the unified Filtering ANN search interface and the two-axis taxonomy. The interface fixes the input/output semantics so different algorithms can be compared head-to-head; the taxonomy organizes the field by attribute types (for example, range predicates versus categorical predicates) and by when filtering happens relative to the search. A component-level analysis of index structures, pruning strategies, and entry-point selection then explains why methods differ in speed and recall on the same workload.","core_discovery":"The full-text authors claim that the scattered filtering-ANN methods of the last few years can be brought under one interface and one taxonomy. Their taxonomy is built on two axes—the attribute types being filtered and the filtering strategy used—and they identify index structure, pruning strategy, and entry-point selection as the components that separate methods and explain their tradeoffs. On four datasets with up to 10 million items and selectivity levels from 0.1% to 100%, the authors measure how each component affects efficiency and quality and condense the results into practical method-selection guidelines. None of this appears in the abstract, which describes an unrelated survey.","pith_inferences":["If the intended paper is the FedLLM survey named in the abstract, no assessment of that survey's claims is possible from this submission; the abstract and full text would need to be reconciled before its white-box/gray-box/black-box taxonomy can be taken seriously.","The filtering-ANN benchmark's practical conclusions are likely to transfer to workloads with similar attribute distributions; testing the same ten algorithms on datasets with mixed predicate types or very low selectivity would be a direct extension.","The same unified-interface approach could be extended to learned or hybrid indexes, but the paper gives no evidence that its taxonomy already includes them; a reader should not assume coverage beyond what is listed.","A mismatch of this kind, if not resolved, means the public record will carry two different claims under one entry; resolving the discrepancy is a prerequisite for any downstream reader who wants to cite either the survey or the benchmark."],"forward_implications":["Any new filtering-ANN index can be positioned against ten existing algorithms and twelve method variants under a common interface, making method comparisons replicable.","Because the taxonomy separates attribute type from filtering strategy, system builders can match a workload's predicate shape to the algorithm family most likely to handle it.","The component analysis points to pruning and entry-point selection as the main levers of cost; tuning those components is thus the most direct route to speedups.","The practical guidelines give vector-database and RAG developers a starting choice set for selectivity regimes between 0.1% and 100%, rather than relying on unbenchmarked defaults."],"supporting_citations":[{"why":"supplies the hierarchical navigable small-world graph structure that many compared graph methods build on.","marker":"[41]"},{"why":"provides the large-scale graph-based index used as a baseline and design reference.","marker":"[31]"},{"why":"defines product quantization, the quantization-based index family included in the comparison.","marker":"[32]"},{"why":"is the earlier graph-based ANN survey and experimental comparison this work extends with attribute filtering.","marker":"[60]"},{"why":"contributes a predicate-agnostic filtering approach evaluated as one of the benchmark's methods.","marker":"[50]"},{"why":"provides a range-dedicated graph method that the taxonomy must cover.","marker":"[62]"},{"why":"contributes a segment-graph filtering strategy for range-constrained search.","marker":"[67]"},{"why":"proposes a unified index for range-filtered approximate nearest neighbors, giving the closest prior form of the paper's own unified interface.","marker":"[38]"}],"fun_headline_variants":["FedLLM survey missing, filtering-ANN taxonomy delivered","Paper promises FedLLM survey, delivers ANN filtering benchmarks","Abstract claims FedLLM survey, real content: filtering-ANN taxonomy","Mismatch: FedLLM survey promise, actual paper on ANN filtering","The promised FedLLM survey is absent; here's a filtering-ANN taxonomy"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The announced survey has no body to stand on, and the benchmark actually present assumes that its ten algorithms and four datasets span the filtering-ANN design space well enough for the taxonomy and guidelines to generalize.","fun_headline_variants_meta":{"raw":{"variants":["FedLLM survey missing, filtering-ANN taxonomy delivered","Paper promises FedLLM survey, delivers ANN filtering benchmarks","Abstract claims FedLLM survey, real content: filtering-ANN taxonomy","Mismatch: FedLLM survey promise, actual paper on ANN filtering","The promised FedLLM survey is absent; here's a filtering-ANN taxonomy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000501,"raw_usage":{"total_tokens":2246,"prompt_tokens":659,"completion_tokens":1587,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":403,"completion_tokens_details":{"reasoning_tokens":1494}},"tokens_in":403,"tokens_out":1587,"duration_ms":13148,"temperature":1.0,"reasoning_tokens":1494,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:24:30.985562+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the full text's title and content with the abstract: the mismatch is already visible without any experiment. For the benchmark itself, a concrete falsifier would be adding a filtering-ANN method whose filtering strategy falls outside the proposed taxonomy and showing that it wins on the same four datasets, or showing that the relative ranking of the ten algorithms changes when a different 10-million-item dataset is used.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the hierarchical navigable small-world graph structure that many compared graph methods build on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the large-scale graph-based index used as a baseline and design reference."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines product quantization, the quantization-based index family included in the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"is the earlier graph-based ANN survey and experimental comparison this work extends with attribute filtering."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"contributes a predicate-agnostic filtering approach evaluated as one of the benchmark's methods."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"contributes a segment-graph filtering strategy for range-constrained search."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"proposes a unified index for range-filtered approximate nearest neighbors, giving the closest prior form of the paper's own unified interface."}],"review_version":1}