{"id":"1100978f-f9ea-42c9-ad39-1767d39f9f58","arxiv_id":"2505.07891","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"A GPT-4-based fact-checking system with graph retrieval reports 88.5% binary accuracy on PolitiFact health claims, but the improvement over GPT-4 is not shown to come from the graph component.","lead":"TrumorGPT combines GPT-4 with semantic health knowledge graphs and graph-based retrieval to fact-check health claims, reporting 88.5% binary accuracy on 600 PolitiFact statements. A generalist reader might look at this as a test of whether retrieval-augmented LLMs can be trusted for public-health verification, but the paper does not release code or show an ablation that isolates the retrieval component.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported 88.5% accuracy gain is attributed to GraphRAG, but the paper provides no ablation, no threshold θ for its Equation (5) retrieval rule, and no code or data; the same GPT-4 builds the graphs and then reads them, so the causal role of graph retrieval is unverified.","rationale":"The reader's weakest_assumption is correct and my concern is the same one, sharpened: the paper claims a mechanism but never varies it. I checked the manuscript text for any ablation, sensitivity analysis, or retrieval log. Section IV-C gives only aggregate metrics in Table I; Section IV-D evaluates graph size and error on 50 altered articles, but no condition removes or corrupts the retrieved graphs. Section III-E is the entire retrieval algorithm and it stops at Eq. (5) without values or implementation. Because GPT-4 constructs the graphs from the same claim or article and then consumes them, the pipeline is inherently confounded: any graph built to support the conclusion can be 'retrieved' with high similarity. The sign of the threshold in III-E, if taken literally, would predict the opposite of the intended behavior, which suggests the method description was not validated against code. The paper also gives no standard errors on Table I, but a 5.2-point gap over 600 balanced items has standard error roughly 1.3 points; even so, the absence of ablation means the gap cannot be assigned to GraphRAG. I therefore agree with the REJECT verdict; the concrete test would resolve the attribution with one controlled experiment, and the reported 88.5% would survive only if disabling or randomizing retrieval degrades accuracy.","tokens_in":20222,"tokens_out":4243,"duration_ms":44660,"concrete_test":"Reproduce Table I on the same 600 statements under three conditions: (A) TrumorGPT as described, with θ and per-claim retrieved graph IDs logged; (B) the same pipeline with retrieval disabled, i.e., plain GPT-4 with the few-shot graph-construction prompt but no graph database; (C) the same pipeline with each query paired with a random, unrelated knowledge graph from the DBpedia health subset. If accuracy in condition B or C remains within the 95% confidence interval of the reported 88.5%, the GraphRAG mechanism is not responsible for the claimed advantage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section IV-C's headline result (88.5% vs. 83.3% for GPT-4) is presented as evidence that GraphRAG over semantic health knowledge graphs improves fact-checking. The load-bearing assumption is that retrieved graphs, not GPT-4's parametric knowledge or the PolitiFact article text, produce the correct answer. Nothing in the paper tests this. Section III-E defines the similarity score in Eq. (5) but never reports the threshold θ, never specifies the graph embeddings or subgraph matcher used in practice, and never shows a single retrieved graph for the 600 test claims; the conditional rule 'S(Gx,Gi) ≤ θ ⇒ True' is also inconsistent with Jaccard similarity, where low overlap should not indicate truth. Section IV-A describes filtering DBpedia triples but not how the triples enter prompts or retrieval. The knowledge graphs themselves are built by GPT-4 with Advanced Data Analysis and few-shot examples, and the final verdict is also produced by GPT-4, so the evidence source and the reasoner are the same model; high accuracy could therefore reflect parametric memory or common-sense reasoning rather than graph-based retrieval. The reported six-way accuracy of 49.3% and the stated tendency to fall back to 'Half True' or 'Mostly False' reinforce that the binary numbers may be an artifact of collapsing categories rather than of graph content. This is not a minor reproducibility gap: if retrieval is not causal, the paper's central contribution reduces to a GPT-4 prompting wrapper, and the comparison against baseline LLMs is uninterpretable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes TrumorGPT, a fact-checking framework in which GPT-4 constructs semantic health knowledge graphs from input documents using topic-enhanced sentence centrality and topic-specific TextRank, and then uses graph-based retrieval-augmented generation (GraphRAG) over a DBpedia-derived graph database to verify health-related claims. The authors report 88.5% binary accuracy on 600 PolitiFact \"Health Care\" and \"Coronavirus\" statements, compared with 72.7–83.3% for six LLM baselines, and they also report 49.3% accuracy on the original six-category PolitiFact scale. The manuscript includes formal convergence statements for the proposed topic-specific TextRank and a discussion of related RAG methods.","tokens_in":20568,"tokens_out":5250,"duration_ms":58344,"significance":"If the central claim were established, a 5.2-point improvement over GPT-4 with shorter responses could be practically valuable for automated fact-checking, and the convergence analysis of the modified TextRank would be a useful supporting contribution. The paper also addresses a timely problem and compares several contemporary LLMs. However, the current evidence does not support the central claim: the evaluation cannot separate the effect of graph retrieval from the underlying GPT-4 model, key components of the retrieval rule are unspecified, and the graph construction pipeline can encode the answer into the very graph that is later used for reasoning. At best, this is a preliminary proof-of-concept, not a validated contribution as stated.","major_comments":[{"comment":"The central claim that GraphRAG over semantic health knowledge graphs causes the accuracy gain is not supported by the experiments. TrumorGPT is compared with raw LLM baselines, but there is no ablation that removes the graph retrieval component, no prompt-matched baseline with the same few-shot instructions, and no report of error bars or significance tests. The 5.2-point gap over GPT-4 could be due to the added instructions, the extra DBpedia data, or the binary label construction rather than to the graph retrieval mechanism. An ablation such as TrumorGPT without retrieval, or GPT-4 with the same external triples in plain text, is essential before Table I can be interpreted.","section":"Section IV-C, Table I"},{"comment":"The retrieval rule is underspecified and the stated condition is internally questionable. The threshold theta in the rule \"if there exists a knowledge graph Gi such that S(Gx,Gi) <= theta, then output True\" is never reported, and if S is the Jaccard similarity defined in Eq. (5), a low similarity score indicates low overlap, not evidence of truth. The paper also does not specify the graph embeddings, the subgraph matcher, or the triple weighting function f(t) used in practice, nor does it show a single retrieved graph for the 600 test claims. In addition, the function definition in Section III-E allows the output \"Undetermined\", but the binary evaluation in Section IV-C gives no mapping from Undetermined cases to True/False.","section":"Section III-E, Eq. (5)"},{"comment":"The evidence source and the reasoner are the same model, which creates a circularity concern. GPT-4 with Advanced Data Analysis constructs the knowledge graphs from the query text, and GPT-4 then produces the verdict by reasoning over those graphs. Table II illustrates the problem: the first example's semantic health knowledge graph contains the conclusion \"More Deaths in 2021 than 2020\" as a node connected by \"implies\", meaning the graph can encode the answer during construction. Without a counterfactual test (for example, retrieving a fixed external graph or a randomly corrupted graph), the reported 88.5% accuracy may reflect GPT-4's parametric memory or post-hoc graph construction rather than the contribution of graph-based retrieval.","section":"Section III-D and Table II"},{"comment":"The experimental setup is not described at a level that permits replication or causal attribution. The manuscript does not state how many DBpedia triples remain after filtering, how those triples are assembled into knowledge graphs, how the triples enter the GPT-4 prompt, or what retrieval database size N is used. The multi-class result further weakens the binary claim: the six-category accuracy drops to 49.3%, and the authors acknowledge that the model tends to select middle categories such as \"Half True\" or \"Mostly False\". Because the binary labels are obtained by collapsing three coarse categories into True and three into False, the high binary accuracy could be an artifact of coarse label construction rather than of graph content. Per-category cell counts for the six-class confusion matrix are not provided.","section":"Section IV-A and Section IV-D"}],"minor_comments":[{"comment":"The text says \"d is the damping vector\" but d is a scalar damping factor; this should be corrected. The chosen values eta = 0.7 and alpha = 1.5 are given without sensitivity analysis or justification beyond a brief sentence.","section":"Section III-C"},{"comment":"The proof of Theorem 2 invokes eigenvalue perturbation theory [49] but does not verify the conditions under which the rank-one update leaves all eigenvalues except the Perron eigenvalue unchanged. The displayed convergence bound also uses a strict inequality without stating the required assumptions on the initial distribution and the spectrum of P.","section":"Section III-C, Theorem 2"},{"comment":"The six-category confusion matrix is shown only as a normalized heatmap; reporting numeric cell counts would allow readers to see how errors are distributed and whether the \"middle category\" tendency dominates the binary accuracy.","section":"Figure 8 and Section IV-D"},{"comment":"No code, data, prompt templates, or API access dates are provided, which makes the proprietary-model comparisons difficult to reproduce. At minimum, the exact prompts used for the baselines and for TrumorGPT should be included in an appendix.","section":"Throughout"},{"comment":"Reference [22] is cited for TextRank but has the title \"An introduction to information retrieval\", which appears mismatched; please check the citation. The reference to [50] should identify the specific theorem used for the convergence bound.","section":"References"}],"recommendation":"reject","confidential_remarks":"The central evaluation is uncontrolled and partially circular, and the missing details (threshold theta, prompt construction, retrieval examples, ablation) are not presentation issues but concern the validity of the main claim. I do not see how the manuscript can be made publishable without substantially new experiments that separate the graph retrieval mechanism from GPT-4's parametric knowledge. A future submission with a controlled ablation, a non-circular graph construction procedure, and released prompts/data could be reconsidered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the claimed GraphRAG advantage is not supported. The paper never isolates retrieval from GPT-4's parametric memory, and the example graphs already contain the conclusions. That said, the topic-specific TextRank part includes a clean convergence proof and a small numerical check; that's the one piece I'd actually use.\n\nWhat's new: the TST adaptation of TextRank with a topic-dependent teleportation vector, plus the proof of convergence (Theorems 1 and 2) and a Watts-Strogatz sanity check. That is a legitimate, if modest, algorithmic contribution. Applying graph retrieval to health fact-checking is a natural extension of published work, not a new phenomenon.\n\nThe problems are where the reader's stress test lands. No ablation: we never see TrumorGPT without GraphRAG, so the 88.5% vs GPT-4's 83.3% could be prompt wording. No threshold θ, no embedding or matcher details, no retrieved graphs for the test set, no error bars. The paper's own Table II shows a knowledge graph whose edge is literally \"More Deaths in 2021 than 2020\" – the conclusion of the query. That is a red flag that the graph is constructed after the fact from the source article, not independent evidence. And the decision rule S(Gx, Gi) ≤ θ ⇒ True is backwards for Jaccard similarity: small overlap should not mean truth. If it's a typo, the paper needs to fix it; as written it makes the mechanism incoherent.\n\nThe six-way result (49.3%) is honestly reported, but the authors explain it away by saying the model falls back to middle categories. That, combined with the binary collapse, means the 88.5% is likely an artifact of coarse labels rather than graph reasoning. I don't see a citation problem; they've cited the relevant TextRank variants and RAG work.\n\nBottom line: the formal TST analysis is worth keeping, but the empirical core is not reproducible or causally valid. This needs code, data, an ablation, and a corrected decision rule before it can be seriously reviewed. I would not accept it for peer review as-is; I'd send it back for a major revision or reject it.","headline":"The topic-specific TextRank proof is solid, but the GraphRAG accuracy claim is untested and the retrieval rule is internally contradictory.","tokens_in":21122,"tokens_out":3884,"would_cite":false,"duration_ms":39111,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"TrumorGPT claims that adding graph-based retrieval over semantic health knowledge graphs to GPT-4 fact-checks health statements with 88.5% binary accuracy, beating six general-purpose LLMs on PolitiFact health-care and coronavirus claims.","keywords":["fact-checking","GraphRAG","semantic health knowledge graph","topic-specific TextRank","topic-enhanced sentence centrality","large language models","health misinformation","PolitiFact"],"falsifier":"Run the same 600 PolitiFact claims through TrumorGPT with the graph-retrieval step disabled or with randomly selected health graphs; if accuracy stays near 88.5%, the graphs are not load-bearing. The paper reports no such ablation, no value for the threshold $\\theta$ used to declare a match, and no account of how DBpedia triples enter the prompt, so the mechanism is untested as stated.","tokens_in":20014,"feed_emoji":"🩺","tokens_out":8374,"duration_ms":71390,"temperature":0.7,"pith_summary":"TrumorGPT claims that a fact-checker can be made more accurate by giving GPT-4 access to a structured, up-to-date health knowledge graph and retrieving evidence from that graph before answering. On 600 PolitiFact 'Health Care' and 'Coronavirus' statements, it reports 88.5% binary accuracy, above GPT-4's 83.3% and five other LLMs. The paper's contribution is a graph-construction pipeline, topic-enhanced sentence centrality plus topic-specific TextRank with few-shot learning, feeding a GraphRAG layer that is meant to let the model verify claims against recent health facts instead of its static training data. A sympathetic reader would care because health misinformation is high-stakes and the reported gain is concrete; the authors also report that the graph layer keeps verdicts concise, averaging 2.8 sentences per response.","feed_headline":"TrumorGPT fact-checks health claims at 88.5 percent accuracy","feed_subtitle":"A graph-retrieval layer lifts GPT-4 past six rival models on PolitiFact health and coronavirus claims.","key_machinery":"The load-bearing object is the semantic health knowledge graph, a directed graph $G = \\{E,R,F\\}$ whose vertices are entities, edges are relations, and facts are triples $(h,r,t)$. To build these graphs, the paper uses topic-enhanced sentence centrality (BERT embeddings weighted with an LDA topic vector, $\\eta = 0.7$) to pick key sentences and topic-specific TextRank, a PageRank variant with a topic-relevance teleportation vector and a health boost factor $\\alpha = 1.5$, to rank the sentences; GPT-4 with few-shot examples then constructs the graph. GraphRAG retrieves stored graphs by comparing a query graph to each candidate with a weighted Jaccard score over consecutive triples, with subgraph isomorphism as the match criterion, and feeds the retrieved evidence into GPT-4's semantic reasoning, which outputs True, False, or Undetermined. The Markov-chain convergence theorems guarantee the ranking iteration converges, but the fact-checking verdict itself is carried by the match between the query graph and the knowledge-base graph.","core_discovery":"The central claim is that TrumorGPT, a GPT-4-based framework augmented by graph-based retrieval-augmented generation over semantic health knowledge graphs, can separate true from false health-related statements with 88.5% accuracy on 600 PolitiFact claims (300 true, 300 false), beating GPT-3.5, GPT-4, LLaMA 3.2, PaLM 2, Claude 3.5 Sonnet, and Gemini 1.5 on accuracy, precision, recall, and F1. The knowledge graphs are built from DBpedia health triples, using few-shot GPT-4 with the proposed topic-enhanced sentence centrality and topic-specific TextRank; retrieval matches the query graph to stored graphs through Jaccard/subgraph-isomorphism scoring, and the retrieved evidence grounds the final true/false decision. The paper also reports that TrumorGPT is concise (2.8 sentences on average) and that its six-way PolitiFact classification is only 49.3% accurate, so its strength is specifically the binary true/false decision.","pith_inferences":["Because the paper reports no ablation disabling GraphRAG, the cleanest test of its causal story is TrumorGPT with and without graph retrieval; until that run exists, part of the 5.2-point gap over GPT-4 could come from prompt design or few-shot example choice rather than from the graphs.","If the graph layer is genuinely causal, its advantage should be largest on claims about events after GPT-4's December 2023 cutoff; the paper's use of post-2023 DBpedia triples makes that an empirical check.","The 49.3% six-category accuracy suggests the knowledge graphs encode enough evidence for a coarse true/false split but not for graded distinctions such as 'Half True' versus 'Mostly False'; a fact-checker aimed at nuanced verdicts would need a finer relation or evidence representation.","The reported 2.8-sentence average suggests the graph retrieval acts as an evidence filter; an extension could measure whether conciseness and accuracy are separately attributable to the graphs versus to the few-shot instruction style."],"forward_implications":["A periodically updated health knowledge base lets a fact-checker answer claims that postdate the base LLM's December 2023 training cutoff.","Graph-grounded verdicts are concise: TrumorGPT averages 2.8 sentences per response, shorter than every compared LLM.","The binary true/false formulation is where the system works; on PolitiFact's full six-category scale accuracy drops to 49.3%, so the framework is a two-way verifier rather than a fine-grained rater.","Retrieval from a curated health-only graph filters off-topic information, which the authors credit for precision (91.4%) exceeding recall (85.0%)."],"supporting_citations":[{"why":"Supplies the underlying GPT-4 model whose parametric knowledge and Advanced Data Analysis capability are used for graph construction and final verdicts.","marker":"[19]"},{"why":"Defines retrieval-augmented generation, the paradigm that GraphRAG extends.","marker":"[25]"},{"why":"Provides the graph-based RAG approach that TrumorGPT adapts for health knowledge graphs.","marker":"[26]"},{"why":"Supplies the RDF triples that are filtered by health keywords to build the semantic health knowledge graph database.","marker":"[53]"},{"why":"Provides the sentence centrality baseline that topic-enhanced sentence centrality modifies.","marker":"[20]"},{"why":"Provides the TextRank algorithm that topic-specific TextRank adapts with topic relevance weights.","marker":"[21]"},{"why":"Serves as a comparison baseline (LLaMA 3.2) whose lower accuracy the paper contrasts with TrumorGPT.","marker":"[15]"},{"why":"Serves as a comparison baseline (Gemini 1.5) in the evaluation table.","marker":"[54]"}],"fun_headline_variants":["Graph-retrieval LLM fact-checks health claims at 88.5% accuracy","TrumorGPT: graph-augmented GPT-4 beats rivals on health fact-checking","GraphRAG fact-checker hits 88.5% on health claims","LLM with graph retrieval fact-checks health at 88.5%","TrumorGPT uses graph retrieval to fact-check health claims"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claimed accuracy gain comes from the retrieved semantic health graphs rather than from GPT-4's own parametric knowledge or the prompt wording.","fun_headline_variants_meta":{"raw":{"variants":["Graph-retrieval LLM fact-checks health claims at 88.5% accuracy","TrumorGPT: graph-augmented GPT-4 beats rivals on health fact-checking","GraphRAG fact-checker hits 88.5% on health claims","LLM with graph retrieval fact-checks health at 88.5%","TrumorGPT uses graph retrieval to fact-check health claims"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000648,"raw_usage":{"total_tokens":3000,"prompt_tokens":993,"completion_tokens":2007,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":1904}},"tokens_in":609,"tokens_out":2007,"duration_ms":12904,"temperature":1.0,"reasoning_tokens":1904,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:26:14.052963+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 600 PolitiFact claims through TrumorGPT with the graph-retrieval step disabled or with randomly selected health graphs; if accuracy stays near 88.5%, the graphs are not load-bearing. The paper reports no such ablation, no value for the threshold $\\theta$ used to declare a match, and no account of how DBpedia triples enter the prompt, so the mechanism is untested as stated.","supporting_citations":[{"cited_title":"Retrieval- augmented generation for knowledge-intensive NLP tasks,","cited_arxiv_id":null,"evidence_quote":"Defines retrieval-augmented generation, the paradigm that GraphRAG extends."},{"cited_title":"DBpedia: A nucleus for a web of open data,","cited_arxiv_id":null,"evidence_quote":"Supplies the RDF triples that are filtered by health keywords to build the semantic health knowledge graph database."},{"cited_title":"Sentence Centrality Revisited for Unsupervised Summarization","cited_arxiv_id":"1906.03508","evidence_quote":"Provides the sentence centrality baseline that topic-enhanced sentence centrality modifies."},{"cited_title":"TextRank: Bringing order into text,","cited_arxiv_id":null,"evidence_quote":"Provides the TextRank algorithm that topic-specific TextRank adapts with topic relevance weights."}],"review_version":1}