{"id":"f091420e-caec-4cc1-9f75-5ad6c5e5840a","arxiv_id":"2502.08109","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A fine-tuned 8B Llama model reports higher hallucination-detection accuracy than GPT-4 and Llama3 70B on four benchmark test sets, while its generated explanations are comparable but not consistently better.","lead":"A team trained a smaller model, HuDEx, to flag when an AI answer contains made-up information and to explain the flag. The authors report it beats much larger models on four hallucination benchmarks, though the evaluation needs stronger checks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 5.1's wording that test sets 'were also used during training' leaves open that Table 3's detection gains reflect memorization, not generalization.","rationale":"The reader's weakest assumption is exactly the most load-bearing concern: Section 5.1's ambiguous sentence about using test sets 'also used during training' threatens the validity of the primary detection comparison. I agree with the conditional verdict because the paper otherwise presents plausible results, but the train/test separation must be verified before the headline claim is trustworthy. The explanation evaluation is secondary; the detection accuracy claim is the foundation. No independent evidence (e.g., released code, data splits, error bars) currently offsets the ambiguity, so the conditional verdict stands: require clarification of the split separation, add significance measures, and ideally release artifacts. If the contamination concern is resolved, the central claim may hold; if not, the reported margins are not evidence of generalization.","tokens_in":10246,"tokens_out":3247,"duration_ms":28379,"concrete_test":"Audit the fine-tuning data pipeline for exact or near-duplicate overlap between every test instance in the four benchmark test splits and every training instance used for LoRA, including the Llama-3-70B-generated explanation data. If overlap is nonzero, retrain with strictly disjoint splits and recompute Table 3. The concern is settled if the authors provide the overlap count or reproduce the Table 3 margins using only the Table 1 training splits; matching margins clear the concern, while materially lower margins confirm contamination.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that a fine-tuned Llama-3.1-8B (HuDEx) beats Llama-3-70B and GPT-4o on hallucination detection. The evidence is Table 3, evaluated on HaluEval dialogue/QA, FactCHD, and FaithDial. Section 5.1 states: 'we used the test sets from HaluEval dialogue, HaluEval QA, FaithDial and FactCHD, which were also used during training.' If the test splits were literally included in LoRA fine-tuning, the reported accuracies (80.6, 89.6, 70.3, 58.8) are contaminated: the model could memorize labels, and the comparison against Llama3 70B and GPT-4o is not a fair generalization test. This is load-bearing because the headline 'surpasses larger LLMs' depends entirely on these four numbers. The zero-shot results in Table 4 do not rescue the claim: HuDEx wins summarization (77.9 vs 69.55/61.9) but loses on HaluEval general (72.6 vs GPT-4o 78.0). The paper never states that only the training splits were used, and no code or data are released to verify split separation. This ambiguity must be resolved before the central claim can be accepted.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes HuDEx, a Llama-3.1-8B model fine-tuned with LoRA on the HaluEval dialogue/QA, FactCHD, and FaithDial datasets to perform binary hallucination detection and generate natural-language explanations for its decisions. The central empirical claim is that HuDEx surpasses larger models such as Llama3 70B and GPT-4o in detection accuracy on four in-distribution test sets (Table 3) and produces explanations nearly as good as the human-authored FactCHD explanations (Tables 5 and 6). The paper also reports zero-shot detection results on HaluEval summarization and HaluEval general (Table 4). The authors frame the contribution as an integration of hallucination detection with explanation generation, aimed at practical reliability.","tokens_in":10467,"tokens_out":3788,"duration_ms":32823,"significance":"If the claims hold, the paper demonstrates a useful direction: a compact 8B fine-tuned model can serve as a specialized hallucination detector with explanations, potentially offering a more interpretable and cheaper alternative to frontier general-purpose LLMs for model evaluation. The integration of detection and explanation is a worthwhile goal, and the training/inference pipeline is described in enough detail to be replicable in principle. The paper also attempts a zero-shot evaluation on unseen HaluEval subsets, which is a positive step. However, the significance is currently tempered by the lack of statistical rigor (single runs, no confidence intervals), the ambiguity about whether the test sets were used in training, and the heavy reliance on LLM-generated explanations and LLM-as-judge evaluation without reported human validation. The explanation results are only competitive (97\\u201399% conversion ratio on FactCHD) rather than demonstrating a clear advantage, so the main contribution rests on the detection accuracy numbers.","major_comments":[{"comment":"The sentence 'we used the test sets from HaluEval dialogue, HaluEval QA, FaithDial and FactCHD, which were also used during training' is ambiguous and potentially self-incriminating. If the exact test instances appeared in the LoRA fine-tuning data, the Table 3 accuracies (80.6, 89.6, 70.3, 58.8) reflect memorization rather than generalization, and the headline claim of surpassing Llama3 70B and GPT-4o is invalid. This is load-bearing because the abstract's central claim depends entirely on these four numbers. The authors must state unambiguously that only the official training splits were used for fine-tuning and that the test splits were held out, and they should provide evidence (e.g., data split configuration, code, or a clarification in the text) to rule out contamination.","section":"\\u00a75.1"},{"comment":"The abstract says HuDEx 'surpasses larger LLMs, such as Llama3 70B and GPT-4, in hallucination detection accuracy,' but the zero-shot results in Table 4 show GPT4o achieving 78.0 on HaluEval general versus HuDEx's 72.6. The claim is therefore too broad. The abstract and conclusion should be qualified to the in-distribution benchmarks or to the HaluEval summarization zero-shot case, and the paper should discuss why performance does not transfer to the general user-query setting.","section":"\\u00a76.1.2"},{"comment":"All detection results in Tables 3 and 4 are point estimates from a single evaluation run, with no error bars, confidence intervals, or significance tests. For instance, the HaluEval QA difference between HuDEx (89.6) and GPT4o (86.6) may be within noise, especially given that accuracy on these benchmarks typically varies with decoding settings and prompt details. The paper should report multiple runs (e.g., different seeds or repeated sampling) and provide variance estimates, or at least explicitly acknowledge the lack of statistical significance testing and temper the superiority claims accordingly.","section":"\\u00a76.1 and \\u00a75.2.1"},{"comment":"The explanation-quality evaluation is not sufficiently grounded. For HaluEval and FaithDial, the explanations were generated by Llama3 70B (the same model used as a detection baseline), and for all datasets the judge is GPT4o. The paper states that human evaluation was conducted on a statistically sampled subset, but no results of that human evaluation are reported, and the sampling parameters (99% confidence, p=0.5, 2% margin of error) do not specify the achieved sample size or the measured defect rate. Additionally, the training-data filtering step that retains only the 93.7% of examples where Llama3 70B's predicted labels matched the gold labels could bias the model toward easy or idiosyncratic cases. The reliability of the explanation claim needs either reported human-evaluation numbers or a critical discussion of these limitations.","section":"\\u00a73.2 and \\u00a76.2"}],"minor_comments":[{"comment":"The model name is spelled inconsistently: 'HuDEx' in the title and most of the paper, but 'HuDex' in Section 1. Please unify the spelling.","section":"\\u00a71"},{"comment":"There is a typo in 'categorize hallucinations into two broad two types' - 'two' is repeated.","section":"\\u00a72.2"},{"comment":"The persona generation process ('we provided ChatGPT with task details') is under-specified. It is unclear which ChatGPT model was used and how the persona selection was performed; more detail would aid reproducibility.","section":"\\u00a74.2 and Figure 3"},{"comment":"The phrase 'we used GPT4o as the judge for experiment' is missing an article; also, the version of GPT4o and decoding parameters are not reported, which are relevant for a judge-based evaluation.","section":"\\u00a75.2.2"},{"comment":"The terms 'anomalies' (4.2%) and 'responses that failed to understand the prompt' (0.5%) are not defined. What constitutes an anomaly, and how were these cases filtered out?","section":"\\u00a73.2"},{"comment":"The sentence 'On the FactCHD dataset, HuDEx outperformed Llama3 70B by ~11%' is ambiguous: 70.3 vs 59.4 is a difference of 10.9 percentage points, which is about 18% relative improvement. Please specify which measure is being used.","section":"\\u00a76.1.1"},{"comment":"Some reference entries are incomplete or inconsistently formatted (e.g., [25] lacks venue information, and [29] lists only page numbers without a publisher). Please verify the bibliography.","section":"References"},{"comment":"The tables would benefit from confidence intervals or at least the number of samples per benchmark. The test sizes are partially given in Table 1 (e.g., 1,000 for HaluEval) but not repeated in the results tables, making the practical significance of the differences hard to judge.","section":"Tables 3 and 4"}],"recommendation":"major_revision","confidential_remarks":"The most pressing issue is the Section 5.1 wording about test sets being 'also used during training.' I do not believe this is necessarily an admission of leakage, but it is a serious ambiguity that the authors must resolve explicitly. If they cannot confirm that the official test splits were excluded from fine-tuning, the detection results should be withdrawn. Note also that the zero-shot HaluEval general result undermines the unqualified 'surpasses larger LLMs' claim, and the explanation evaluation lacks the reported human-assessment numbers promised by the statistical sampling description. This paper may be suitable for publication after these issues are addressed, but it currently requires substantive revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look, but the central claim is only as good as one ambiguous sentence. The paper fine-tunes Llama-3.1-8B with LoRA to both detect hallucinations and produce explanations, and reports higher accuracy than GPT-4o and Llama3 70B on HaluEval dialogue/QA, FactCHD, and FaithDial. That is a genuinely useful empirical result if it holds: a compact model that can serve as a guardrail, with explanations attached, is practical for deployment.\n\nWhat is actually new is the specific package—detection plus explanation from one small fine-tuned model, with a prompt design that adapts to whether background knowledge is present. The data construction is careful: they filtered out explanations that did not align with existing labels, kept only verified matches, and ran a human validation on a statistically sampled subset of the generated training explanations. That is real work, and the detection scoreboard in Table 3 is consistent across all four benchmarks.\n\nThe soft spots are real. Section 5.1 says they used the test sets from HaluEval, FactCHD, and FaithDial, \"which were also used during training.\" That wording, taken literally, means the test instances were in the training data. The stress-test note is right to flag this: it is load-bearing, because the entire 'surpasses larger LLMs' claim sits on those four accuracy numbers. It is probably just sloppy phrasing—they mean the datasets were used for training, not the test splits—but with no code or data released, the reader has no way to verify separation. This must be fixed before the abstract claim can be accepted.\n\nThe other weaknesses are more ordinary: no error bars or significance tests, so a 1.5–3 point gap over GPT-4o on QA could be noise; the abstract overclaims, since GPT-4o wins on HaluEval general and HuDEx explanation factuality trails Llama3 70B on two datasets; and explanation quality is judged only by GPT-4o, with no human evaluation of the final explanations. Those are fixable with revision, not fatal flaws.\n\nI would send this to peer review. The method is sound, the comparisons are mostly fair, and the potential utility is clear. But the authors must clarify train/test separation explicitly, release artifacts, add variance reporting, and soften the abstract to match what the data actually show.","headline":"A practical 8B hallucination detector with explanations, but the Section 5.1 wording about test sets used during training must be clarified before the headline claim can be trusted.","tokens_in":11039,"tokens_out":1741,"would_cite":false,"duration_ms":16322,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that HuDEx, an 8-billion-parameter LoRA fine-tune of Llama 3.1, detects LLM hallucinations more accurately than Llama3 70B and GPT-4o on four benchmarks, with explanations judged 97–99% as good as human-written ones.","keywords":["hallucination detection","explainability","LLM reliability","LoRA fine-tuning","HaluEval","FactCHD","FaithDial","LLM-as-judge"],"falsifier":"Ask the authors for the exact training split and re-run the detection experiments with any test instances that also appear in fine-tuning removed; if accuracy drops substantially, the central claim is an artifact. Independently, re-score a random sample of HuDEx's explanations with human raters instead of an LLM judge: the reliable-explanations claim would be falsified if human factuality and clarity scores fall well below the reported 97–99% conversion ratios.","tokens_in":10013,"feed_emoji":"🔍","tokens_out":11666,"duration_ms":93776,"temperature":0.7,"pith_summary":"HuDEx is built around a simple idea: a hallucination detector should say why, not just that. The authors fine-tune a compact Llama 3.1 8B model with LoRA on the HaluEval, FactCHD, and FaithDial datasets so that a single pass produces both a verdict — hallucinated or not — and a plain-language explanation of the reasoning. Because two of those datasets lack explanations, the authors had Llama3 70B generate them and kept only the cases where the generated label agreed with the gold label (93.7% of samples), then spot-checked the explanations by human review. The central claim is that this two-task model beats much larger frontier models — Llama3 70B and GPT-4o — in detection accuracy on all four test sets (for example, 89.6% versus 86.6% and 82.7% on HaluEval QA), and produces explanations scored by an LLM judge at 97–99% of the human-written FactCHD gold explanations. The stated motivation is that pairing detection with explanation improves reliability: both users and the generating model itself can see and act on the mistake.","feed_headline":"8B model beats GPT-4o and Llama 70B at hallucination checks","feed_subtitle":"HuDEx also writes plain-language reasons, rated 97–99% as good as human labels.","key_machinery":"The load-bearing object is HuDEx itself: Llama 3.1 8B adapted with LoRA and trained on two tasks at once — hallucination detection and hallucination explanation — on the same examples. Three pieces carry the argument. The data pipeline: HaluEval and FaithDial explanations are machine-generated by Llama3 70B and filtered by requiring the generated label to match the gold label (93.7% agreement), with a statistically sampled human check at a 99% confidence level. The inference prompt: a fixed hallucination-expert persona plus adaptive task stages that branch on whether the prompt supplies background knowledge, forcing the model to reason from the source when it exists and from its own knowledge when it does not. The evaluation protocol: detection is scored as binary accuracy against four test sets, while explanation quality is graded by an LLM judge on two 3-point criteria, factuality and clarity.","core_discovery":"On its own terms, the paper's discovery is that interpreting a hallucination and detecting it are the same task, and that training them together makes a small model a stronger judge than a much larger general-purpose one. HuDEx takes a response and optional background knowledge, and with a hallucination-expert persona and stage-structured prompt it first classifies the response and then justifies the classification in natural language. The authors report that this 8B model outperforms Llama3 70B and GPT-4o in binary detection accuracy on the HaluEval dialogue (80.6%), HaluEval QA (89.6%), FactCHD (70.3%), and FaithDial (58.8%) test sets, and that in zero-shot settings it leads on HaluEval summarization (77.9%) while trailing GPT-4o on HaluEval general (72.6% versus 78.0%). For explanations, an LLM judge scores HuDEx's explanations within 2–3% of the original FactCHD human-written explanations and better than Llama3 70B on clarity in two of three datasets. The paper frames this as a move beyond evaluation-only benchmarks: a detector that explains itself can be used both by end users and as a tool to evaluate other LLMs.","pith_inferences":["A stress test the paper does not run: evaluate HuDEx against freshly annotated, never-published samples from the same task distributions; if its lead over GPT-4o evaporates outside the four fixed test sets, the generality claim would need revision.","Because the explanations were taught by Llama3 70B, the model's stated reasons will inherit that teacher's blind spots; training the same pipeline with a different explanation source (human annotations, or a smaller model) would isolate how much of the detection gain comes from the explanation data itself.","The conclusion sketches an automated feedback loop; a natural next experiment is feeding HuDEx's explanations back to the generating model and measuring whether rewriting its answer reduces hallucination rates on a second pass.","The zero-shot gap on HaluEval general suggests the detector transfers best to knowledge-grounded tasks; adding ungrounded user-query hallucination data to training would test whether the approach extends to the hardest, most open-ended cases."],"forward_implications":["A specialized 8B detector can replace frontier models for hallucination screening in knowledge-grounded settings, cutting cost and latency while improving accuracy.","Detection and justification come from one forward pass, so an end user or an auditing pipeline gets the reason along with the verdict, not an unexplained flag.","The same model can serve as a hallucination-focused judge for evaluating other LLMs, giving a more granular assessment than generic correctness scoring.","Zero-shot results on HaluEval summarization indicate that the detection skills transfer to at least some unseen task types without retraining.","The 93.7% label-agreement filter offers a concrete recipe for building explanation-annotated hallucination training sets from unlabeled benchmarks."],"supporting_citations":[{"why":"Supplies the HaluEval QA and dialogue training data and the two zero-shot test sets (summarization, general) that anchor most detection results.","marker":"[24]"},{"why":"Supplies the FactCHD training and test examples and the gold explanations used as the 100% quality ceiling in the explanation comparison.","marker":"[25]"},{"why":"Supplies the FaithDial training and test data, whose labels let the authors build binary hallucination instances from the dialogue responses.","marker":"[26]"},{"why":"Provides the BEGIN label scheme the authors rely on to preprocess FaithDial responses into hallucination and non-hallucination classes.","marker":"[27]"},{"why":"The Llama 3 model family: Llama 3.1 8B is the base being fine-tuned, and Llama3 70B is the teacher that generated the explanation training data and serves as a detection baseline.","marker":"[28]"},{"why":"LoRA, the parameter-efficient fine-tuning method that keeps training feasible on an 8B base model.","marker":"[29]"},{"why":"The GPT-4 technical report; GPT-4o serves as the frontier detection baseline that HuDEx must beat and as the LLM judge scoring explanation quality.","marker":"[30]"}],"fun_headline_variants":["8B model outscores GPT-4o and Llama 70B on hallucination detection","Detector that explains: 8B model beats 70B rivals on hallucinations","Small model detects hallucinations and writes human-grade explanations","HuDEx: one 8B model to catch and justify LLM hallucinations","Hallucination checks with reasons: 8B model outperforms GPT-4o"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole detection scoreboard rests on the assumption that the test examples were not seen during fine-tuning; the paper's wording about using test sets 'which were also used during training' leaves this ambiguous, so the reported lead over GPT-4o depends on the training and test splits being cleanly separated.","fun_headline_variants_meta":{"raw":{"variants":["8B model outscores GPT-4o and Llama 70B on hallucination detection","Detector that explains: 8B model beats 70B rivals on hallucinations","Small model detects hallucinations and writes human-grade explanations","HuDEx: one 8B model to catch and justify LLM hallucinations","Hallucination checks with reasons: 8B model outperforms GPT-4o"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00043,"raw_usage":{"total_tokens":2248,"prompt_tokens":1050,"completion_tokens":1198,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":666,"completion_tokens_details":{"reasoning_tokens":1095}},"tokens_in":666,"tokens_out":1198,"duration_ms":9522,"temperature":1.0,"reasoning_tokens":1095,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T10:24:54.990650+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask the authors for the exact training split and re-run the detection experiments with any test instances that also appear in fine-tuning removed; if accuracy drops substantially, the central claim is an artifact. Independently, re-score a random sample of HuDEx's explanations with human raters instead of an LLM judge: the reliable-explanations claim would be falsified if human factuality and clarity scores fall well below the reported 97–99% conversion ratios.","supporting_citations":[{"cited_title":"Halueval: A large-scale hallucination evaluation benchmark for large language models","cited_arxiv_id":null,"evidence_quote":"Supplies the HaluEval QA and dialogue training data and the two zero-shot test sets (summarization, general) that anchor most detection results."},{"cited_title":"Factchd: Benchmarking fact-conflicting hallucination detection, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the FactCHD training and test examples and the gold explanations used as the 100% quality ceiling in the explanation comparison."},{"cited_title":"Faithdial: A faithful benchmark for information-seeking dialogue","cited_arxiv_id":null,"evidence_quote":"Supplies the FaithDial training and test data, whose labels let the authors build binary hallucination instances from the dialogue responses."},{"cited_title":"Evaluating attribution in dialogue systems: The begin benchmark","cited_arxiv_id":null,"evidence_quote":"Provides the BEGIN label scheme the authors rely on to preprocess FaithDial responses into hallucination and non-hallucination classes."},{"cited_title":"The llama 3 herd of models, 2024","cited_arxiv_id":null,"evidence_quote":"The Llama 3 model family: Llama 3.1 8B is the base being fine-tuned, and Llama3 70B is the teacher that generated the explanation training data and serves as a detection baseline."},{"cited_title":"Lora: Low-rank adaptation of large language models","cited_arxiv_id":null,"evidence_quote":"LoRA, the parameter-efficient fine-tuning method that keeps training feasible on an 8B base model."},{"cited_title":"Gpt-4 technical report, 2024","cited_arxiv_id":null,"evidence_quote":"The GPT-4 technical report; GPT-4o serves as the frontier detection baseline that HuDEx must beat and as the LLM judge scoring explanation quality."}],"review_version":1}