Pith. sign in

REVIEW 7 cited by

Lynx: An Open Source Hallucination Evaluation Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.08488 v2 pith:XJTUURHH submitted 2024-07-11 cs.AI cs.CL

classification cs.AIcs.CL
keywords lynxhallucinationevaluationhalubenchllmsmodelsreal-worldaccess
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Retrieval Augmented Generation (RAG) techniques aim to mitigate hallucinations in Large Language Models (LLMs). However, LLMs can still produce information that is unsupported or contradictory to the retrieved contexts. We introduce LYNX, a SOTA hallucination detection LLM that is capable of advanced reasoning on challenging real-world hallucination scenarios. To evaluate LYNX, we present HaluBench, a comprehensive hallucination evaluation benchmark, consisting of 15k samples sourced from various real-world domains. Our experiment results show that LYNX outperforms GPT-4o, Claude-3-Sonnet, and closed and open-source LLM-as-a-judge models on HaluBench. We release LYNX, HaluBench and our evaluation code for public access.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers

    cs.AI 2026-05 conditional novelty 7.0 of 10

    Typed quantity verification exposes a canonical-equivalence blind spot in neural fact-checkers; Symbolic Augmentation fixes it (36.5%→98.2%) and transfers to SciFact-Open (+0.037 binary macro-F1).

  2. Retromorphic Testing with Hierarchical Verification for Hallucination Detection in RAG

    cs.CL 2026-03 conditional novelty 6.0 of 10

    A claim-by-claim hierarchical verifier improves RAG hallucination detection over baselines, and a re-annotated benchmark finds 1.68x more hallucinated cases than the original labels.

  3. Reconsidering LLM Uncertainty Estimation Methods in the Wild

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Most LLM uncertainty estimates degrade under distribution shift and adversarial prompts, but simple ensembling of scores at test time improves reliability.

  4. TruthTorchLM: A Comprehensive Library for Predicting Truthfulness in LLM Outputs

    cs.CL 2025-07 conditional novelty 5.0 of 10

    TruthTorchLM is a new open-source library that standardizes 30+ LLM truthfulness prediction methods and benchmarks them on three datasets.

  5. TUM-MiKaNi at SemEval-2025 Task 3: Towards Multilingual and Knowledge-Aware Non-factual Hallucination Identification

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A retrieval-based fact checker and a BERT classifier, combined by support vector regression, reach top-10 multilingual hallucination detection in eight of fourteen languages on the Mu-SHROOM shared task.

  6. Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Using hallucinated benchmark answers as rejected DPO pairs, ordered by an external fact-checker's grounding score, improves hallucination detection in 1B-3B Llama models.

  7. Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models

    cs.CL 2025-06 reject novelty 3.0 of 10

    A survey of LLM hallucination research that formalizes hallucination types and argues, via incompleteness and undecidability arguments, that hallucinations cannot be fully eliminated.

Pith tools