Pith. sign in

Paper Citation Record · LEDGER

FELM: Benchmarking Factuality Evaluation of Large Language Models

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2310.00741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.00741 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:41:35.442979Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T12:56:24.475550Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f24bea38-ebb2-4a4d-a456-cf5a7b856883 · inbound

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions cites this paper.

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions FELM: Benchmarking Factuality Evaluation of Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:46:27.452995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T02:46:26.957539Z digest=sha256:f14e2d14ed0dc3dfb37889c9df8c49b43aac6bde44ea08ba6dcc69c491a0336e

Observation f67ea904-82ad-45f6-96fe-888ce408dc48 · inbound

Measuring short-form factuality in large language models cites this paper.

Measuring short-form factuality in large language models FELM: Benchmarking Factuality Evaluation of Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.295212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:bfd446c251b8506d6b59f5e9f5191c7062754b5ac5a3fd72f4db5dbcbc70498f

Observation d7b2f14d-b2d5-4eff-90e9-e66ef016ccdf · inbound

Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation cites this paper.

Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation FELM: Benchmarking Factuality Evaluation of Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:35.442979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:41:35.442979Z digest=sha256:c01a34d40b1876310c6b1719ed31360cf1a5373924e396b5d505f2c2545d60f3

Observation 813cdce5-ddb9-4dbc-b57a-58910839ddd6 · inbound

Beyond Facts: Evaluating Intent Hallucination in Large Language Models cites this paper.

Beyond Facts: Evaluating Intent Hallucination in Large Language Models FELM: Benchmarking Factuality Evaluation of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:43.265833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:59:43.265833Z digest=sha256:2e824522f9863d4297b974427845335f6b243d86c4b31b2fc0d8be5fa005d027

Observation b10f13f3-359f-49e1-b274-876e00aa10c5 · inbound

ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations cites this paper.

ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations FELM: Benchmarking Factuality Evaluation of Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:56:24.477667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T12:54:01.717015Z digest=sha256:8fe85f721034aa0566705754a35a8c8073557454169390e575a8a7c41c649391

Observation 2f05b71f-d49a-43e9-9018-d26af0e805ae · inbound

HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification cites this paper.

HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification FELM: Benchmarking Factuality Evaluation of Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T14:47:41.336316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:47:41.336316Z digest=sha256:0734e4c9527f38baff9b7392649db4a95f51180fb2f4fca99bade389397727d3