Pith. sign in

Paper Citation Record · LEDGER

Long-form factuality in large language models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2403.18802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.18802 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:58:23.220001Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T06:44:18.784391Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1efbd665-8e61-4f59-9a39-601f5debd3a3 · inbound

Preserving Knowledge in Large Language Model with Model-Agnostic Self-Decompression cites this paper.

Preserving Knowledge in Large Language Model with Model-Agnostic Self-Decompression Long-form factuality in large language models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:13:39.462264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-24T00:09:52.093810Z digest=sha256:0773dc430facb32c23c0e7941217532da5bb861115d0420b002fb156d83567d3

Observation 06b1d0f2-7278-4747-b188-f7c739200a4b · inbound

Measuring short-form factuality in large language models cites this paper.

Measuring short-form factuality in large language models Long-form factuality in large language models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:50.288896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-15T06:45:50.219157Z digest=sha256:0a7a61c7c9125c701266c924b06800509d6c854c8273dc25d75e2f877c6fe51b

Observation c897e224-bf87-4b6e-ab14-3f6ae69407a8 · inbound

Towards Robust Evaluation of Unlearning in LLMs via Data Transformations cites this paper.

Towards Robust Evaluation of Unlearning in LLMs via Data Transformations Long-form factuality in large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:18:31.629667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:18:31.629667Z digest=sha256:cbb41dd592e155ff484b008162291d6916ea2403c8b7e733526e8d7db67d529b

Observation 6002745f-df62-44e0-a74b-388e571d69eb · inbound

RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models cites this paper.

RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models Long-form factuality in large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:16.330494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:08:16.330494Z digest=sha256:fe1c26631eaf5c6804a1b3372cce6cf380fc9d8934e5cf815239dfcef28f2b20

Observation eb5678ae-067b-4754-9ba4-46a0538a337e · inbound

Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM cites this paper.

Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM Long-form factuality in large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:53:12.361255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:53:12.361255Z digest=sha256:d35a6748b441c2de2183e091a71be8e39019a6d20ee30559f4df949fd890dad8

Observation 435aeff9-4beb-40f2-abd4-64dea21e67fe · inbound

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input cites this paper.

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input Long-form factuality in large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:11.920951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:11.920951Z digest=sha256:3a88eb5a93b42882016de6d62ff5fcd9e976c41af6f00c9112e5ae7227e34b6a

Observation 68e4e313-b78e-46bd-afcc-62d90a6680bd · inbound

Phare: A Safety Probe for Large Language Models cites this paper.

Phare: A Safety Probe for Large Language Models Long-form factuality in large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:58:23.220001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:58:23.220001Z digest=sha256:90cce22368916e0817a0943871d29ba3e7727ea90992a52dfa337952261606e1

Observation 0c1cca02-e050-42b6-81ca-fa039b8bc8a7 · inbound

Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation cites this paper.

Learning Auxiliary Tasks Improves Reference-Free Hallucination Detection in Open-Domain Long-Form Generation Long-form factuality in large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:35.414809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:41:35.414809Z digest=sha256:138e42763c13c21bac0286574b37ac51d93ab4871fe86b2918f0f0baa2622f88

Observation 93e1b83e-2620-44ce-86b5-de939e95e8df · inbound

HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs cites this paper.

HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs Long-form factuality in large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:34:07.814259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:34:07.814259Z digest=sha256:a1225c3c3bc904ca306e1b40ab8befac1f3953e31f4cb880dccb619c9f988120

Observation c2dcf01b-4e35-4276-95f5-0e034b3270a4 · inbound

RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking cites this paper.

RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking Long-form factuality in large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:51:47.184863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:51:47.184863Z digest=sha256:76ad889d5b61140a35498f8d4f3293fa9e0fce40f64d6c5e8f0c76ed8146d2c3

Observation 0f995fdc-e477-4550-9282-17b2f2db3e86 · inbound

Veracity: An Open-Source AI Fact-Checking System cites this paper.

Veracity: An Open-Source AI Fact-Checking System Long-form factuality in large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:39.856611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:39.856611Z digest=sha256:f64654dc74192ab213f23c31074ec9efdc24bf6f0b29ce1dd0e582cc3bbc8294

Observation 2fce3b73-9d30-4887-902c-f4726afa870e · inbound

Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? cites this paper.

Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? Long-form factuality in large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:21.746960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:02:21.746960Z digest=sha256:fd0f3d4dcb27775e1801fa76e5ecf613b717968e45e91ec8a9594d2efda1234e

Observation 9e1a384a-f6b2-4472-b578-acd638b17247 · inbound

JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring cites this paper.

JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring Long-form factuality in large language models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T14:51:03.793719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:51:03.793719Z digest=sha256:0d4b5f43204c7ad82a34e280f30b60a2bea5948dc107998cb5fa9ba9bf33c1c0

Observation cf27c228-ad09-4a55-bc2a-242a563f198e · inbound

DecMetrics: Structured Claim Decomposition Scoring for Factually Consistent LLM Outputs cites this paper.

DecMetrics: Structured Claim Decomposition Scoring for Factually Consistent LLM Outputs Long-form factuality in large language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T13:18:40.851754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:18:40.851754Z digest=sha256:fe3c16adc7c860bfcd86c743b91e7ba3e9f13f02d76519c61efa1c1e7931d9c8

Observation 0da5c309-6527-475d-86f9-86c56defb8bc · inbound

Human-AI Complementarity: A Goal for Amplified Oversight cites this paper.

Human-AI Complementarity: A Goal for Amplified Oversight Long-form factuality in large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:22.548978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:22.548978Z digest=sha256:c081384904124ba83d3be55edddc9a4f8ce1eed3768524b831a8c9884f87dd61

Observation 2829bc8c-14c0-46d3-8f85-8ca0e8a248ff · inbound

All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting cites this paper.

All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting Long-form factuality in large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T22:21:36.837859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:21:36.837859Z digest=sha256:4c116038985b37eaa229116c906bf9e171bedb3e4d1c17893d3a2ec68efe8dc7

Observation b23702a5-67fc-4a5b-992c-45f6d364fca5 · inbound

Beyond Precision: Importance-Aware Recall for Factuality Evaluation in Long-Form LLM Generation cites this paper.

Beyond Precision: Importance-Aware Recall for Factuality Evaluation in Long-Form LLM Generation Long-form factuality in large language models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:23:13.860761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T20:20:16.299929Z digest=sha256:c74a9fd74148dd562ab1fef25ac8ca59a3565e337f9c299306c4deab45a22a24

Observation 69ce7a89-cf1a-4c6a-9687-9ed399fd4f38 · inbound

VerifAI: A Verifiable Open-Source Search Engine for Biomedical Question Answering cites this paper.

VerifAI: A Verifiable Open-Source Search Engine for Biomedical Question Answering Long-form factuality in large language models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:07:58.734248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T14:03:58.706599Z digest=sha256:bf047e536bf33c83573672b76f6652d59ea9a62cdc43b8cb5c0f22e23e73d3e8

Observation e771731a-56e5-4628-ae6f-3b69b81154bb · inbound

Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systems cites this paper.

Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systems Long-form factuality in large language models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:02.238245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T05:31:11.718051Z digest=sha256:4637fad7c13fec6a5f25364ddea979a485e1ee2a873e3be151bb4788915362d3

Observation 1257812d-bb5c-4dab-a9a8-fe38ec1b1653 · inbound

Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systems cites this paper.

Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systems Long-form factuality in large language models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:59:15.228903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T23:56:44.979471Z digest=sha256:c8eee83d27368bb432032a3cc257c7828954cee0b9e913ed384bf7ded51787d3

Observation 3133827c-6439-48ad-bf6a-4851398e89b2 · inbound

FinGround: Detecting and Grounding Financial Hallucinations via Atomic Claim Verification cites this paper.

FinGround: Detecting and Grounding Financial Hallucinations via Atomic Claim Verification Long-form factuality in large language models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:11:20.023407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T06:24:43.021602Z digest=sha256:34361a01de71f84adcf4f9a0ba4abba89e94bed18b2b0a8cc9f50cc3b0b05b70

Observation 7ed87e6d-64d2-4283-816f-b245dccae29e · inbound

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs cites this paper.

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs Long-form factuality in large language models

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:27.235175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:50:17.399580Z digest=sha256:e81db8552ea62f4c95886b0a0cd789efa486541a757675982cb0ef744d037441

Observation b24693e0-49c3-4527-9bf8-95a34920c2dd · inbound

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs cites this paper.

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs Long-form factuality in large language models

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:22:28.962869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T07:20:32.494840Z digest=sha256:e536479003f83958216b3e75eae3f9c66a8ebc522b0bbec6d9e8a7bd44bf0ba9

Observation cb1d8c87-bea7-45d8-8671-28a21c766c1a · inbound

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios? cites this paper.

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios? Long-form factuality in large language models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:44:18.785878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T06:42:30.226214Z digest=sha256:c6f8a732af055b6599a3512b6873cdaf284fd88913e95f0cb54f21c256e6abcc

Observation dc54ef0c-6cf5-47bc-b6ae-eea8cc20e1f2 · inbound

Answer-Reconstruction Search Density: Measuring the Query and Source Work Compressed by Conversational Answers cites this paper.

Answer-Reconstruction Search Density: Measuring the Query and Source Work Compressed by Conversational Answers Long-form factuality in large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T14:03:17.719461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:03:17.719461Z digest=sha256:09248dc60cca6a69c6ae96f41a745a6e2ba71806c8e04da5191fa348779a1cd0