Pith. sign in

Paper Citation Record · LEDGER

RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2311.08147.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.08147 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T12:00:32.617408Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T05:13:56.507219Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 41fc09dd-7012-45fa-9c9a-1865ca71d008 · inbound

Retrieval-Augmented Generation for Large Language Models: A Survey cites this paper.

Retrieval-Augmented Generation for Large Language Models: A Survey RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:13:56.510163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T05:10:25.171044Z digest=sha256:5a0898104b5f3989252850608d1146eeacef98c0c1d832f5d9cba147f573ca51

Observation 47dd87a8-cc73-45cb-bb5c-9f56620c9942 · inbound

MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries cites this paper.

MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:54:19.697179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T13:54:19.264893Z digest=sha256:c19703185a14886c7ddfad18991b9cce3ccd3489ad585d105bea6dee9005462c

Observation 15cbf505-2427-41dd-b33b-7d393d2f5b4b · inbound

A Survey on Retrieval-Augmented Text Generation for Large Language Models cites this paper.

A Survey on Retrieval-Augmented Text Generation for Large Language Models RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:15:55.504139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T02:15:05.379583Z digest=sha256:975354269a6a23547e36008545ac4770d24e1bfb58d0c76506ed104080d01f0a

Observation 97719dd1-7d00-4b95-baa4-d59979e22716 · inbound

Trustworthiness in Retrieval-Augmented Generation Systems: A Survey cites this paper.

Trustworthiness in Retrieval-Augmented Generation Systems: A Survey RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:08:25.947076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T21:08:11.787013Z digest=sha256:433d0a74fe71ec92366bc209cffd8f4ed4355b127a3c47944f2b5c87cc724d55

Observation 6e73874d-07a9-468f-b954-113bad16f86f · inbound

Investigating the Robustness of Deductive Reasoning with Large Language Models cites this paper.

Investigating the Robustness of Deductive Reasoning with Large Language Models RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T12:00:32.617408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:00:32.617408Z digest=sha256:88719ae3b79b2b26660e7cf7301a46f4e4b96f4c48569e43add8cda43098fd89

Observation 86e94ea8-943c-4b6c-acdb-cd6ff8e4f462 · inbound

Magic Mushroom: A Customizable Benchmark for Fine-grained Analysis of Retrieval Noise Erosion in RAG Systems cites this paper.

Magic Mushroom: A Customizable Benchmark for Fine-grained Analysis of Retrieval Noise Erosion in RAG Systems RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:57:26.924212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:57:26.924212Z digest=sha256:32e9fa0d2a757027fbb82b4b81a7429eb5f8b2a41ba58cbc383793d799dbcfdb

Observation 64d03a24-522f-4ce1-a251-8bc9e3d6c841 · inbound

GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation cites this paper.

GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge

Reference 9474

Resolution
unresolved
no resolver link, observed 2026-08-07T05:33:18.221584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:33:18.221584Z digest=sha256:f53be29e2077c162d4c3a01ce16c9af1ab6308150b1c9250334580d42f2e3307

Observation 725a453f-3efc-440f-a043-964bee82a586 · inbound

"Lost-in-the-Later": Framework for Quantifying Contextual Grounding in Large Language Models cites this paper.

"Lost-in-the-Later": Framework for Quantifying Contextual Grounding in Large Language Models RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:32:20.719247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:32:20.719247Z digest=sha256:e0c448f105404e8ecc3744c2df991e02ac4d6ce205ee73647665a7554ca65533

Observation 146e528a-1dca-4fae-81ed-72ac81a42570 · inbound

Quasiparticle interference in LiFeAs: Signature of inelastic tunneling through spin fluctuations cites this paper.

Quasiparticle interference in LiFeAs: Signature of inelastic tunneling through spin fluctuations RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T19:50:49.876347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:50:49.876347Z digest=sha256:06fd4d052ac2cae4f6b4e54a73f4b33f8eaa99561857bd848b1ade61d4e0fc4f

Observation e82784f5-1350-4159-94cf-5032c79e0449 · inbound

Do Audio-Visual Large Language Models Really See and Hear? cites this paper.

Do Audio-Visual Large Language Models Really See and Hear? RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:15.854047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:56:19.815569Z digest=sha256:cc8428eec68a5a299376605dbe2db3e0b7859024393b5a3ec3caa8072f17653b

Observation 9c69728c-b56c-43f2-af83-36ccdeb70603 · inbound

Overcoming the "Impracticality" of RAG: Proposing a Real-World Benchmark and Multi-Dimensional Diagnostic Framework cites this paper.

Overcoming the "Impracticality" of RAG: Proposing a Real-World Benchmark and Multi-Dimensional Diagnostic Framework RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:28:13.946502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T20:27:57.033386Z digest=sha256:ddf1dc4762d81867c69a4a56d6ffa048f8dd3c929289d075429b2c3b155c65b6

Observation 80095f73-ebec-4846-a8d4-6b456bdb5e14 · inbound

EHRAG: Bridging Semantic Gaps in Lightweight GraphRAG via Hybrid Hypergraph Construction and Retrieval cites this paper.

EHRAG: Bridging Semantic Gaps in Lightweight GraphRAG via Hybrid Hypergraph Construction and Retrieval RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge

Reference 155

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:51:10.610246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T05:43:04.813867Z digest=sha256:e5fb754595fb51ef71e2b30e6f6d9efe02319d51868e19b77220b91a4a469bde