Pith. sign in

Paper Citation Record · LEDGER

RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2409.12558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.12558 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:52:50.668532Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T13:31:25.571744Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e1c01c1c-5214-47e0-b87a-379eddca2b78 · inbound

MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems cites this paper.

MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:56:58.583654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:56:58.583654Z digest=sha256:b23c93044dd26532801980e799b8b685016940ee6a7e72183fa1bdd3765a25f8

Observation 037d4ad9-20fd-4f75-85d7-27851b8d649f · inbound

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets cites this paper.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.668532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.668532Z digest=sha256:9dd2a42fba261460e8bd7ce2f21731b5aaff25ca537eaab55295c745ef7ff48e

Observation 3748f086-724c-41fd-a4f5-1bb7ec0b7c11 · inbound

Benchmarking Poisoning Attacks against Retrieval-Augmented Generation cites this paper.

Benchmarking Poisoning Attacks against Retrieval-Augmented Generation RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:57.345015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:57.345015Z digest=sha256:0f40c609d7853a97cd20339a37c50e28793caf01e17f4f01fb74009217bcd718

Observation 4be80c42-00fb-4945-af81-cc9ffa3eab15 · inbound

RAGAPHENE: A RAG Annotation Platform with Human Enhancements and Edits cites this paper.

RAGAPHENE: A RAG Annotation Platform with Human Enhancements and Edits RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:13:34.417415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:13:34.417415Z digest=sha256:485235b0e749c9629ed797c3eae3ff5cbd917e4c268a5b7acfdc72ddebeb51b6

Observation 1a605631-59ab-4aaa-a8df-63e29c5b7dfa · inbound

Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks cites this paper.

Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:31:25.574510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T13:26:58.309566Z digest=sha256:4ef85faad2275bf2be18884c6faa504afde24d81a313ae0edcb6b070a7996f2e

Observation e4d9157b-1140-4e9a-9ad5-0fd05e5f10e8 · inbound

H-RAG at SemEval-2026 Task 8: Hierarchical Parent-Child Retrieval for Multi-Turn RAG Conversations cites this paper.

H-RAG at SemEval-2026 Task 8: Hierarchical Parent-Child Retrieval for Multi-Turn RAG Conversations RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:46:10.409356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-09T19:22:10.806765Z digest=sha256:359106bd51e38f109450f8546257b40eb1d5eca2b3c42e48a172727ca0e4c099