Pith. sign in

Paper Citation Record · LEDGER

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets

As of 19 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 2 inbound Pith citation observations for arXiv:2504.20119.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.20119 v2

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:52:50.682341Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:39:30.304142Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T20:58:45.171226Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact10
  • verified fuzzy14
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f34183de-1a7b-4400-8d65-9579e70d1bfe · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.301627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.301627Z digest=sha256:e90eb20f476e9105fe8e6c48f54ca0013755e22e23fdcdef97700c2ae3d3b8ba

Observation 760f13f0-68a5-48fb-95e7-c919bd96b33e · outbound

This paper cites Si ren’s Song in the AI Ocean: A Survey on Hallucination in Large Langu age Models,.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Si ren’s Song in the AI Ocean: A Survey on Hallucination in Large Langu age Models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:52.242603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.307198Z digest=sha256:2c7fb452449cea72825edf66337575757a7d10e1ae28dbe7bc1ce5deadb4c3e1

Observation 664fcab1-507d-4238-8aa8-5ad4f7ba42e7 · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.317258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.317258Z digest=sha256:9142c3d163beee49925d3ff52ff1be05a84f22dc36160fc4a0a8cf93807cc8ed

Observation bf273c94-36fe-497d-949c-4126e62d3c64 · outbound

This paper cites CRUD-RAG: A comprehensive chines e benchmark for retrieval-augmented generation of large lan guage models.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets CRUD-RAG: A comprehensive chines e benchmark for retrieval-augmented generation of large lan guage models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:52.227058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.322909Z digest=sha256:e10c4becb19ecfcba42afb7d352dc8ae8d71319b84472970af92cebb2bccf6eb

Observation 4d9c402b-f066-4d11-b359-ff0ee5633b6e · outbound

This paper cites Performance evaluation of vector embeddings with retrieval- augmented generation,.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Performance evaluation of vector embeddings with retrieval- augmented generation,

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-16T05:52:51.946105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.330430Z digest=sha256:2a1d4bb39bd1732322b334507a3e9477d2182dd784cb8f806896e73ac5f50dc2

Observation 5ea4758c-fc90-43f0-82b2-6e0c427eb72c · outbound

This paper cites Ragas: Automated Evaluation of Retrieval Augmented Generation.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Ragas: Automated Evaluation of Retrieval Augmented Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.338817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.338817Z digest=sha256:b2f43671f32d1707a4decc652f0ab6b21f5f9372289c0e55921b4017081db5e8

Observation abc963ba-231c-4279-8b8d-e507e433487e · outbound

This paper cites AR ES: An automated evaluation framework for retrieval-augmented g eneration systems.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets AR ES: An automated evaluation framework for retrieval-augmented g eneration systems

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:52.212350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.344510Z digest=sha256:b989731d7877574a71a95d3a28b75a1049fbbacb70ab9e5854f1b98395356e9e

Observation 433b7aa9-45aa-446f-aee8-a5d9bcb1e677 · outbound

This paper cites Evaluating Retrieval Quality in Retrieval-Augmented Generation.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluating Retrieval Quality in Retrieval-Augmented Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.349252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.349252Z digest=sha256:3dbaa2a782a531ae1d9fe54e172f3ff87fd52f3d447d35568718f573108e5cf6

Observation b8e54e03-e717-43b0-bc5f-257ab86bb0a2 · outbound

This paper cites Beyond Benchmarks: Evaluating Embedding Model Similarity for Retrieval Augmented Generation Systems.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Beyond Benchmarks: Evaluating Embedding Model Similarity for Retrieval Augmented Generation Systems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.354394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.354394Z digest=sha256:35f75d9a3a5b59c8b9685766e4808929537708c015cdaabce8372c14ad20df0d

Observation 7d407466-2bfd-41fb-80d6-07fd166235ea · outbound

This paper cites MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.359772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.359772Z digest=sha256:8cb401eda8c655fbded0c6aee7ab9e053d9c7c3e22a5cd40602594ac31ac14c6

Observation 863e1ce2-011b-4fb1-8a21-91269fc194f2 · outbound

This paper cites DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.364571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.364571Z digest=sha256:011bd092d79257bab07345d1a90c9416ca2e90e03a910eda679913df967939fa

Observation 222d1de1-a190-4c71-9c61-0062011222cd · outbound

This paper cites Customized retriev al augmented generation and benchmarking for EDA tool documen tation QA.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Customized retriev al augmented generation and benchmarking for EDA tool documen tation QA

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:52.194419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.369399Z digest=sha256:42f33a756ec5c660d575d187add1d5151280f51a1dcb74871489e5d19ba0fd9c

Observation 4649fa2f-5b3f-433e-9322-26de9bf3a533 · outbound

This paper cites Evaluation of Retrieval-Augmented Generation: A Survey.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluation of Retrieval-Augmented Generation: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.373822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.373822Z digest=sha256:66d206fddef17728c9783074bac8de5b013be6836630d164d02fb70fa3af6b1d

Observation 31900d05-d94d-4eba-888f-9a5cfc055c2b · outbound

This paper cites Benchmarking of retrieval augmented gene ration: A comprehensive systematic literature review on evaluatio n dimensions, evaluation metrics and datasets,.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Benchmarking of retrieval augmented gene ration: A comprehensive systematic literature review on evaluatio n dimensions, evaluation metrics and datasets,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:52.177982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.378843Z digest=sha256:4c6e5ad1895f808851840b75547ccc8ab6ca9cbdb7a274146658127f134dccfa

Observation c598f289-f925-4191-8d7e-1abdc0611479 · outbound

This paper cites Guidelines for perform ing systematic literature reviews in software engineering,.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Guidelines for perform ing systematic literature reviews in software engineering,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:52.164412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.384845Z digest=sha256:7bd0c5274215c1807096ab010430893f8a400c4afb53e03b394d6d33e21adb74

Observation 2d54bdb2-d04a-46c6-88bb-2f0b9b1c984d · outbound

This paper cites LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.389629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.389629Z digest=sha256:66ad67263445ab49d047a0e3ceb34878493d1ec4845a93e8abf3052a0210cac7

Observation c2ea125d-8cba-4a47-9ea9-2ccb7c7b084a · outbound

This paper cites Benchmarking Retrieval-Augmented Generation for Medicine.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Benchmarking Retrieval-Augmented Generation for Medicine

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.395011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.395011Z digest=sha256:7ef9ce187bc199fb3c4ecbeffbfefb0e4d39af8bd20f3ee211581f3d92a781fc

Observation 1c568a59-c57e-47db-92e1-ead6206945cf · outbound

This paper cites WeQA: A Benchmark for Retrieval Augmented Generation in Wind Energy Domain.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets WeQA: A Benchmark for Retrieval Augmented Generation in Wind Energy Domain

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-16T05:52:51.735818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.400147Z digest=sha256:d8388013691f092f92822412fda7b512be1fa5101bd4e8143da9237b7ee07aa3

Observation f09c2fc8-abf3-49d0-a589-d5b725207f05 · outbound

This paper cites an unresolved cited work.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:52:52.150219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.404935Z digest=sha256:30b8fce5796e38731c26960de082ad8eb275f8fb005f9719fc8dafa00392f0c0

Observation 70b752f4-32c9-4b6a-a9bf-b3289421f9e2 · outbound

This paper cites RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.409946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.409946Z digest=sha256:9a5d553d7a422e40f609b9362fca2f0a4f47ccfda5f0161bf9f05a645541dc65

Observation 528dcb6e-1b04-4567-a02e-235eb98cd349 · outbound

This paper cites Lynx: An Open Source Hallucination Evaluation Model.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Lynx: An Open Source Hallucination Evaluation Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.415581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.415581Z digest=sha256:0083078a52c62af80f87bd08c45178959ff8537e4f9bf112d1856c8ce7e2affd

Observation 17d82c36-4ea8-4aa7-ba5c-32476f4b41c4 · outbound

This paper cites Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.420495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.420495Z digest=sha256:e8760de6187a6a024a46b90dda5e287dec67014c0f94df7e78695a8f0ddbae68

Observation 4f85a63a-5ac9-456f-b9fe-d4ebdb5ecde8 · outbound

This paper cites ELOQ: Resources for Enhancing LLM Detection of Out-of-Scope Questions.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets ELOQ: Resources for Enhancing LLM Detection of Out-of-Scope Questions

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-16T05:52:51.665496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.426002Z digest=sha256:307344e14df69300c062c2c2b00a7061e11011f2293dfe6ec73c7301d1b100eb

Observation 5bcf4cb5-4836-4161-ab98-03647423129b · outbound

This paper cites Retrieval Augmented Generation Systems: Automatic Dataset Creation, Evaluation and Boolean Agent Setup.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Retrieval Augmented Generation Systems: Automatic Dataset Creation, Evaluation and Boolean Agent Setup

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-16T05:52:51.646049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.431940Z digest=sha256:9bc0a49f1370d19fbc1a2c9e86fb2801603f8cddf44a5345d9c7fa7b1440cff8

Observation 56e0374f-bf28-4473-a81d-f0a1073ab108 · outbound

This paper cites ClimRetrieve: A Benchmarking Dataset for Information Retrieval from Corporate Climate Disclosures.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets ClimRetrieve: A Benchmarking Dataset for Information Retrieval from Corporate Climate Disclosures

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.437094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.437094Z digest=sha256:fc1098f4dc1b9f682ac110f1ef0dc7444724c81f04f65c541353cbd2d7f71125

Observation b2fc9442-9968-4114-aba5-5599d87a2b86 · outbound

This paper cites RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.442866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.442866Z digest=sha256:23a2a2fa3bba10a61e4a6a9216fcb2d35675a9fe81d4a4710d22635a7e871ab9

Observation 329397d0-8c65-4954-bf5d-21cff6bafe4e · outbound

This paper cites Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.447905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.447905Z digest=sha256:9c8ca0cde070cf6177632c1be1890457cf0944a86bd0e5f8dd212c5c32f1fe91

Observation 8161b24b-4802-4582-a81d-ac58d95465b0 · outbound

This paper cites Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-16T05:52:51.563878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.452577Z digest=sha256:fecfad911fc8695b57fd595cfee7dbd4b833949fe57bfee143b8eb605ed3beb2

Observation 88d04068-5e49-4db6-9aec-b2fd03b14dce · outbound

This paper cites FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.457455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.457455Z digest=sha256:e844e7d131d5b693d9874e86f857fc16aa015f4f2fdc2b9ba90c7dfb228cc8e0

Observation 7fb67258-e95d-47ff-9d84-333c815b8a3b · outbound

This paper cites Long$^2$RAG: Evaluating Long-Context & Long-Form Retrieval-Augmented Generation with Key Point Recall.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Long$^2$RAG: Evaluating Long-Context & Long-Form Retrieval-Augmented Generation with Key Point Recall

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.462406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.462406Z digest=sha256:71791c5aa32233d8dfb1cd32bd5030ef61c256df8a1000d7e5932589e89ba174

Observation 5c912b17-48f1-494d-a217-2c0d989ddc56 · outbound

This paper cites CoFE-RAG: A Comprehensive Full-chain Evaluation Framework for Retrieval-Augmented Generation with Enhanced Data Diversity.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets CoFE-RAG: A Comprehensive Full-chain Evaluation Framework for Retrieval-Augmented Generation with Enhanced Data Diversity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.467121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.467121Z digest=sha256:5f6c6f3c1f7fff735c3b3e39ccc43d89d8e0881a6e87ddae8aacf2b9484ced8f

Observation 1774fc7b-69ba-4c3d-966c-1134d127121c · outbound

This paper cites Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.475100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.475100Z digest=sha256:d6eee326f88274fa7d930a958425aa433a21f56f43d2d2485b2ae65e5243d7c4

Observation 2ecba298-6c25-4802-bedb-e60fb3efa999 · outbound

This paper cites RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.482124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.482124Z digest=sha256:b7db66c55f17a333574fc74b4f7a7024fcdbe98fedaff51d18876a7b63a55303

Observation 302c887c-bd6c-4031-8cf4-d1c81a7bfc53 · outbound

This paper cites Multimodal re- trieval augmented generation evaluation benchmark,.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Multimodal re- trieval augmented generation evaluation benchmark,

Reference 34

Resolution
verified exact
raw_fallback, observed 2026-08-16T05:52:51.448336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.487181Z digest=sha256:1229fc273d536856bd090cc99ade710292ceb2f733cf3386aa48f6902556a6c7

Observation 7195bb04-fbbb-4c6e-a375-2ee7cd946089 · outbound

This paper cites Benchmarking Large Language Models in Retrieval-Augmented Generation.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Benchmarking Large Language Models in Retrieval-Augmented Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.492510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.492510Z digest=sha256:f70d3ff1f6464f50dfe54e850eb826b77a1b2a2b907b4175d3e22d82c6e47642

Observation ad4bc8f3-17c6-4878-aa9c-be5e52e03f2d · outbound

This paper cites FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows".

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.497905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.497905Z digest=sha256:c78145e4ba0e67f10c8fcbb10395791e24f56d1575ab4e41c9df0e41af3179c5

Observation 11ec0065-3897-4409-a431-8f516cf8210c · outbound

This paper cites A knowledge-cen tric benchmarking framework and empirical study for retrieval- augmented generation.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets A knowledge-cen tric benchmarking framework and empirical study for retrieval- augmented generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.502664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.502664Z digest=sha256:b3de7e56bc241a3c1fa138e6683a97229187ce965439f44966746595e872cbbb

Observation 21227e61-2181-4306-83d2-504c4cc10c8e · outbound

This paper cites Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.507117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.507117Z digest=sha256:3bb8394a4c615ff88bf9d20a6bc481157d4f6d406845b70d89ca0fddd11b70a9

Observation 1cf8f6ce-944a-4a3e-8e68-5c66ab3e0451 · outbound

This paper cites Enhancing Q&A Text Retrieval with Ranking Models: Benchmarking, fine-tuning and deploying Rerankers for RAG.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Enhancing Q&A Text Retrieval with Ranking Models: Benchmarking, fine-tuning and deploying Rerankers for RAG

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.512404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.512404Z digest=sha256:2eac65f577891546ad68bb153bc493344f7fd047c69524e84cdca3f5a9eff262

Observation 5bddf35e-797d-4b25-814b-5b1f8bc23558 · outbound

This paper cites CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.518272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.518272Z digest=sha256:055fd90d02c23ce6e6343cc236b03156c64c6c1d5961747f55ca30a7d99021ed

Observation c129df0a-dbf9-4101-926e-8ef43fa8674d · outbound

This paper cites IRSC: A Zero-shot Evaluation Benchmark for Information Retrieval through Semantic Comprehension in Retrieval-Augmented Generation Scenarios.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets IRSC: A Zero-shot Evaluation Benchmark for Information Retrieval through Semantic Comprehension in Retrieval-Augmented Generation Scenarios

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-16T05:52:51.200754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.523580Z digest=sha256:5e610043449c8dff1f9ca32277f199e23b8d4c43518a746e64e472ef66294504

Observation 0e88ca3d-683d-4cad-ab11-859631eeb11d · outbound

This paper cites Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.529508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.529508Z digest=sha256:2d9ba7bd05b4fdae255d1d7c85538c102315fd6f2ef62b80d2d594fbed587fd5

Observation 8c3ef1f8-77ff-4c04-a5fc-e60e00bb6253 · outbound

This paper cites UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.534038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.534038Z digest=sha256:ab4656cb3fdd1ccf67cd3acc9e97fd59b9f08bb133852ea04abe1326dad7abc3

Observation 585f5aa8-a99f-48b3-ad4d-5952571535be · outbound

This paper cites Evaluating RAG-Fusion with RAGElo: an Automated Elo-based Framework.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluating RAG-Fusion with RAGElo: an Automated Elo-based Framework

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.538867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.538867Z digest=sha256:7189bb1bd1338facde4cf8876023fe7647c625a1a77ea77bd19143a2c27a43ea

Observation 2de3d60c-1d92-4164-9944-e00e1d8e0be5 · outbound

This paper cites RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.543996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.543996Z digest=sha256:769cc697a6c9b95f806c08bf3e603a7c7599846c01bb4f36f9a778a46d8279be

Observation ba24deba-a572-403b-9415-aa5f05bded64 · outbound

This paper cites VERA: Validation and Evaluation of Retrieval-Augmented Systems.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets VERA: Validation and Evaluation of Retrieval-Augmented Systems

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.548589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.548589Z digest=sha256:79cdc33d0d2e99375871e5a17b7d7144a39383c47aa648284f2521280abe63bf

Observation ec33de68-b87c-4551-abaa-33da5318b6ca · outbound

This paper cites Evaluating the Retrieval Component in LLM-Based Question Answering Systems.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluating the Retrieval Component in LLM-Based Question Answering Systems

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.553454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.553454Z digest=sha256:c17bd9522bda7886dea892f1d510ab3b28593cddfef9edfd318592d54478d20c

Observation 39bea0a5-e93c-4509-929e-2ab72754cdd1 · outbound

This paper cites FaaF: Facts as a Function for the evaluation of generated text.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets FaaF: Facts as a Function for the evaluation of generated text

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.559787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.559787Z digest=sha256:f2a0eda145686fbcdcc3efaa17d99ba83413f941b52bdaacaad7344e0efab191

Observation 3810b43b-48e0-4066-9eb7-f71c72fc9a47 · outbound

This paper cites A Methodology for Evaluating RAG Systems: A Case Study On Configuration Dependency Validation.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets A Methodology for Evaluating RAG Systems: A Case Study On Configuration Dependency Validation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.565643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.565643Z digest=sha256:a7c998d1853d6f8c9757f301d3913781f5e4ef59292200a63350785242444085

Observation 757a155d-ae07-4ede-98a0-602c866bdc70 · outbound

This paper cites BERGEN: A Benchmarking Library for Retrieval-Augmented Generation.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets BERGEN: A Benchmarking Library for Retrieval-Augmented Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.571397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.571397Z digest=sha256:1c23c752646c4986f910879844717562464004c71ee6edb8d2d24cccaff3f813

Observation 08c49097-cf5e-40ef-9e54-db072fa65db5 · outbound

This paper cites Evaluating lar ge language models for arabic sentiment analysis: A comparati ve study using retrieval-augmented generation,.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluating lar ge language models for arabic sentiment analysis: A comparati ve study using retrieval-augmented generation,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:52.135952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.577667Z digest=sha256:a5ffe4982b1f91aab54c94b860be41096e7ac6ef717c8ac562bcf9d00e03baff

Observation a3ce4fb8-9f2b-4f88-a7a9-0c89a5d76270 · outbound

This paper cites ReEval: Automatic hallucination evaluation for retrieval-augmen ted large language models via transferable adversarial attacks,.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets ReEval: Automatic hallucination evaluation for retrieval-augmen ted large language models via transferable adversarial attacks,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:52.113985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.583168Z digest=sha256:86bba4b9ad1a73b47c0f6953db7c584f43170411d5511e8a5ef9282d50fa7e3a

Observation 78fc351b-0be1-4490-98a0-4d3620a01469 · outbound

This paper cites Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam Generation.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam Generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.587880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.587880Z digest=sha256:4a86685ef272b5a3b20840e21a63182f5d9707548f860d4243d759ff8910708e

Observation 99c2ea8c-678c-4857-8c58-56dff351cb75 · outbound

This paper cites CRAG -- Comprehensive RAG Benchmark.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets CRAG -- Comprehensive RAG Benchmark

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.594834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.594834Z digest=sha256:74550c2b6060cc4ab6e82268dd98bfc51277c59fe0d47f95f0e786cde5815f8b

Observation 1e480352-fdb3-4fd9-9b21-ff842d616c8d · outbound

This paper cites Automatic questi on answering for the linguistic domain – an evaluation of LLM knowledge ba se extension with RAG,.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Automatic questi on answering for the linguistic domain – an evaluation of LLM knowledge ba se extension with RAG,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:52.097013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.601436Z digest=sha256:dba640aac16459b2c54655bd6d2b6b633230b91fcf3faf3c0bf9e6152ac7b2ca

Observation ee5418b3-67ec-40e7-9a9a-5fa8faa3f34e · outbound

This paper cites Should We Fine-Tune or RAG? Evaluating Different Techniques to Adapt LLMs for Dialogue.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Should We Fine-Tune or RAG? Evaluating Different Techniques to Adapt LLMs for Dialogue

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-16T05:52:50.998056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.606011Z digest=sha256:220be2ebed6ce238ec130ad7f63c53a45dfe3d502a7c844cbd4235ef0178e93a

Observation 12ad363b-eb04-463d-b03b-51d12e6c5d51 · outbound

This paper cites Benchmarking retrieval augme nted genera- tion in quantitative finance,.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Benchmarking retrieval augme nted genera- tion in quantitative finance,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:52.078932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.611136Z digest=sha256:10f73fcf026132a4012e4cd5406479685ed77cdf165cb4d97bad0db6fadce7d6

Observation f495307d-e889-4653-a3f6-362212845869 · outbound

This paper cites an unresolved cited work.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:52:52.061193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.615167Z digest=sha256:ef44eafea0224928c7afd181bcc5c3af272d3677b2d6562666f2b4caadbfa1b7

Observation ae415f85-0fc8-427f-9958-03f8fabe28f7 · outbound

This paper cites Evaluating Quality of Answers for Retrieval-Augmented Generation: A Strong LLM Is All You Need.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluating Quality of Answers for Retrieval-Augmented Generation: A Strong LLM Is All You Need

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.619509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.619509Z digest=sha256:f33f8383190d8f2bed1b587c20431202efce7465d5c1c0834f516ea4432c8e6a

Observation 5d49749a-4131-49d1-8729-74cc71a8e6fe · outbound

This paper cites MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.625795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.625795Z digest=sha256:3f165c2b907591a8df7878a7d2963951782afb96db0d7bc8544ba8d9865c8ffc

Observation 4764630f-bec1-4f0a-bc9f-3d31dfa11ba9 · outbound

This paper cites Evaluating the Efficacy of Open-Source LLMs in Enterprise-Specific RAG Systems: A Comparative Study of Performance and Scalability.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluating the Efficacy of Open-Source LLMs in Enterprise-Specific RAG Systems: A Comparative Study of Performance and Scalability

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.631621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.631621Z digest=sha256:76226af7e9063b00f0b72a6232538900449623bfa0c8b08737fdb05798fce869

Observation b796f565-4997-4db8-b319-27dba7745f4a · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understan ding,.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets BERT: Pre-training of deep bidirectional transformers for language understan ding,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:52.044193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.636369Z digest=sha256:be3c06e7c5ba7777c1e8aa16b2c01482e7dbf5474f2385fe47c032c802f960f5

Observation 2a4accce-5f40-4819-bd6b-87fb3b6afcad · outbound

This paper cites Available: http://arxiv.org/abs/1810.0480 5.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Available: http://arxiv.org/abs/1810.0480 5

Reference 63

Resolution
verified exact
raw_fallback, observed 2026-08-16T05:52:50.923054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.640689Z digest=sha256:a11ec991d6be287521a53f84eaa68b0d7bd5dcddc09df2d46772f13f1ff30df9

Observation d1282e4a-4315-4d5e-867f-4957748d6e58 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.644874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.644874Z digest=sha256:27aa7affb4332e73595ed9f14dc3337059f6e4b80848e906862471683b89c2ce

Observation 248157ed-26fb-4d5c-9058-47a950d09450 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets BERTScore: Evaluating Text Generation with BERT

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.649252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.649252Z digest=sha256:5aef67a98e47d41806f3d7e2386882305c704f1c422431bcfc83b32fd6915307

Observation 69d2c35a-ad0d-4dc0-b2eb-2ebf193af01a · outbound

This paper cites Towards a Unified Multi-Dimensional Evaluator for Text Generation.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Towards a Unified Multi-Dimensional Evaluator for Text Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.653425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.653425Z digest=sha256:eafb2076b44f718cfe3839ac0d4d9c13bf1e130c2828bc66b1b44d2f22e20d09

Observation 5e2ada04-eafe-4b87-8235-6a879d8215aa · outbound

This paper cites [Online].

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets [Online]

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:52.029329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.659121Z digest=sha256:f28c7dceb9f637c1c46c9333fd40a1f4a9c3f81df5361a8aa5cfcc8658ae4af8

Observation e214b3a4-4bca-4ef3-b9e7-5ba3a0a558df · outbound

This paper cites Evaluation of orca 2 against other LLMs for retrieval augmented generation,.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluation of orca 2 against other LLMs for retrieval augmented generation,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:52.013493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.663099Z digest=sha256:d81855fb2e9021b3d529845238e757bcea509583df8aca2151487ddc82dfc2c2

Observation 037d4ad9-20fd-4f75-85d7-27851b8d649f · outbound

This paper cites RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.668532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.668532Z digest=sha256:9dd2a42fba261460e8bd7ce2f21731b5aaff25ca537eaab55295c745ef7ff48e

Observation 2d0c1ad4-d94f-428c-92e0-ba674e36cdb5 · outbound

This paper cites RAGProbe: An Automated Approach for Evaluating RAG Applications.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets RAGProbe: An Automated Approach for Evaluating RAG Applications

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:50.673328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:50.673328Z digest=sha256:0200721f1490b355e4fbefd850923d2be1803f612cbcd907e6f84c66f44f7fce

Observation 356a5e99-87fe-4f18-83a6-bdbf0ac13744 · outbound

This paper cites Intrinsic Evaluation of RAG Systems for Deep-Logic Questions.

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Intrinsic Evaluation of RAG Systems for Deep-Logic Questions

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-16T05:52:50.730844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.677948Z digest=sha256:3c72809758e127b845cd36d224fb6c68b858540d84dbbf0f4065615575d118c4

Observation fc822c81-0989-41fc-8227-947ea4baa4ce · outbound

This paper cites [Online].

Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets [Online]

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:52:51.993401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T05:52:50.682341Z digest=sha256:b41720c9ead6f460f12d67b558214a20b358db8849c795ef0bc04bd2494fc6af

Pith citing papers

Observation 065b5508-60fc-413d-9e7c-f652b27a5e9f · inbound

A Survey of Context Engineering for Large Language Models cites this paper.

A Survey of Context Engineering for Large Language Models Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:45.173586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T20:58:45.060041Z digest=sha256:529654dc3fc1e4a842066225aa59c38e112cbc848bdf361a84a1cff7d41a65d3

Observation 8e119205-8ef6-4d02-9bd1-3ce6d6dfaf79 · inbound

Rewrite-to-Rank: Optimizing Ad Visibility via Retrieval-Aware Text Rewriting cites this paper.

Rewrite-to-Rank: Optimizing Ad Visibility via Retrieval-Aware Text Rewriting Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:30.304142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:30.304142Z digest=sha256:f893a989949776a68fa05bc615cba469675a494270b3ccffe07ed67a18f078b4