Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:52:50.682341Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 2 inbound Pith citation observations for arXiv:2504.20119.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:52:50.682341Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:39:30.304142Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-13T20:58:45.171226Z
72 of 72 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f34183de-1a7b-4400-8d65-9579e70d1bfe · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Retrieval-Augmented Generation for Large Language Models: A Survey
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 760f13f0-68a5-48fb-95e7-c919bd96b33e · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Si ren’s Song in the AI Ocean: A Survey on Hallucination in Large Langu age Models,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 664fcab1-507d-4238-8aa8-5ad4f7ba42e7 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf273c94-36fe-497d-949c-4126e62d3c64 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets CRUD-RAG: A comprehensive chines e benchmark for retrieval-augmented generation of large lan guage models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4d9c402b-f066-4d11-b359-ff0ee5633b6e · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Performance evaluation of vector embeddings with retrieval- augmented generation,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5ea4758c-fc90-43f0-82b2-6e0c427eb72c · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Ragas: Automated Evaluation of Retrieval Augmented Generation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abc963ba-231c-4279-8b8d-e507e433487e · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets AR ES: An automated evaluation framework for retrieval-augmented g eneration systems
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 433b7aa9-45aa-446f-aee8-a5d9bcb1e677 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluating Retrieval Quality in Retrieval-Augmented Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e54e03-e717-43b0-bc5f-257ab86bb0a2 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Beyond Benchmarks: Evaluating Embedding Model Similarity for Retrieval Augmented Generation Systems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d407466-2bfd-41fb-80d6-07fd166235ea · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 863e1ce2-011b-4fb1-8a21-91269fc194f2 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 222d1de1-a190-4c71-9c61-0062011222cd · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Customized retriev al augmented generation and benchmarking for EDA tool documen tation QA
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4649fa2f-5b3f-433e-9322-26de9bf3a533 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluation of Retrieval-Augmented Generation: A Survey
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31900d05-d94d-4eba-888f-9a5cfc055c2b · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Benchmarking of retrieval augmented gene ration: A comprehensive systematic literature review on evaluatio n dimensions, evaluation metrics and datasets,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c598f289-f925-4191-8d7e-1abdc0611479 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Guidelines for perform ing systematic literature reviews in software engineering,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2d54bdb2-d04a-46c6-88bb-2f0b9b1c984d · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2ea125d-8cba-4a47-9ea9-2ccb7c7b084a · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Benchmarking Retrieval-Augmented Generation for Medicine
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c568a59-c57e-47db-92e1-ead6206945cf · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets WeQA: A Benchmark for Retrieval Augmented Generation in Wind Energy Domain
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f09c2fc8-abf3-49d0-a589-d5b725207f05 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 70b752f4-32c9-4b6a-a9bf-b3289421f9e2 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 528dcb6e-1b04-4567-a02e-235eb98cd349 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Lynx: An Open Source Hallucination Evaluation Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17d82c36-4ea8-4aa7-ba5c-32476f4b41c4 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f85a63a-5ac9-456f-b9fe-d4ebdb5ecde8 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets ELOQ: Resources for Enhancing LLM Detection of Out-of-Scope Questions
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5bcf4cb5-4836-4161-ab98-03647423129b · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Retrieval Augmented Generation Systems: Automatic Dataset Creation, Evaluation and Boolean Agent Setup
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 56e0374f-bf28-4473-a81d-f0a1073ab108 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets ClimRetrieve: A Benchmarking Dataset for Information Retrieval from Corporate Climate Disclosures
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2fc9442-9968-4114-aba5-5599d87a2b86 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 329397d0-8c65-4954-bf5d-21cff6bafe4e · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8161b24b-4802-4582-a81d-ac58d95465b0 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 88d04068-5e49-4db6-9aec-b2fd03b14dce · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb67258-e95d-47ff-9d84-333c815b8a3b · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Long$^2$RAG: Evaluating Long-Context & Long-Form Retrieval-Augmented Generation with Key Point Recall
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c912b17-48f1-494d-a217-2c0d989ddc56 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets CoFE-RAG: A Comprehensive Full-chain Evaluation Framework for Retrieval-Augmented Generation with Enhanced Data Diversity
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1774fc7b-69ba-4c3d-966c-1134d127121c · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ecba298-6c25-4802-bedb-e60fb3efa999 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 302c887c-bd6c-4031-8cf4-d1c81a7bfc53 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Multimodal re- trieval augmented generation evaluation benchmark,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7195bb04-fbbb-4c6e-a375-2ee7cd946089 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Benchmarking Large Language Models in Retrieval-Augmented Generation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad4bc8f3-17c6-4878-aa9c-be5e52e03f2d · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11ec0065-3897-4409-a431-8f516cf8210c · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets A knowledge-cen tric benchmarking framework and empirical study for retrieval- augmented generation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21227e61-2181-4306-83d2-504c4cc10c8e · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cf8f6ce-944a-4a3e-8e68-5c66ab3e0451 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Enhancing Q&A Text Retrieval with Ranking Models: Benchmarking, fine-tuning and deploying Rerankers for RAG
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bddf35e-797d-4b25-814b-5b1f8bc23558 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c129df0a-dbf9-4101-926e-8ef43fa8674d · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets IRSC: A Zero-shot Evaluation Benchmark for Information Retrieval through Semantic Comprehension in Retrieval-Augmented Generation Scenarios
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0e88ca3d-683d-4cad-ab11-859631eeb11d · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c3ef1f8-77ff-4c04-a5fc-e60e00bb6253 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 585f5aa8-a99f-48b3-ad4d-5952571535be · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluating RAG-Fusion with RAGElo: an Automated Elo-based Framework
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2de3d60c-1d92-4164-9944-e00e1d8e0be5 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba24deba-a572-403b-9415-aa5f05bded64 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets VERA: Validation and Evaluation of Retrieval-Augmented Systems
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec33de68-b87c-4551-abaa-33da5318b6ca · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluating the Retrieval Component in LLM-Based Question Answering Systems
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39bea0a5-e93c-4509-929e-2ab72754cdd1 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets FaaF: Facts as a Function for the evaluation of generated text
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3810b43b-48e0-4066-9eb7-f71c72fc9a47 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets A Methodology for Evaluating RAG Systems: A Case Study On Configuration Dependency Validation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 757a155d-ae07-4ede-98a0-602c866bdc70 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets BERGEN: A Benchmarking Library for Retrieval-Augmented Generation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08c49097-cf5e-40ef-9e54-db072fa65db5 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluating lar ge language models for arabic sentiment analysis: A comparati ve study using retrieval-augmented generation,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a3ce4fb8-9f2b-4f88-a7a9-0c89a5d76270 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets ReEval: Automatic hallucination evaluation for retrieval-augmen ted large language models via transferable adversarial attacks,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 78fc351b-0be1-4490-98a0-4d3620a01469 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam Generation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c2ea8c-678c-4857-8c58-56dff351cb75 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets CRAG -- Comprehensive RAG Benchmark
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e480352-fdb3-4fd9-9b21-ff842d616c8d · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Automatic questi on answering for the linguistic domain – an evaluation of LLM knowledge ba se extension with RAG,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ee5418b3-67ec-40e7-9a9a-5fa8faa3f34e · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Should We Fine-Tune or RAG? Evaluating Different Techniques to Adapt LLMs for Dialogue
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 12ad363b-eb04-463d-b03b-51d12e6c5d51 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Benchmarking retrieval augme nted genera- tion in quantitative finance,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f495307d-e889-4653-a3f6-362212845869 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ae415f85-0fc8-427f-9958-03f8fabe28f7 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluating Quality of Answers for Retrieval-Augmented Generation: A Strong LLM Is All You Need
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d49749a-4131-49d1-8729-74cc71a8e6fe · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4764630f-bec1-4f0a-bc9f-3d31dfa11ba9 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluating the Efficacy of Open-Source LLMs in Enterprise-Specific RAG Systems: A Comparative Study of Performance and Scalability
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b796f565-4997-4db8-b319-27dba7745f4a · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets BERT: Pre-training of deep bidirectional transformers for language understan ding,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2a4accce-5f40-4819-bd6b-87fb3b6afcad · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Available: http://arxiv.org/abs/1810.0480 5
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d1282e4a-4315-4d5e-867f-4957748d6e58 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 248157ed-26fb-4d5c-9058-47a950d09450 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets BERTScore: Evaluating Text Generation with BERT
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69d2c35a-ad0d-4dc0-b2eb-2ebf193af01a · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Towards a Unified Multi-Dimensional Evaluator for Text Generation
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e2ada04-eafe-4b87-8235-6a879d8215aa · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets [Online]
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e214b3a4-4bca-4ef3-b9e7-5ba3a0a558df · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Evaluation of orca 2 against other LLMs for retrieval augmented generation,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 037d4ad9-20fd-4f75-85d7-27851b8d649f · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d0c1ad4-d94f-428c-92e0-ba674e36cdb5 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets RAGProbe: An Automated Approach for Evaluating RAG Applications
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 356a5e99-87fe-4f18-83a6-bdbf0ac13744 · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets Intrinsic Evaluation of RAG Systems for Deep-Logic Questions
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fc822c81-0989-41fc-8227-947ea4baa4ce · outbound
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets [Online]
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 065b5508-60fc-413d-9e7c-f652b27a5e9f · inbound
A Survey of Context Engineering for Large Language Models Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8e119205-8ef6-4d02-9bd1-3ce6d6dfaf79 · inbound
Rewrite-to-Rank: Optimizing Ad Visibility via Retrieval-Aware Text Rewriting Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.