Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 59 inbound Pith citation observations for arXiv:2305.01937.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:56:40.104221Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
33
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 814939cf-b7f6-4969-91c5-9e1bec2f57ff · inbound
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Can Large Language Models Be an Alternative to Human Evaluations?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6a5e10f4-5fc1-42c1-9d2c-97dd29ea5f00 · inbound
DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models Can Large Language Models Be an Alternative to Human Evaluations?
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 858940d4-8f27-42d8-a754-5512b4e099b5 · inbound
A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions Can Large Language Models Be an Alternative to Human Evaluations?
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1587904f-5c1b-4a5b-85f6-e4ca2379b93e · inbound
Instruction-Following Evaluation for Large Language Models Can Large Language Models Be an Alternative to Human Evaluations?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 91bdb0e5-5a3c-4906-9c5b-d5406d963c07 · inbound
The Falcon Series of Open Language Models Can Large Language Models Be an Alternative to Human Evaluations?
Reference 260
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3fa4798e-cce5-4d98-9c6a-7e808df44b3b · inbound
Data-Centric Foundation Models in Computational Healthcare: A Survey Can Large Language Models Be an Alternative to Human Evaluations?
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c30be60d-47ca-4eb8-bd7e-a1d0dfa99e6d · inbound
Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness Can Large Language Models Be an Alternative to Human Evaluations?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 628ca227-a898-4afd-9e45-72f0461c960c · inbound
SRSA: A Cost-Efficient Strategy-Router Search Agent for Real-world Human-Machine Interactions Can Large Language Models Be an Alternative to Human Evaluations?
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc76682c-2a6e-4f13-8195-8cc0f480d730 · inbound
Do LLMs Agree on the Creativity Evaluation of Alternative Uses? Can Large Language Models Be an Alternative to Human Evaluations?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7038c9c-3323-4c76-8d93-48edf064a7e2 · inbound
ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation Can Large Language Models Be an Alternative to Human Evaluations?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24059aa6-2376-47a7-9ebe-1ebf7aeeb741 · inbound
Human-Like Code Quality Evaluation through LLM-based Recursive Semantic Comprehension Can Large Language Models Be an Alternative to Human Evaluations?
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 035363bf-f34a-4f15-9bfe-707d3331831b · inbound
Can Large Language Models Serve as Evaluators for Code Summarization? Can Large Language Models Be an Alternative to Human Evaluations?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c9a7192-0623-4bfc-b300-bbd626bb44d8 · inbound
Multi-Facet Blending for Faceted Query-by-Example Retrieval Can Large Language Models Be an Alternative to Human Evaluations?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f030f7-35dc-4570-97ee-63008b05c85a · inbound
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Can Large Language Models Be an Alternative to Human Evaluations?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef857637-bd29-48a3-be1b-afe3df0fbd8c · inbound
The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance? Can Large Language Models Be an Alternative to Human Evaluations?
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9074c8f0-f0f8-4576-a9f0-12d9021e1673 · inbound
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models Can Large Language Models Be an Alternative to Human Evaluations?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 803fb350-3e91-419d-bb69-400314a11faf · inbound
Generative Adversarial Reviews: When LLMs Become the Critic Can Large Language Models Be an Alternative to Human Evaluations?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30aa11da-13e1-4ed8-ab8d-c1c7a56d4f0e · inbound
EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Can Large Language Models Be an Alternative to Human Evaluations?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90df8baa-0b88-4629-973b-b0de585f97bf · inbound
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Can Large Language Models Be an Alternative to Human Evaluations?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21417018-4d95-476d-aa08-613aab922500 · inbound
The Impostor is Among Us: Can Large Language Models Capture the Complexity of Human Personas? Can Large Language Models Be an Alternative to Human Evaluations?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c757a55-1016-4b16-9f7b-7da69d37863e · inbound
Can Large Language Models Predict the Outcome of Judicial Decisions? Can Large Language Models Be an Alternative to Human Evaluations?
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0fe5bfd-b6eb-4b82-9d8e-d5ee28e22c34 · inbound
Automated Assignment Grading with Large Language Models: Insights From a Bioinformatics Course Can Large Language Models Be an Alternative to Human Evaluations?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d837da3-c0d3-443b-909c-2ddfa06d830f · inbound
Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline Can Large Language Models Be an Alternative to Human Evaluations?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2044898d-6f81-4aa9-bd60-dcbe72f24209 · inbound
A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs Can Large Language Models Be an Alternative to Human Evaluations?
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbdd4d8d-e2d6-4a3b-a199-1cbd65dcac8e · inbound
Stop Overvaluing Multi-Agent Debate -- We Must Rethink Evaluation and Embrace Model Heterogeneity Can Large Language Models Be an Alternative to Human Evaluations?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ea38642-fee4-454e-b912-af096972d4fa · inbound
Image Embedding Sampling Method for Diverse Captioning Can Large Language Models Be an Alternative to Human Evaluations?
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9799cb4-970b-4e35-ba8e-c3918c478614 · inbound
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Can Large Language Models Be an Alternative to Human Evaluations?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f1316c7-326e-40a5-b1d3-2277807b866d · inbound
OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models Can Large Language Models Be an Alternative to Human Evaluations?
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65579d38-8691-4034-aace-16374b079272 · inbound
OSS-UAgent: An Agent-based Usability Evaluation Framework for Open Source Software Can Large Language Models Be an Alternative to Human Evaluations?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1933a8e-b3d0-405c-be0a-094c8a48f047 · inbound
How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG Can Large Language Models Be an Alternative to Human Evaluations?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7578f62b-ed4b-4f24-bfe9-852e61ef515b · inbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Can Large Language Models Be an Alternative to Human Evaluations?
Reference 195
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54074dbd-16a8-45eb-a163-1ff7bc0d7026 · inbound
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics Can Large Language Models Be an Alternative to Human Evaluations?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59cc03f1-7b1c-493d-99d9-46e76a6d7f5b · inbound
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation Can Large Language Models Be an Alternative to Human Evaluations?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e38fc80c-7501-41f6-b8ba-bf5a816c9c98 · inbound
Counterfactual Voting Adjustment for Quality Assessment and Fairer Voting in Online Platforms with Helpfulness Evaluation Can Large Language Models Be an Alternative to Human Evaluations?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec2ed023-6672-463e-b71c-a4a9d244423d · inbound
CitySim: Modeling Urban Behaviors and City Dynamics with Large-Scale LLM-Driven Agent Simulation Can Large Language Models Be an Alternative to Human Evaluations?
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1d6d28a-d741-4c4b-8ab3-f1fac97bed19 · inbound
Balancing Information Accuracy and Response Timeliness in Networked LLMs Can Large Language Models Be an Alternative to Human Evaluations?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f2966d4-d7e1-4d72-a744-79d8b4c41f34 · inbound
Generative Artificial Intelligence Extracts Structure-Function Relationships from Plants for New Materials Can Large Language Models Be an Alternative to Human Evaluations?
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 170f9278-a099-449f-8b80-b2227c26e93e · inbound
Rule2Text: A Framework for Generating and Evaluating Natural Language Explanations of Knowledge Graph Rules Can Large Language Models Be an Alternative to Human Evaluations?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e55e6375-5023-4525-930f-9762a6c464aa · inbound
Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation Can Large Language Models Be an Alternative to Human Evaluations?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53dfcc92-1497-46ac-b364-d3a2c9f783b0 · inbound
LaQual: An Automated Framework for LLM App Quality Evaluation Can Large Language Models Be an Alternative to Human Evaluations?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65ab60d8-40af-4688-830c-3bfff04da3a1 · inbound
Integrating Generative AI into Cybersecurity Education: A Study of OCR and Multimodal LLM-assisted Instruction Can Large Language Models Be an Alternative to Human Evaluations?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0b98ba9-1358-48cd-b90a-8588516087d6 · inbound
AdsQA: Towards Advertisement Video Understanding Can Large Language Models Be an Alternative to Human Evaluations?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14006fc8-2a4b-4f0d-952e-2f5837b31537 · inbound
Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration Can Large Language Models Be an Alternative to Human Evaluations?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61632217-2d50-409d-b6e5-6a8d682ed4cc · inbound
RESBev: Making BEV Perception More Robust Can Large Language Models Be an Alternative to Human Evaluations?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a41f30db-f2bb-4359-a9c6-29dfe2f88bb4 · inbound
The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning Can Large Language Models Be an Alternative to Human Evaluations?
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1852c7c6-a82c-4588-971c-8b94274622cf · inbound
Agent-Agnostic Evaluation of SQL Accuracy in Production Text-to-SQL Systems Can Large Language Models Be an Alternative to Human Evaluations?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9832c778-716b-484b-b9f1-5cc675053839 · inbound
Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement Can Large Language Models Be an Alternative to Human Evaluations?
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 677c1e49-b3a3-4693-b411-37b7c88188ff · inbound
Margin-Adaptive Confidence Ranking for Reliable LLM Judgement Can Large Language Models Be an Alternative to Human Evaluations?
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5db47bd7-7218-4732-b59d-ef9a38d09c68 · inbound
Benchmark Everything Everywhere All at Once Can Large Language Models Be an Alternative to Human Evaluations?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 60e96050-0a3d-431a-b6aa-d38f7ce63405 · inbound
StanceNakba Shared Task: Actor and Topic-Aware Stance Detection in Public Discourse Can Large Language Models Be an Alternative to Human Evaluations?
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0276a202-001e-4943-b76e-e5f0ab254512 · inbound
Iterating Toward Better Search: A Two-Agent Simulation Framework for Evaluating Agentic Search Architectures in E-Commerce Can Large Language Models Be an Alternative to Human Evaluations?
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7adfcf79-b2b2-488c-8df2-3c3ed7903b59 · inbound
Game Theory Driven Multi-Agent Framework Mitigates Language Model Hallucination Can Large Language Models Be an Alternative to Human Evaluations?
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 421e9056-0d74-47db-801a-056fae582872 · inbound
Scaling Point-in-Time Language Models Can Large Language Models Be an Alternative to Human Evaluations?
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87af01b4-7f64-4967-b3e3-62c7fad841f3 · inbound
Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models Can Large Language Models Be an Alternative to Human Evaluations?
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbe6a12e-e9cb-41b0-b30d-e0229ad69db4 · inbound
Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable Can Large Language Models Be an Alternative to Human Evaluations?
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22d0fa57-b970-4538-8cce-554746fac9ca · inbound
(Towards) Scalable Reliable Automated Evaluation with Large Language Models Can Large Language Models Be an Alternative to Human Evaluations?
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1511161-fd7d-413a-83e0-d220d59c6e99 · inbound
TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation Can Large Language Models Be an Alternative to Human Evaluations?
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d72eedcd-f0f0-402f-b029-4a84f6caaacc · inbound
VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP Can Large Language Models Be an Alternative to Human Evaluations?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f360cad3-a1a7-4b5c-820d-9cabe5ff4e23 · inbound
Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Can Large Language Models Be an Alternative to Human Evaluations?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.