Pith. sign in

Paper Citation Record · LEDGER

Can Large Language Models Be an Alternative to Human Evaluations?

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 59 inbound Pith citation observations for arXiv:2305.01937.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.01937 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:56:40.104221Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

33
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 814939cf-b7f6-4969-91c5-9e1bec2f57ff · inbound

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena cites this paper.

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena Can Large Language Models Be an Alternative to Human Evaluations?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:52:59.159832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T18:52:59.033645Z digest=sha256:0352b194a93906860749a3a22d1a7727d73cbba5d84b1fa1faa231bd94221ee3

Observation 6a5e10f4-5fc1-42c1-9d2c-97dd29ea5f00 · inbound

DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models cites this paper.

DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models Can Large Language Models Be an Alternative to Human Evaluations?

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:21:24.419213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-16T13:21:24.297836Z digest=sha256:fba63df24c657d6639c57d023a8a8054f4d66365c72aa0f129e80894146196f9

Observation 858940d4-8f27-42d8-a754-5512b4e099b5 · inbound

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions cites this paper.

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions Can Large Language Models Be an Alternative to Human Evaluations?

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:46:27.484024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T02:46:26.957539Z digest=sha256:1233c4afc12c1c538f3aa4f8fe5cf9343683a747fda791edcb27d08bf820be57

Observation 1587904f-5c1b-4a5b-85f6-e4ca2379b93e · inbound

Instruction-Following Evaluation for Large Language Models cites this paper.

Instruction-Following Evaluation for Large Language Models Can Large Language Models Be an Alternative to Human Evaluations?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:36:00.925197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-24T05:34:04.002648Z digest=sha256:b3320e3c59c5e9ab45b363c8e5bcf606cce61abe9cb76751624819a1ca207a3b

Observation 91bdb0e5-5a3c-4906-9c5b-d5406d963c07 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Can Large Language Models Be an Alternative to Human Evaluations?

Reference 260

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:46:10.036499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:905c992c7e668626755b10e6c1a23babff19a8688ca7b7e5e01ad28a23fc2074

Observation 3fa4798e-cce5-4d98-9c6a-7e808df44b3b · inbound

Data-Centric Foundation Models in Computational Healthcare: A Survey cites this paper.

Data-Centric Foundation Models in Computational Healthcare: A Survey Can Large Language Models Be an Alternative to Human Evaluations?

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:13:53.198466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-24T04:13:05.328492Z digest=sha256:02d252cd0eb3185a7ca8fe77923e2e31ba22023ff3eb7df2be096ec2a019acde

Observation c30be60d-47ca-4eb8-bd7e-a1d0dfa99e6d · inbound

Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness cites this paper.

Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness Can Large Language Models Be an Alternative to Human Evaluations?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T21:16:25.212599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:16:25.212599Z digest=sha256:0ac286a62b851ffecf8c8e6af7038667d2f123816d09c61e759fa1cb60e80ab4

Observation 628ca227-a898-4afd-9e45-72f0461c960c · inbound

SRSA: A Cost-Efficient Strategy-Router Search Agent for Real-world Human-Machine Interactions cites this paper.

SRSA: A Cost-Efficient Strategy-Router Search Agent for Real-world Human-Machine Interactions Can Large Language Models Be an Alternative to Human Evaluations?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:12:03.398720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:12:03.398720Z digest=sha256:73a624da5c9c7f05e55a698533dbb030491221d9951799712f579fa97bf40922

Observation cc76682c-2a6e-4f13-8195-8cc0f480d730 · inbound

Do LLMs Agree on the Creativity Evaluation of Alternative Uses? cites this paper.

Do LLMs Agree on the Creativity Evaluation of Alternative Uses? Can Large Language Models Be an Alternative to Human Evaluations?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:12:47.277502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:12:47.277502Z digest=sha256:88907714165b619a18478a720468845f16253a6655a05eb2c320efcd34dc5bbe

Observation c7038c9c-3323-4c76-8d93-48edf064a7e2 · inbound

ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation cites this paper.

ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation Can Large Language Models Be an Alternative to Human Evaluations?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:21:08.583992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:21:08.583992Z digest=sha256:855021bc1b076bd62f708c984e626cc608ca537eb18ad17a73cc14346f2f5042

Observation 24059aa6-2376-47a7-9ebe-1ebf7aeeb741 · inbound

Human-Like Code Quality Evaluation through LLM-based Recursive Semantic Comprehension cites this paper.

Human-Like Code Quality Evaluation through LLM-based Recursive Semantic Comprehension Can Large Language Models Be an Alternative to Human Evaluations?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T05:35:51.335834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:35:51.335834Z digest=sha256:b57766bc267db5dd2f3ea275dd5a2031c06ab2e53e258f2d290c73bd1b2f1da1

Observation 035363bf-f34a-4f15-9bfe-707d3331831b · inbound

Can Large Language Models Serve as Evaluators for Code Summarization? cites this paper.

Can Large Language Models Serve as Evaluators for Code Summarization? Can Large Language Models Be an Alternative to Human Evaluations?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:32:03.124672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:32:03.124672Z digest=sha256:b0de4ffa01dddd9f03a19be7bdabdf5c51d0523d0509a91faeaee2832f6a4885

Observation 3c9a7192-0623-4bfc-b300-bbd626bb44d8 · inbound

Multi-Facet Blending for Faceted Query-by-Example Retrieval cites this paper.

Multi-Facet Blending for Faceted Query-by-Example Retrieval Can Large Language Models Be an Alternative to Human Evaluations?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:27:02.267505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:27:02.267505Z digest=sha256:1ef0eac6965efe1fdaa0591b5baea598fbafb17fd7b74bfdcb8d8636de1411fc

Observation e7f030f7-35dc-4570-97ee-63008b05c85a · inbound

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models cites this paper.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Can Large Language Models Be an Alternative to Human Evaluations?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.563407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.563407Z digest=sha256:2c54714d6b49ce41740a206a06b6c6b1c0d36531681493a5e2d2a71ca631bab0

Observation ef857637-bd29-48a3-be1b-afe3df0fbd8c · inbound

The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance? cites this paper.

The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance? Can Large Language Models Be an Alternative to Human Evaluations?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:55.423085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:55.423085Z digest=sha256:36c60cbbfb81350f02905acb969982e0f5a93427bdc84ecfd6717ad989130500

Observation 9074c8f0-f0f8-4576-a9f0-12d9021e1673 · inbound

Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models cites this paper.

Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models Can Large Language Models Be an Alternative to Human Evaluations?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:56.195540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:35:56.195540Z digest=sha256:3a4e7130aa0309ebb5a3228becdbec2526c952392427aa6c6c25f961282d640d

Observation 803fb350-3e91-419d-bb69-400314a11faf · inbound

Generative Adversarial Reviews: When LLMs Become the Critic cites this paper.

Generative Adversarial Reviews: When LLMs Become the Critic Can Large Language Models Be an Alternative to Human Evaluations?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:57:43.472128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:57:43.472128Z digest=sha256:4c8a20d7461cd51c6b197581354a21b5c5c99c1f207d74464e326057a525dfa3

Observation 30aa11da-13e1-4ed8-ab8d-c1c7a56d4f0e · inbound

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta cites this paper.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Can Large Language Models Be an Alternative to Human Evaluations?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.024146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.024146Z digest=sha256:d59d26d97db484a588ccc24704a141b177d2cba8048926d4726add079d67b470

Observation 90df8baa-0b88-4629-973b-b0de585f97bf · inbound

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges cites this paper.

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Can Large Language Models Be an Alternative to Human Evaluations?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:15.501335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:15.501335Z digest=sha256:a7cc26b4aafd6a5d52ea34dfe26c4d3e4b26973b0f8eda5df290cdeff7fba153

Observation 21417018-4d95-476d-aa08-613aab922500 · inbound

The Impostor is Among Us: Can Large Language Models Capture the Complexity of Human Personas? cites this paper.

The Impostor is Among Us: Can Large Language Models Capture the Complexity of Human Personas? Can Large Language Models Be an Alternative to Human Evaluations?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:35:14.027629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:35:14.027629Z digest=sha256:5dedd1a0aa916097e6bde2e5205b842e22fa6f1fce61444548e6724dd1194a5e

Observation 0c757a55-1016-4b16-9f7b-7da69d37863e · inbound

Can Large Language Models Predict the Outcome of Judicial Decisions? cites this paper.

Can Large Language Models Predict the Outcome of Judicial Decisions? Can Large Language Models Be an Alternative to Human Evaluations?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:38.447618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:38.447618Z digest=sha256:62830d95ab7101791e1d837b628f7a9de45f5d1fd87ed0cf6da189bd2ac2303f

Observation f0fe5bfd-b6eb-4b82-9d8e-d5ee28e22c34 · inbound

Automated Assignment Grading with Large Language Models: Insights From a Bioinformatics Course cites this paper.

Automated Assignment Grading with Large Language Models: Insights From a Bioinformatics Course Can Large Language Models Be an Alternative to Human Evaluations?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T15:08:51.682665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:08:51.682665Z digest=sha256:52a25c451eb1c3b859a188f69dd7ca57be098e6989789157450cd80369905fce

Observation 9d837da3-c0d3-443b-909c-2ddfa06d830f · inbound

Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline cites this paper.

Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline Can Large Language Models Be an Alternative to Human Evaluations?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T23:21:01.959032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:21:01.959032Z digest=sha256:357564d195a99d52431b530425456b0d0ead1ebef4861635a4064aa5ce1a1c92

Observation 2044898d-6f81-4aa9-bd60-dcbe72f24209 · inbound

A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs cites this paper.

A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs Can Large Language Models Be an Alternative to Human Evaluations?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T10:44:41.079145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:44:41.079145Z digest=sha256:8e56e2675e560da34781965e99fde30fbac5a090070d32728f38f7911ac5101b

Observation bbdd4d8d-e2d6-4a3b-a199-1cbd65dcac8e · inbound

Stop Overvaluing Multi-Agent Debate -- We Must Rethink Evaluation and Embrace Model Heterogeneity cites this paper.

Stop Overvaluing Multi-Agent Debate -- We Must Rethink Evaluation and Embrace Model Heterogeneity Can Large Language Models Be an Alternative to Human Evaluations?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:44:11.154001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:44:11.154001Z digest=sha256:347e0f29ecc816176e93e4f87f3c8f9db6c40b00d9e0b013a4d788199e35f6fb

Observation 3ea38642-fee4-454e-b912-af096972d4fa · inbound

Image Embedding Sampling Method for Diverse Captioning cites this paper.

Image Embedding Sampling Method for Diverse Captioning Can Large Language Models Be an Alternative to Human Evaluations?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T19:25:33.558270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:25:33.558270Z digest=sha256:be12a870a251697bf7db81b8c5255db7baa05a9c524e0107ff1ce08edcd6a6a2

Observation f9799cb4-970b-4e35-ba8e-c3918c478614 · inbound

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge cites this paper.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Can Large Language Models Be an Alternative to Human Evaluations?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:08.594167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:08.594167Z digest=sha256:58be484923f389be9e7c0a4ef6db0797e933b5afc378154c976ea3736c0ea8e6

Observation 3f1316c7-326e-40a5-b1d3-2277807b866d · inbound

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models cites this paper.

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models Can Large Language Models Be an Alternative to Human Evaluations?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:36.992431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:36.992431Z digest=sha256:b62f58c21ac92ea532655d1286b3e6d22b2163452b2b6e4eee49e62756754257

Observation 65579d38-8691-4034-aace-16374b079272 · inbound

OSS-UAgent: An Agent-based Usability Evaluation Framework for Open Source Software cites this paper.

OSS-UAgent: An Agent-based Usability Evaluation Framework for Open Source Software Can Large Language Models Be an Alternative to Human Evaluations?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:38.647469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:38.647469Z digest=sha256:27c5679f1c5f8e1b470975c846314491536b04052b703a3a477b092c275f3ccb

Observation c1933a8e-b3d0-405c-be0a-094c8a48f047 · inbound

How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG cites this paper.

How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG Can Large Language Models Be an Alternative to Human Evaluations?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:53.297574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:53.297574Z digest=sha256:dfc49c0e7545f56855fe4671a76fd16c86df38ad14bc6ecdace826aa434d0b67

Observation 7578f62b-ed4b-4f24-bfe9-852e61ef515b · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Can Large Language Models Be an Alternative to Human Evaluations?

Reference 195

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:27.194667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:27.194667Z digest=sha256:6c06e056bf3c786682044c78ec3a92f0339d27c229e1113b0c653b57a5a92775

Observation 54074dbd-16a8-45eb-a163-1ff7bc0d7026 · inbound

Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics cites this paper.

Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics Can Large Language Models Be an Alternative to Human Evaluations?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:04.677606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:04.677606Z digest=sha256:c244256be0070cbb1cdbc86e4c717b83c51e73ad88a9b0eaf2750f6c8687ce72

Observation 59cc03f1-7b1c-493d-99d9-46e76a6d7f5b · inbound

CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation cites this paper.

CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation Can Large Language Models Be an Alternative to Human Evaluations?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:35.183306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:35.183306Z digest=sha256:3bbfb4bd79fe09f0b863b82db3dfcc74f54277dbb90aa47b341499013de335c2

Observation e38fc80c-7501-41f6-b8ba-bf5a816c9c98 · inbound

Counterfactual Voting Adjustment for Quality Assessment and Fairer Voting in Online Platforms with Helpfulness Evaluation cites this paper.

Counterfactual Voting Adjustment for Quality Assessment and Fairer Voting in Online Platforms with Helpfulness Evaluation Can Large Language Models Be an Alternative to Human Evaluations?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:40:14.996568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:40:14.996568Z digest=sha256:7007485c21e97a4b6b6db5303ff579159658bb441bdce13ab07617288865640e

Observation ec2ed023-6672-463e-b71c-a4a9d244423d · inbound

CitySim: Modeling Urban Behaviors and City Dynamics with Large-Scale LLM-Driven Agent Simulation cites this paper.

CitySim: Modeling Urban Behaviors and City Dynamics with Large-Scale LLM-Driven Agent Simulation Can Large Language Models Be an Alternative to Human Evaluations?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:39.652733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:22:39.652733Z digest=sha256:6b418f86c1132710e09739cf8427ba0b93660accdbe5c31e14b7ad588fbb9f75

Observation b1d6d28a-d741-4c4b-8ab3-f1fac97bed19 · inbound

Balancing Information Accuracy and Response Timeliness in Networked LLMs cites this paper.

Balancing Information Accuracy and Response Timeliness in Networked LLMs Can Large Language Models Be an Alternative to Human Evaluations?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T05:12:32.559295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:12:32.559295Z digest=sha256:cfbdf7eb0f968906e2ce00587e1f440ee624302a0eb4ebceba40b5282d621477

Observation 5f2966d4-d7e1-4d72-a744-79d8b4c41f34 · inbound

Generative Artificial Intelligence Extracts Structure-Function Relationships from Plants for New Materials cites this paper.

Generative Artificial Intelligence Extracts Structure-Function Relationships from Plants for New Materials Can Large Language Models Be an Alternative to Human Evaluations?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T22:56:19.869599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:56:19.869599Z digest=sha256:161ac19d664eb587a76e1bfbf4f0e422c448fef817bd7e070da2fe5df615b453

Observation 170f9278-a099-449f-8b80-b2227c26e93e · inbound

Rule2Text: A Framework for Generating and Evaluating Natural Language Explanations of Knowledge Graph Rules cites this paper.

Rule2Text: A Framework for Generating and Evaluating Natural Language Explanations of Knowledge Graph Rules Can Large Language Models Be an Alternative to Human Evaluations?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:19:15.326550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:19:15.326550Z digest=sha256:1b9cfec3c95215d4a7e4e67394b8191b907c019dfcfdaeda4a5c2132e9fc7216

Observation e55e6375-5023-4525-930f-9762a6c464aa · inbound

Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation cites this paper.

Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation Can Large Language Models Be an Alternative to Human Evaluations?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T17:19:10.145233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:19:10.145233Z digest=sha256:f4cebcc7cdb5264ab2daecf45dc89ec5757228b04c580b14b41c159b68ffefea

Observation 53dfcc92-1497-46ac-b364-d3a2c9f783b0 · inbound

LaQual: An Automated Framework for LLM App Quality Evaluation cites this paper.

LaQual: An Automated Framework for LLM App Quality Evaluation Can Large Language Models Be an Alternative to Human Evaluations?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T16:25:18.214750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:25:18.214750Z digest=sha256:35662487ec61d2562ee6b51c29615331a179087346461e47942c4f5b950ec468

Observation 65ab60d8-40af-4688-830c-3bfff04da3a1 · inbound

Integrating Generative AI into Cybersecurity Education: A Study of OCR and Multimodal LLM-assisted Instruction cites this paper.

Integrating Generative AI into Cybersecurity Education: A Study of OCR and Multimodal LLM-assisted Instruction Can Large Language Models Be an Alternative to Human Evaluations?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T11:15:57.397068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:15:57.397068Z digest=sha256:1dd9e34afcdd48ef74226c0aedd3988add3b711ca66012330f3bd57e3af9b2ff

Observation c0b98ba9-1358-48cd-b90a-8588516087d6 · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding Can Large Language Models Be an Alternative to Human Evaluations?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.683371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.683371Z digest=sha256:7e738bfc865c70b893fe9cd8d05ac0906413648846bf4b242050619af97187fe

Observation 14006fc8-2a4b-4f0d-952e-2f5837b31537 · inbound

Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration cites this paper.

Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration Can Large Language Models Be an Alternative to Human Evaluations?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T17:40:26.398037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:40:26.398037Z digest=sha256:862557a782289461c903f79154453a1baaab7bb50b08a95fd0fcd5ef14f75c37

Observation 61632217-2d50-409d-b6e5-6a8d682ed4cc · inbound

RESBev: Making BEV Perception More Robust cites this paper.

RESBev: Making BEV Perception More Robust Can Large Language Models Be an Alternative to Human Evaluations?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-15T00:10:48.378080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T00:10:48.378080Z digest=sha256:9c82b826d5a0f25333cc07cc0311ce526be855942a3d2d0b1d88bd5ba26a97c2

Observation a41f30db-f2bb-4359-a9c6-29dfe2f88bb4 · inbound

The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning cites this paper.

The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning Can Large Language Models Be an Alternative to Human Evaluations?

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:41:36.710801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T06:38:26.930020Z digest=sha256:7f4cb0b5947381200dbb13104d253399018c7b474261287e616a305e4183c7df

Observation 1852c7c6-a82c-4588-971c-8b94274622cf · inbound

Agent-Agnostic Evaluation of SQL Accuracy in Production Text-to-SQL Systems cites this paper.

Agent-Agnostic Evaluation of SQL Accuracy in Production Text-to-SQL Systems Can Large Language Models Be an Alternative to Human Evaluations?

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:11:28.693906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T07:05:57.220730Z digest=sha256:36e5461f55a4ccf6f0ae1d7d443b8a159ed084743622f54ca87dcc0e518fdd96

Observation 9832c778-716b-484b-b9f1-5cc675053839 · inbound

Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement cites this paper.

Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement Can Large Language Models Be an Alternative to Human Evaluations?

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:01:10.260954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T10:32:15.938575Z digest=sha256:4a49c6faa4640ab1499c56cf30d57f82e5a5c719b3f61b074d248dd4dbd9e811

Observation 677c1e49-b3a3-4693-b411-37b7c88188ff · inbound

Margin-Adaptive Confidence Ranking for Reliable LLM Judgement cites this paper.

Margin-Adaptive Confidence Ranking for Reliable LLM Judgement Can Large Language Models Be an Alternative to Human Evaluations?

Reference 117

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:12:39.511563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-19T16:05:41.505091Z digest=sha256:72d85c73dd2dd1c1e64576dd16208dbdc289222bc74fe7a2370a6fecd2c648f4

Observation 5db47bd7-7218-4732-b59d-ef9a38d09c68 · inbound

Benchmark Everything Everywhere All at Once cites this paper.

Benchmark Everything Everywhere All at Once Can Large Language Models Be an Alternative to Human Evaluations?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.078136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:3a1c95c9bec729a68b6a3e7bc70fa2dea84bd4676e26303db6e51d15aea656ba

Observation 60e96050-0a3d-431a-b6aa-d38f7ce63405 · inbound

StanceNakba Shared Task: Actor and Topic-Aware Stance Detection in Public Discourse cites this paper.

StanceNakba Shared Task: Actor and Topic-Aware Stance Detection in Public Discourse Can Large Language Models Be an Alternative to Human Evaluations?

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-27T10:10:48.752648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T10:05:02.300949Z digest=sha256:c17c7ff7e2e613f0cf2179ee233e4d1ac5c92e06f858acf266efbeee44a54209

Observation 0276a202-001e-4943-b76e-e5f0ab254512 · inbound

Iterating Toward Better Search: A Two-Agent Simulation Framework for Evaluating Agentic Search Architectures in E-Commerce cites this paper.

Iterating Toward Better Search: A Two-Agent Simulation Framework for Evaluating Agentic Search Architectures in E-Commerce Can Large Language Models Be an Alternative to Human Evaluations?

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:18:23.040793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T07:05:31.737761Z digest=sha256:a29ae130c1831be88867f8bc0f4a7576026d8b467d0e74f6dcce1a5c73dcf673

Observation 7adfcf79-b2b2-488c-8df2-3c3ed7903b59 · inbound

Game Theory Driven Multi-Agent Framework Mitigates Language Model Hallucination cites this paper.

Game Theory Driven Multi-Agent Framework Mitigates Language Model Hallucination Can Large Language Models Be an Alternative to Human Evaluations?

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-10T08:06:57.889561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T07:59:27.313453Z digest=sha256:326951f07f57ca3f9328fecfebd44a9b42cc07150798b3e451f0660f0ca1ac3e

Observation 421e9056-0d74-47db-801a-056fae582872 · inbound

Scaling Point-in-Time Language Models cites this paper.

Scaling Point-in-Time Language Models Can Large Language Models Be an Alternative to Human Evaluations?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T15:39:36.027708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:39:36.027708Z digest=sha256:b35e65bb2dadfc310257c25efe014520cdcdb5fdb7c8503c32e6e7e37b007497

Observation 87af01b4-7f64-4967-b3e3-62c7fad841f3 · inbound

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models cites this paper.

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models Can Large Language Models Be an Alternative to Human Evaluations?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T23:43:09.099487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T23:43:09.099487Z digest=sha256:83c8b48d691a9ed4c73e36a1c62607c21bea6ea6274e96c550be87a72ec955c5

Observation dbe6a12e-e9cb-41b0-b30d-e0229ad69db4 · inbound

Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable cites this paper.

Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable Can Large Language Models Be an Alternative to Human Evaluations?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T07:48:36.319338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:48:36.319338Z digest=sha256:8a0cdfc9ccb2a0fe946884b396d3c1e01b94879189a4f06a5213da461ee4f1eb

Observation 22d0fa57-b970-4538-8cce-554746fac9ca · inbound

(Towards) Scalable Reliable Automated Evaluation with Large Language Models cites this paper.

(Towards) Scalable Reliable Automated Evaluation with Large Language Models Can Large Language Models Be an Alternative to Human Evaluations?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T12:20:07.027430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T12:20:07.027430Z digest=sha256:5dc6a9b590fbcc627ca480e659a9c1edf16705ac098ebe1c5e58b193a73fc602

Observation e1511161-fd7d-413a-83e0-d220d59c6e99 · inbound

TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation cites this paper.

TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation Can Large Language Models Be an Alternative to Human Evaluations?

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:33.909559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:27:33.909559Z digest=sha256:c3389536000745cc3ad64f362bd3372fd4d4fdf63afa8dc290dcd9ee24095e49

Observation d72eedcd-f0f0-402f-b029-4a84f6caaacc · inbound

VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP cites this paper.

VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP Can Large Language Models Be an Alternative to Human Evaluations?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:56:40.104221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:56:40.104221Z digest=sha256:7d429369284dd6da088f6380836c2364cd0ba44f63f44d7830c51034d091d21a

Observation f360cad3-a1a7-4b5c-820d-9cabe5ff4e23 · inbound

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models cites this paper.

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Can Large Language Models Be an Alternative to Human Evaluations?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T13:10:34.444699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:10:34.444699Z digest=sha256:815e9032b5c5966d610e1f323f6bedb8c27c2c4c42c9845267565df8283c98db