Pith. sign in

Paper Citation Record · LEDGER

A Closer Look into Automatic Evaluation Using Large Language Models

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2310.05657.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.05657 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:23:16.291779Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T13:21:24.424456Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cff20c7d-4cb6-46c8-a067-effaa1d92336 · inbound

DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models cites this paper.

DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models A Closer Look into Automatic Evaluation Using Large Language Models

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:21:24.426800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-16T13:21:24.297836Z digest=sha256:a54f6e1e82a2321ad95b0acdbc3d7a3a769a6222cd4e023f0d04b25974df95ff

Observation dc38f512-19d1-4b03-9cfa-bb5d867f70bd · inbound

SRSA: A Cost-Efficient Strategy-Router Search Agent for Real-world Human-Machine Interactions cites this paper.

SRSA: A Cost-Efficient Strategy-Router Search Agent for Real-world Human-Machine Interactions A Closer Look into Automatic Evaluation Using Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:12:03.395779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:12:03.395779Z digest=sha256:b47c1d15526792809bc772a362425dbf847d5f9201806209b8906ce26e64b913

Observation 53e2fd14-1328-4777-9240-8d802af2a910 · inbound

Do LLMs Agree on the Creativity Evaluation of Alternative Uses? cites this paper.

Do LLMs Agree on the Creativity Evaluation of Alternative Uses? A Closer Look into Automatic Evaluation Using Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:12:47.293766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:12:47.293766Z digest=sha256:f54f7ccb7729ff34a109de19b8486169fb6c5ffb34586dc278c7820e1c2400e1

Observation 2c6c2898-cd37-47a3-971f-33cbb43c2730 · inbound

Can Large Language Models Serve as Evaluators for Code Summarization? cites this paper.

Can Large Language Models Serve as Evaluators for Code Summarization? A Closer Look into Automatic Evaluation Using Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T04:32:03.129691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:32:03.129691Z digest=sha256:db498cb2eb8ab70532414e5a2987e4afa4fbd582aa8b93c626baff3c0d5e03c1

Observation e163aa16-5511-4ac9-a9ee-7262dde01c41 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods A Closer Look into Automatic Evaluation Using Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.213435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:661b9b791c458891200119498f7d186ff66cafb4e2523616fb06f851acd21c55

Observation b392c54d-b376-4cc4-ad89-4c8829e860a4 · inbound

Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources cites this paper.

Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources A Closer Look into Automatic Evaluation Using Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T15:51:48.547875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:51:48.547875Z digest=sha256:c72d0ec0dc39a70ac8a1be55d70f92871cfc652e45768497c4a53ae5ac0ba5f4

Observation 93ef57d5-2b0a-4b62-a680-e1442ec51b7b · inbound

Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach cites this paper.

Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach A Closer Look into Automatic Evaluation Using Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:16.291779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:23:16.291779Z digest=sha256:83efa84f55624d75d0fd639f8c2d895309803b1f02e91ca1e23c7b918c74a929

Observation 47647def-fcef-4404-9f48-38fca9d1d730 · inbound

An Empirical Study of Evaluating Long-form Question Answering cites this paper.

An Empirical Study of Evaluating Long-form Question Answering A Closer Look into Automatic Evaluation Using Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.922444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.922444Z digest=sha256:a074186c6cb95ca6cc8ca229ce376ac14e24873f4b9774ba0ac195c0d40daeec

Observation ab4f6210-4423-471c-a539-eaf97fb7ffdd · inbound

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation cites this paper.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation A Closer Look into Automatic Evaluation Using Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.963076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.963076Z digest=sha256:7b12d1f4aae8e92331e0cd028f9ddfc57fdb4457a10037a53bd5128deec8a056

Observation 52f0b447-b60c-45b0-9fb7-fff4ca4b4a4a · inbound

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts cites this paper.

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts A Closer Look into Automatic Evaluation Using Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:44.035946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:44.035946Z digest=sha256:9c02c2cdd18bf108748aa2137ca9aba3d3f02c739242a2256655f5148ff5ce6b

Observation 2a51c0da-2df3-497b-909d-c47dd448e45c · inbound

Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models cites this paper.

Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models A Closer Look into Automatic Evaluation Using Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T16:59:53.321474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:59:53.321474Z digest=sha256:d51f972dd3e8ff3897a9e627ddfbbaf6b0201f304d9934a84bbc493b57d05121

Observation d63e359a-b106-4ba5-863b-6b3ffcb2ccbe · inbound

PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses cites this paper.

PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses A Closer Look into Automatic Evaluation Using Large Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:00:02.948664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T13:57:41.428695Z digest=sha256:4aac0285137920d3ab3abddaeee74160c6d6a7f6a405f5671597fbb6f531fb55

Observation ad93f109-7739-4ce9-9fde-edd5fa5eda8e · inbound

Towards Annotation-Free Validation of MLLMs: A Vision-Language Logical Consistency Metric cites this paper.

Towards Annotation-Free Validation of MLLMs: A Vision-Language Logical Consistency Metric A Closer Look into Automatic Evaluation Using Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:11.378608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T10:18:37.901403Z digest=sha256:1e33615898434f7965849176f9196725e96659ce83a1965f86f9f7e5cb2db5a3