Pith. sign in

Paper Citation Record · LEDGER

A Closer Look into Automatic Evaluation Using Large Language Models

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2310.05657.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.05657 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:23:16.291779Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T13:21:24.424456Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cff20c7d-4cb6-46c8-a067-effaa1d92336 · inbound

DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models cites this paper.

DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models A Closer Look into Automatic Evaluation Using Large Language Models

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:21:24.426800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-16T13:21:24.297836Z digest=sha256:28ae6eabfc1c658b26bbaa5bfd9adb043e8db06a6ed104884b604f67c1db81ab

Observation dc38f512-19d1-4b03-9cfa-bb5d867f70bd · inbound

SRSA: A Cost-Efficient Strategy-Router Search Agent for Real-world Human-Machine Interactions cites this paper.

SRSA: A Cost-Efficient Strategy-Router Search Agent for Real-world Human-Machine Interactions A Closer Look into Automatic Evaluation Using Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:12:03.395779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:12:03.395779Z digest=sha256:fc431ecab5bc6f62661c7f19b550ba3e613e30a0076f30690a17c7995a7fd036

Observation 53e2fd14-1328-4777-9240-8d802af2a910 · inbound

Do LLMs Agree on the Creativity Evaluation of Alternative Uses? cites this paper.

Do LLMs Agree on the Creativity Evaluation of Alternative Uses? A Closer Look into Automatic Evaluation Using Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:12:47.293766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:12:47.293766Z digest=sha256:e59922f228de5047c82dada753f223a64705c173d2c0e91cd560e6000c2632a8

Observation 2c6c2898-cd37-47a3-971f-33cbb43c2730 · inbound

Can Large Language Models Serve as Evaluators for Code Summarization? cites this paper.

Can Large Language Models Serve as Evaluators for Code Summarization? A Closer Look into Automatic Evaluation Using Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T04:32:03.129691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:32:03.129691Z digest=sha256:cba6e2a4c3be04d4fdcdc9de2db9f6fe37d2f98388442ec0e569ac612c5a689d

Observation e163aa16-5511-4ac9-a9ee-7262dde01c41 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods A Closer Look into Automatic Evaluation Using Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.213435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:6cab869f84c47e7f57a8fab1edcee759bacb15144c464cf7c62c132889487f74

Observation b392c54d-b376-4cc4-ad89-4c8829e860a4 · inbound

Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources cites this paper.

Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources A Closer Look into Automatic Evaluation Using Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T15:51:48.547875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:51:48.547875Z digest=sha256:b5560efdc606e83563f7aa806bd0ec880cb05c3ac57a7b2134f4bd178ad528ea

Observation 93ef57d5-2b0a-4b62-a680-e1442ec51b7b · inbound

Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach cites this paper.

Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach A Closer Look into Automatic Evaluation Using Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:23:16.291779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:23:16.291779Z digest=sha256:5fd096e7c97dedcafe91b59343151541ad6dac62d41f9907afae2df064465ff9

Observation 47647def-fcef-4404-9f48-38fca9d1d730 · inbound

An Empirical Study of Evaluating Long-form Question Answering cites this paper.

An Empirical Study of Evaluating Long-form Question Answering A Closer Look into Automatic Evaluation Using Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.922444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.922444Z digest=sha256:b3890b482c6c8fcd317e728cd76f96a78723b87f81472cab540f5857be286cf4

Observation ab4f6210-4423-471c-a539-eaf97fb7ffdd · inbound

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation cites this paper.

Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation A Closer Look into Automatic Evaluation Using Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:46.963076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:46.963076Z digest=sha256:c3aa9728608dc838b2e44ec0af6ae2fb7f8408856d70f7f2e5abc2adf7aefc24

Observation 52f0b447-b60c-45b0-9fb7-fff4ca4b4a4a · inbound

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts cites this paper.

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts A Closer Look into Automatic Evaluation Using Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:44.035946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:44.035946Z digest=sha256:6cac66da3fde532ae99cb1202b53c1621f9bbe1a5495c05f0c28559186ba0f02

Observation 2a51c0da-2df3-497b-909d-c47dd448e45c · inbound

Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models cites this paper.

Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models A Closer Look into Automatic Evaluation Using Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T16:59:53.321474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:59:53.321474Z digest=sha256:ddebd53c327cea048715b0ba612aa865b878af94d1f7212b54cb930e0f7e7cec

Observation d63e359a-b106-4ba5-863b-6b3ffcb2ccbe · inbound

PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses cites this paper.

PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses A Closer Look into Automatic Evaluation Using Large Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:00:02.948664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T13:57:41.428695Z digest=sha256:f6b8c59ccd139171436025fcb339627e1898399b45e1bd5c22507187ddb8687e

Observation ad93f109-7739-4ce9-9fde-edd5fa5eda8e · inbound

Towards Annotation-Free Validation of MLLMs: A Vision-Language Logical Consistency Metric cites this paper.

Towards Annotation-Free Validation of MLLMs: A Vision-Language Logical Consistency Metric A Closer Look into Automatic Evaluation Using Large Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:11.378608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T10:18:37.901403Z digest=sha256:80913ddb5d39cbdee84d20db395fd484f887a6ce541d8a30d94dcfb814059736