Pith. sign in

Paper Citation Record · LEDGER

LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2305.13711.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.13711 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:21:26.014390Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

13
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5845ceaf-2264-486b-8a47-37a80cded7f4 · inbound

The Rise and Potential of Large Language Model Based Agents: A Survey cites this paper.

The Rise and Potential of Large Language Model Based Agents: A Survey LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:47:48.089570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T10:47:44.152066Z digest=sha256:ad4b55220ed02a75add050b757cc05f1863b4bd7e872136d2d293adbcd1b7d15

Observation 4ef38040-ddcb-4866-8b9f-7509d29866da · inbound

The Prompt Report: A Systematic Survey of Prompt Engineering Techniques cites this paper.

The Prompt Report: A Systematic Survey of Prompt Engineering Techniques LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:16:17.959486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T02:16:17.875268Z digest=sha256:3886beabf2d85d227335ed18f08760e73e0552990d2defc2643b864aa44c8de9

Observation 0b709059-0e66-46cb-adbc-02615707342e · inbound

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models cites this paper.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.233656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.233656Z digest=sha256:799e8baa8bbdf52a465b563c29e7ac15e4e5058217ac1bfd500449cc5cc4316e

Observation 7fed3f83-a174-4bed-9481-4cb7bfbe194d · inbound

Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks cites this paper.

Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:28:14.345801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:28:14.345801Z digest=sha256:fe24e701aac7afbdd3e880cf6302f3fa2ecbb0ae2287c867f0e19e131b1ffd2f

Observation 8297fc60-5b9b-4af3-b422-632528bf69a5 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.265522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:6bb8154ee0df363fc470888453569940a20efe4f25b509f619d2ac6a355ee3d3

Observation d0c1f4d9-65d2-4510-8dac-c096b93179b6 · inbound

RAG Playground: A Framework for Systematic Evaluation of Retrieval Strategies and Prompt Engineering in RAG Systems cites this paper.

RAG Playground: A Framework for Systematic Evaluation of Retrieval Strategies and Prompt Engineering in RAG Systems LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T14:15:59.533685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:15:59.533685Z digest=sha256:ba52c9c254768d4d32c9cc9a7d674460b075878908022523593f9fb166e78b44

Observation 2d4ddd91-5fb0-447d-974f-b8c6d610fc06 · inbound

LLMs to Support a Domain Specific Knowledge Assistant cites this paper.

LLMs to Support a Domain Specific Knowledge Assistant LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T23:35:38.372362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:35:38.372362Z digest=sha256:d1115a30ebfde97221c61ade6db1180f1b4b75a474990040243ce69b6c2c0218

Observation 8dc45675-9d25-4c49-be35-f7f3e5710963 · inbound

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs cites this paper.

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T11:32:47.892435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:32:47.892435Z digest=sha256:35e18d8c77e20e5f2696bb4c2fad4877998bb2b8d673de843270279152e0fa75

Observation 0959be33-1ff3-4d99-bf62-e0700bccf69b · inbound

An Empirical Study of Evaluating Long-form Question Answering cites this paper.

An Empirical Study of Evaluating Long-form Question Answering LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:26.014390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:26.014390Z digest=sha256:6c6ccc4c1ec41850bef66200566583b8edfe9175127c3cffdf2482e73aeef793

Observation 32005ca5-4794-4a72-ba54-e22a26fc7317 · inbound

Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks cites this paper.

Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:09.780825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:18:09.780825Z digest=sha256:7dabe2723df46e80c6d9692ba8f063b3049f5de28884b24f34925b54987e5d04

Observation 47ff5ca5-de94-4afb-8dc1-8e649e45097d · inbound

LLM-based Evaluation Policy Extraction for Ecological Modeling cites this paper.

LLM-based Evaluation Policy Extraction for Ecological Modeling LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:05.451620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:05.451620Z digest=sha256:59f1264c6841440a90ca830f75e21328af251a6c603a15d26da52cce8bfc5609

Observation 4068be9d-418f-4d2d-8c04-8c37d0b311d4 · inbound

How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG cites this paper.

How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:55.224300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:55.224300Z digest=sha256:866b0af5b74b7e4e68d3380ace7cad66f3902229547594a194d1e6de95d50912

Observation d57e1e38-1499-44f6-92d6-cd04e346b757 · inbound

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead cites this paper.

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 202

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:38.138381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:38.138381Z digest=sha256:69d3115e759822728c49d237c672211672bc74161a14bafb051a6e4b71844dda

Observation aabaf863-da9c-452c-a8c8-9587455fc849 · inbound

How Stylistic Similarity Shapes Preferences in Dialogue Dataset with User and Third Party Evaluations cites this paper.

How Stylistic Similarity Shapes Preferences in Dialogue Dataset with User and Third Party Evaluations LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:27:08.221285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:27:08.221285Z digest=sha256:18342eb167491efa7f8141ac97232c648f822cc71fd4a1b47dd811e0ed19e8b1

Observation 958f4ec2-d612-42a4-95e4-d117d75e8720 · inbound

AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data cites this paper.

AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:37.679270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:37:37.679270Z digest=sha256:9fcb20efdd81746a08b8768cc733794a3f5730023d351e7ddb5cbd8c15c6b7a0

Observation 378b3328-2179-4531-887c-88d7a069cb16 · inbound

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution cites this paper.

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T22:54:24.564592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:54:24.564592Z digest=sha256:e7f385642e2279761926bfa4349264fc698a675d8daff6997dc8aac45091949a

Observation 188ac597-16b6-4ce8-b2f9-11e8718ccef4 · inbound

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models cites this paper.

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:30:23.267894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T22:29:36.960961Z digest=sha256:67af30cfb27befb9b947aea3ed2cbce06aea880a5404166c6565ff037117fdd1

Observation 6f9bc56a-9784-479c-be9d-002aa8b041c9 · inbound

Data Selection for Multi-turn Dialogue Instruction Tuning cites this paper.

Data Selection for Multi-turn Dialogue Instruction Tuning LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:55:59.683275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:54:25.546834Z digest=sha256:1cb704159cc9264394f85a9077d032530ca7b2d807ce8b0112b7049f1e7b88be

Observation 97735431-dc0d-43ed-9288-07d9533441c0 · inbound

PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models cites this paper.

PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:18:15.753099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T23:17:29.025222Z digest=sha256:241f6a4fadec218c3410abd77b04fad3a5edb1b8a4b1bca2fa4b542c16d8bf2e

Observation c1532c2d-ded9-44c0-9eed-4c33cc0ad124 · inbound

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why cites this paper.

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:53:58.861365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-29T21:43:59.462151Z digest=sha256:80c9847e5828a4f577084ac2bb72d09df75355d16859ea94f7b61bf3ae0b979b

Observation f4f2a142-e68a-49d4-8bf3-393017b314b5 · inbound

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why cites this paper.

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T02:20:17.847138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T02:20:17.847138Z digest=sha256:105154e28bb5f066ba2ffc3bc321e085be09ebe87b0343d6da5989ea464606ca

Observation 68e43071-7ef2-4fca-a085-aed917c8fdf3 · inbound

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? cites this paper.

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.536998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-30T05:59:58.183264Z digest=sha256:0d68df03773afcbd9695888177064c47f04532c669f25561e10e40c50af4d213

Observation 0fb1fe54-f9c0-486b-880a-5d035780e7d0 · inbound

OmniPresent: Generating Coherent Presentation Suites from Scientific Papers cites this paper.

OmniPresent: Generating Coherent Presentation Suites from Scientific Papers LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T09:30:41.159870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:30:41.159870Z digest=sha256:5cba77c72ba5a0a1991b1a33fcde85de07645988c9ae20b7e5f9b6fda34f252d