Pith. sign in

Paper Citation Record · LEDGER

LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2305.13711.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.13711 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:32:47.892435Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

13
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5845ceaf-2264-486b-8a47-37a80cded7f4 · inbound

The Rise and Potential of Large Language Model Based Agents: A Survey cites this paper.

The Rise and Potential of Large Language Model Based Agents: A Survey LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:47:48.089570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:47:44.152066Z digest=sha256:2e64855e513c60eea88020d222fbadfb5218bb0f8944d1704b5f30c84287b7a0

Observation 4ef38040-ddcb-4866-8b9f-7509d29866da · inbound

The Prompt Report: A Systematic Survey of Prompt Engineering Techniques cites this paper.

The Prompt Report: A Systematic Survey of Prompt Engineering Techniques LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:16:17.959486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:16:17.875268Z digest=sha256:64aa7d30ca4c56d3f30c472fc6ec2d1a6a2269e20018b6dff5dd2b810582573d

Observation 8297fc60-5b9b-4af3-b422-632528bf69a5 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.265522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:12641f583a16f954994b442ede1e825fa03442eda7e7edb64890063672d7c5ca

Observation 2d4ddd91-5fb0-447d-974f-b8c6d610fc06 · inbound

LLMs to Support a Domain Specific Knowledge Assistant cites this paper.

LLMs to Support a Domain Specific Knowledge Assistant LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T23:35:38.372362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:35:38.372362Z digest=sha256:f324c0b1aea005ce5c5d775713a8d4b53571b715a5852e633e57097ab9c42abd

Observation 8dc45675-9d25-4c49-be35-f7f3e5710963 · inbound

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs cites this paper.

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T11:32:47.892435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:32:47.892435Z digest=sha256:cd54e854f9624183d284b6c3134ce520f7814f30038408cd1381b03df821244d

Observation 4068be9d-418f-4d2d-8c04-8c37d0b311d4 · inbound

How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG cites this paper.

How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:55.224300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:55.224300Z digest=sha256:0f44ce8aaabc612258d43d118dbbd7dcbd553442ae9091b1d4330cdf9a0b9579

Observation d57e1e38-1499-44f6-92d6-cd04e346b757 · inbound

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead cites this paper.

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 202

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:38.138381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:38.138381Z digest=sha256:e7478626389bc57a1b4aabefd70b6dd8103787fcf85461ce901031ead769bef1

Observation aabaf863-da9c-452c-a8c8-9587455fc849 · inbound

How Stylistic Similarity Shapes Preferences in Dialogue Dataset with User and Third Party Evaluations cites this paper.

How Stylistic Similarity Shapes Preferences in Dialogue Dataset with User and Third Party Evaluations LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:27:08.221285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:27:08.221285Z digest=sha256:f26f4ed9e375e01a65a412624369a7f6607fec3dd25a6c64a4243fb2b6b78fd0

Observation 958f4ec2-d612-42a4-95e4-d117d75e8720 · inbound

AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data cites this paper.

AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:37.679270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:37:37.679270Z digest=sha256:3c5f29785dd9329a3a38a1a21558b0514037a33d519ff59e0548ba82ca681d11

Observation 378b3328-2179-4531-887c-88d7a069cb16 · inbound

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution cites this paper.

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T22:54:24.564592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:54:24.564592Z digest=sha256:5ac344126e5a689804083730daf0bfd686538f31a11e458c8e28e8d568f4f096

Observation 188ac597-16b6-4ce8-b2f9-11e8718ccef4 · inbound

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models cites this paper.

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:30:23.267894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T22:29:36.960961Z digest=sha256:f1ecc494281e3b92240df5814ffa22d979adfb9527c5cfbfef5597d2367230e4

Observation 6f9bc56a-9784-479c-be9d-002aa8b041c9 · inbound

Data Selection for Multi-turn Dialogue Instruction Tuning cites this paper.

Data Selection for Multi-turn Dialogue Instruction Tuning LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:55:59.683275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:54:25.546834Z digest=sha256:6b689cc52e27257dcac773e34be1f6f361fa3ff698aafbfcdef42bbb6ae95d3a

Observation 97735431-dc0d-43ed-9288-07d9533441c0 · inbound

PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models cites this paper.

PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:18:15.753099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T23:17:29.025222Z digest=sha256:86d18b6b593eb4b7b580bd40a0309be519006bf934f018604f5c2e8e40b20201

Observation c1532c2d-ded9-44c0-9eed-4c33cc0ad124 · inbound

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why cites this paper.

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:53:58.861365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T21:43:59.462151Z digest=sha256:2daee9f288a2c0c1629cf0198ec4b16c6aaae533538365355d40bd7fd3a95898

Observation f4f2a142-e68a-49d4-8bf3-393017b314b5 · inbound

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why cites this paper.

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T02:20:17.847138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T02:20:17.847138Z digest=sha256:d6daa168238fd9ddfa9274dca68f19a780e4549a305da7b04cbcffe1529d0368

Observation 68e43071-7ef2-4fca-a085-aed917c8fdf3 · inbound

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? cites this paper.

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.536998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T05:59:58.183264Z digest=sha256:8890c735435596b000050229399eb6e70fb24d36f72b819abc9847ee5e9d47e2

Observation 0fb1fe54-f9c0-486b-880a-5d035780e7d0 · inbound

OmniPresent: Generating Coherent Presentation Suites from Scientific Papers cites this paper.

OmniPresent: Generating Coherent Presentation Suites from Scientific Papers LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T09:30:41.159870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:30:41.159870Z digest=sha256:72c43564c0b023e8dfd8ed8f43dfbdc0aa9caffd84522f6fc7b0a38afcb035c4