Pith. sign in

Paper Citation Record · LEDGER

Evaluation of Text Generation: A Survey

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2006.14799.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2006.14799 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:11:58.311715Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

198
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a2a8a936-2901-4bf7-aed2-d32891c2bd79 · inbound

ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate cites this paper.

ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate Evaluation of Text Generation: A Survey

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T13:03:18.831146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T13:03:18.765496Z digest=sha256:3563f2b5c180c5b4bbcfceb2f0eabdf18c15b937e14bb1fa5322bcd50ac8c941

Observation efb345bd-5e0a-46eb-a688-8598dcdaefd9 · inbound

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions cites this paper.

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions Evaluation of Text Generation: A Survey

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-24T00:53:40.707687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T00:52:52.056076Z digest=sha256:15f6fc57472ef23f27f067daa5a784b475521bef68fb764192dcd9c0083349bb

Observation 9cec52b8-d6d1-4973-9769-01f4e119bf6b · inbound

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? cites this paper.

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? Evaluation of Text Generation: A Survey

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T22:42:13.667706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T22:38:35.969273Z digest=sha256:5d897301dabb2ec29ee5e6317804c4c6b26122ae1dd218ba97782056b26432c1

Observation 143a904f-2d5f-4f5a-9e60-0ef3fc36568d · inbound

The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing cites this paper.

The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing Evaluation of Text Generation: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:58.311715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:11:58.311715Z digest=sha256:9fcdeddce93bf9f3684c5c4235b7e34ced7bdafba5901a04551a463d0b501359

Observation ce75d046-338e-4a26-8104-92b995978220 · inbound

COGENT: A Curriculum-oriented Framework for Generating Grade-appropriate Educational Content cites this paper.

COGENT: A Curriculum-oriented Framework for Generating Grade-appropriate Educational Content Evaluation of Text Generation: A Survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:56:45.220857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:56:45.220857Z digest=sha256:468b1ad2f7f43cd57abca5f8b010cd98971757adc9776d290a058b2d9436c60f

Observation c300e68f-6016-4616-baa5-125b9b337477 · inbound

From Multimodal Perception to Strategic Reasoning: A Survey on AI-Generated Game Commentary cites this paper.

From Multimodal Perception to Strategic Reasoning: A Survey on AI-Generated Game Commentary Evaluation of Text Generation: A Survey

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:21.956909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:21.956909Z digest=sha256:e11ff6e8aa309c60ac1eb6cdd20c275f0f37cd6da30e2c7ab09b7e9bbed1f131

Observation 50479774-0e37-4d24-9f6f-4533ae249535 · inbound

Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks cites this paper.

Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks Evaluation of Text Generation: A Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:35:50.545027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T03:32:53.109030Z digest=sha256:24a4eecda708a28d580411d5817e158d1e8dc984355191621fb3841a6958e63a

Observation 007b65bd-009e-47b0-a59b-d46d02cf808c · inbound

What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025 cites this paper.

What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025 Evaluation of Text Generation: A Survey

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T11:07:09.410983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:07:09.410983Z digest=sha256:ba59c8747f0c42a700e6cd961eada9ee653a263bfede23cc0753d3ea47e06ee8

Observation 0e5a751a-a9a4-46ff-b9e6-e283b33d1e13 · inbound

FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge cites this paper.

FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge Evaluation of Text Generation: A Survey

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T03:55:51.714813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:55:51.714813Z digest=sha256:1e9adf70b953fe0942bf639540ae25fca365727affa6ac6bc72226a27baa705c

Observation 88c2bb92-11a6-43dc-9004-87cb3914fba4 · inbound

Synthetic Reflections on Resource Extraction cites this paper.

Synthetic Reflections on Resource Extraction Evaluation of Text Generation: A Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:57:15.853158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T03:54:55.533944Z digest=sha256:fba2e1ffcea1a59cdc86479a96e98c87dbd7c763afe07dac5b34d5596e2ba1f9

Observation 617c7367-66b1-45de-9ff0-84e8271b4aa0 · inbound

Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation cites this paper.

Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation Evaluation of Text Generation: A Survey

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:15:34.545671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T12:13:24.420641Z digest=sha256:762de4c6539ba2f66f8f9ace69fd5234c5318fc340e2d34b595a81deb020052c

Observation a7868629-76f9-46c2-b2f2-abefed8d414a · inbound

Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation cites this paper.

Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation Evaluation of Text Generation: A Survey

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T02:35:45.191487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:35:45.191487Z digest=sha256:8389d2002f5c6b3bb63462bf246eb45ecbbe4b0afb1601fc3f8a3465753f9c17

Observation 7cba6b9d-5c4e-493e-bfd5-284663c2206d · inbound

AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis cites this paper.

AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis Evaluation of Text Generation: A Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:05:48.300289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T19:22:08.306946Z digest=sha256:c287683a5248def09571cf4499cf8f51d56d6464a932d3a014283d9b5471e25a

Observation 773ab666-1d23-4a4b-97ab-a14db984d069 · inbound

Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation cites this paper.

Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation Evaluation of Text Generation: A Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:01:00.374583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:42:35.824427Z digest=sha256:c4944a74ae68365b08d3eb49ca483a8319031aecf13eb82c02d07e78fe3346a7

Observation 859b8115-85cd-4cde-98e8-3dd14d3159fd · inbound

Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest cites this paper.

Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest Evaluation of Text Generation: A Survey

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:41:01.810530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T03:20:12.777878Z digest=sha256:5072f8b7283a1cd8cc020732021991a6d3fa5e57099006ed9a5f1c396acd71ca

Observation 3d0da929-fb2c-439d-8cba-fe3d48bdc309 · inbound

From World-Gen to Quest-Line: A Dependency-Driven Prompt Pipeline for Coherent RPG Generation cites this paper.

From World-Gen to Quest-Line: A Dependency-Driven Prompt Pipeline for Coherent RPG Generation Evaluation of Text Generation: A Survey

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:46:23.304340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T16:22:33.542081Z digest=sha256:ba82a7500c50b11aa0699e380a342325070221ee9402d88f777c5f950b643676

Observation 015d2145-0090-4c70-8c2a-e1b952ffb4b4 · inbound

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs cites this paper.

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs Evaluation of Text Generation: A Survey

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:16:07.273886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T16:26:54.068997Z digest=sha256:32bdfa4e526730bfbc045a347f85b5d571e3ce3f22a5cec1ac25cad0e6e6b029

Observation f9c2670c-78e1-4163-ab12-bf5469935630 · inbound

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs cites this paper.

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs Evaluation of Text Generation: A Survey

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:16:16.112644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:13:48.672711Z digest=sha256:ac6a0d3f3d1d4326a605291d76f053a4eeba38e62408d726f6d1436d62853daf

Observation b47a783a-d4fc-442d-b115-fc96af605418 · inbound

Convex Optimization for Alignment and Preference Learning on a Single GPU cites this paper.

Convex Optimization for Alignment and Preference Learning on a Single GPU Evaluation of Text Generation: A Survey

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:05:23.148540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-25T05:01:31.560963Z digest=sha256:2368316f3b0f4d976e3f40eafb89acd1b62ebcd61aa91b87848365949c9d4921

Observation e88dd9d5-af7c-4532-8131-fb9a56f46150 · inbound

ExCAM: Explainable Cultural Awareness Metrics cites this paper.

ExCAM: Explainable Cultural Awareness Metrics Evaluation of Text Generation: A Survey

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:43:13.260618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:42:43.805674Z digest=sha256:b823ec27fa074acb8fdd88ea04038ead206d5cc99b4148cd702619c14eece16c

Observation bbd39bfd-7cfe-46dc-b98d-a3dd669da1f9 · inbound

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation cites this paper.

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation Evaluation of Text Generation: A Survey

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:37:22.704121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T20:20:08.996005Z digest=sha256:204e0fadc9bc74d3aca6a9d2e0dd5831d330fce080563e9a9a58a0a0a7ad8c98

Observation 24c20b2b-7a21-4c61-8d33-96d4d2031529 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Evaluation of Text Generation: A Survey

Reference 226

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.717285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:01e88c202d4a2c201acb68241e15313b8858744eefdefceba60159f595e7e397

Observation 7692ee91-fc88-446a-b53a-8d491efa1769 · inbound

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning cites this paper.

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning Evaluation of Text Generation: A Survey

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.839875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T19:04:45.062426Z digest=sha256:a30ed9dd134be5cf19ae6d68c80337ead8b4ea953659f52efcfc27a868398335

Observation f1743dbb-a58a-4bba-a85e-423bbb92f935 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Evaluation of Text Generation: A Survey

Reference 198

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.750686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:f3d5a05859f912cda9ce30165d030e99d779a6d225401aa181735bfe5d4b85bd

Observation b4d60861-c9df-42a0-8558-524570f03689 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Evaluation of Text Generation: A Survey

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:55:59.679519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:c09ce05860a5d8afc97eac42682a534709b6c5e3b6dba5940dd24ba3550fe211

Observation aea004cf-30af-4fc3-9326-98ba3f2ff614 · inbound

SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework cites this paper.

SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework Evaluation of Text Generation: A Survey

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T18:57:16.346686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T18:48:09.531329Z digest=sha256:b2a1d1c48305b5028b1df82d2a1bf47b02f69e8f10ebedbd5bbda316fc0f960b

Observation 0c3d5759-7a87-45ce-a85c-37d305a1e334 · inbound

Fidelity-Diversity Metrics for Text cites this paper.

Fidelity-Diversity Metrics for Text Evaluation of Text Generation: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T17:13:44.985215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:13:44.985215Z digest=sha256:9fd9f6d03646056bc76aa7e1484c385d4b60ae209f6b2fdc6f4012599122ed80

Observation ca7c4a61-66e9-4aad-b528-250fd0392090 · inbound

grapheme-kit: Grapheme-Level Metrics and Text Processing for Multilingual NLP cites this paper.

grapheme-kit: Grapheme-Level Metrics and Text Processing for Multilingual NLP Evaluation of Text Generation: A Survey

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T04:46:36.501877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:46:36.501877Z digest=sha256:ba61c08ed15a3142b03bf10511921e1cd5c71bfc1c302a97968bac486fb4745d

Observation d0aa5349-6573-4528-84e6-c435d6897b56 · inbound

LLM-as-a-Judge for Evaluating System Responses in Conversational Music Recommendation cites this paper.

LLM-as-a-Judge for Evaluating System Responses in Conversational Music Recommendation Evaluation of Text Generation: A Survey

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:41.469277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:52:41.469277Z digest=sha256:8b25099e1c63b8fe89997b19cc06cf45cf585cb216d89fd01b44ea3fb9f2c9d0