Pith. sign in

Paper Citation Record · LEDGER

Is ChatGPT a Good NLG Evaluator? A Preliminary Study

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2303.04048.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.04048 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:59:32.351325Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0cd08ba4-c0c0-4b19-97d5-f78d227a9fff · inbound

G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment cites this paper.

G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T22:55:50.901173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T22:55:50.014540Z digest=sha256:52cafbca45ffe0251ea1a79e1dbb6a61cacc0e86bebe753c6bb5b9e5722fd541

Observation 2a945c7e-7d2a-48b2-bcd2-39f3bffba0e4 · inbound

ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate cites this paper.

ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:03:18.849944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T13:03:18.765496Z digest=sha256:2a774e83e3a2e5ab474760398b292c68683589d633f69f283d20625050fb762d

Observation da3b69b4-b1ad-43a3-a649-ca760b30037e · inbound

GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts cites this paper.

GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:25:21.097589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T06:25:20.966510Z digest=sha256:dcd3aeb8c5aa0177e54821236801f7f8c120410c1bf9d211f389d3840b81d8ed

Observation d53759e9-017d-4afd-a1b8-aa63f2e7c5af · inbound

From Local to Global: A Graph RAG Approach to Query-Focused Summarization cites this paper.

From Local to Global: A Graph RAG Approach to Query-Focused Summarization Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:58.035351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-11T05:10:57.816312Z digest=sha256:a8086718dfff5f32d8953bac7a8ff7ec41c20df911c24263e075198b1dc82aba

Observation bc4e9816-f1ee-4c3e-8b68-81dca979c415 · inbound

Dictionary Insertion Prompting for Multilingual Reasoning on Multilingual Large Language Models cites this paper.

Dictionary Insertion Prompting for Multilingual Reasoning on Multilingual Large Language Models Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:05:44.746639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-23T18:03:19.212619Z digest=sha256:8e491c34195162b382d89bc6b3d476ba4d5205569e9393890694b4d6770948b7

Observation bbf426cf-9d02-4344-b980-43b4eedce01e · inbound

The Performance of the LSTM-based Code Generated by Large Language Models (LLMs) in Forecasting Time Series Data cites this paper.

The Performance of the LSTM-based Code Generated by Large Language Models (LLMs) in Forecasting Time Series Data Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T10:59:32.351325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:59:32.351325Z digest=sha256:297ce9d5061f2eb1224881bfcf31155620d6643c0ce66843dfc2302a65b1b88f

Observation ac43f845-a8d3-4563-997b-b3718a9a75be · inbound

Can Large Language Models Serve as Evaluators for Code Summarization? cites this paper.

Can Large Language Models Serve as Evaluators for Code Summarization? Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:32:03.114529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:32:03.114529Z digest=sha256:457d291203f20c07d9d54a7e4e35d4641252cec0cb2f8c0c5821091513e0f9bb

Observation eeea721b-332f-47bd-b9a5-6a3963aa821d · inbound

Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM cites this paper.

Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:53:12.345574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:53:12.345574Z digest=sha256:5a17fb3732a0eec770e3dc531e0f2a2d9ba69411047681cb296dc15ac5518bf2

Observation ff2479d4-4955-448a-ac03-c54cf16ba38a · inbound

Can LLMs Ask Good Questions? cites this paper.

Can LLMs Ask Good Questions? Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:33.156333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:33.156333Z digest=sha256:ae1a59ac91e2dd6f74f9eb7115ba4adcaad78343cd2ff17f5fbb8e60ced78ad5

Observation 1f0cc3fd-c942-477f-9192-509c990e79c5 · inbound

Optimization is Better than Generation: Optimizing Commit Message Leveraging Human-written Commit Message cites this paper.

Optimization is Better than Generation: Optimizing Commit Message Leveraging Human-written Commit Message Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T19:41:51.246209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:41:51.246209Z digest=sha256:c231f3bbb2381722640010dbbb837fc7f9c3c283f92a14906be6f78aa57accf3

Observation d12a5834-776f-4c48-a43c-de0aa7f17bf7 · inbound

Evaluating Small Language Models for News Summarization: Implications and Factors Influencing Performance cites this paper.

Evaluating Small Language Models for News Summarization: Implications and Factors Influencing Performance Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T18:17:21.169401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:17:21.169401Z digest=sha256:9586a727c9da598e9306bf3a6720a2650d9fae066f684e33abe2a5e3141e3272

Observation 65f07b96-8cf7-42db-8261-40ad4f42a36c · inbound

Reason4Rec: Deliberative User Preference Alignment of Large Language Models for Recommendation cites this paper.

Reason4Rec: Deliberative User Preference Alignment of Large Language Models for Recommendation Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T13:33:57.670287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:33:57.670287Z digest=sha256:de886b74fdd862531a7a524e60a8c7bd2b5807da1eb1271f199be26e02cb377a

Observation 3b2abbf5-d459-4bd4-aa41-eda814481ab7 · inbound

LLMs to Support a Domain Specific Knowledge Assistant cites this paper.

LLMs to Support a Domain Specific Knowledge Assistant Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T23:35:38.452575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:35:38.452575Z digest=sha256:be0576745aec130b91e62e16b3721d951dc279c6636d3def5dc286b9372a9407

Observation fb4babe8-2286-4312-9de9-f81bf65a5075 · inbound

In-depth Analysis of Graph-based RAG in a Unified Framework cites this paper.

In-depth Analysis of Graph-based RAG in a Unified Framework Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:37:22.326507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T01:36:25.057478Z digest=sha256:453faff0a38869550de9dc1f0c8d7467d2faba143d2b650784fe982307deeb7f

Observation 0aec59ac-933b-45c0-97a9-90d788933f67 · inbound

R-TOFU: Unlearning in Large Reasoning Models cites this paper.

R-TOFU: Unlearning in Large Reasoning Models Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:10.122394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:26:10.122394Z digest=sha256:5991d86d47b17fc7fb7cab9c6c7c591841c9841097f8cdb64b7157cb3d9e899d

Observation 9ef6c73d-62fc-4021-b73b-af03aada0b01 · inbound

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge cites this paper.

Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:10.352664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:10.352664Z digest=sha256:e279cbcb939dc7c6a084d5ae5077eab1e6ea48b2785ea667e58b2ece70480ab4

Observation e384682d-fe63-4e4f-81c2-40d84d611c53 · inbound

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion cites this paper.

$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.497625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:41.497625Z digest=sha256:9ff06e120f02670ab827d67bafd3c0c734bec6d96399d6b18f44ce573056c050

Observation d54c5bac-9273-48d6-b8c2-88b91f717d13 · inbound

Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons cites this paper.

Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:27.365332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:27.365332Z digest=sha256:0fbf9c4a27cca17bca349ba9eb7491d19d644b4c28673fb49c27bdb6888c1bdd

Observation 61ca4b3c-b906-4721-b1d8-51e7392ce5bf · inbound

Retrieval-Augmented Recommendation Explanation Generation with Hierarchical Aggregation cites this paper.

Retrieval-Augmented Recommendation Explanation Generation with Hierarchical Aggregation Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:07:37.408102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:07:37.408102Z digest=sha256:ee93f1870df5047caf93efa880931e29e7bc35eb51c010cf84734d732b4e97fe

Observation 7074985d-727f-467a-9264-d029930fb6de · inbound

SLoW: Select Low-frequency Words! Automatic Dictionary Selection for Translation on Large Language Models cites this paper.

SLoW: Select Low-frequency Words! Automatic Dictionary Selection for Translation on Large Language Models Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:34:26.392442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-21T23:31:58.225436Z digest=sha256:0a56c70e586b9bcdd8076b5d452200925cf8ea438c13b7e3f9905c148181f83b

Observation 6b9d5de5-5b61-4411-aa16-363f63b34525 · inbound

Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts cites this paper.

Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:21.334991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:21.334991Z digest=sha256:73cf73f0d4f85e6e4e093c7d526a2c7f8774c8051d1cefb92d64e9901a046854

Observation 0a3412f4-d392-46eb-9f58-4b818546cf79 · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:31.375425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:31.375425Z digest=sha256:94674415991c893ec9fc1b02157c3829d21d052053605f807d426283a861ea08

Observation 37094538-b6cc-4d56-a730-73334d089f0a · inbound

RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model cites this paper.

RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:03.720711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:03.720711Z digest=sha256:d0dcc5de4b8b1e3c2268633e2676b8c5638e310ab3bb6db257b1725f087c6a69

Observation 0c864c14-98df-4288-ae5e-7a1bfb264827 · inbound

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data cites this paper.

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:49.943824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:49.943824Z digest=sha256:b967fe3bc4658426d8995d5937ae6b66906f56ef6cc56d6abb98dccbaee18846

Observation d11aba2e-8ed6-4286-91e1-cc67ff93c667 · inbound

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method cites this paper.

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:42.132070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:42.132070Z digest=sha256:076f3a687b5d33e8809e2c31b03f1e31222baa914015238cefc43d6dc2fff4b5

Observation f1f30f70-58da-4bda-9248-7d40dbb60e40 · inbound

PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses cites this paper.

PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:00:03.058822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T13:57:41.428695Z digest=sha256:46a7e585c6d727cadabcf749105af635bdab3b485b8f5231c72de6c187ec6d4f

Observation 5d918467-b1be-45b2-9bde-d2798909e7d0 · inbound

MMP-Refer: Multimodal Path Retrieval-augmented LLMs For Explainable Recommendation cites this paper.

MMP-Refer: Multimodal Path Retrieval-augmented LLMs For Explainable Recommendation Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:33:02.616446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T17:28:28.480792Z digest=sha256:9bb89ab9959bdcdd1d14f4973b3ad5f0d7efa5dbe8f0d4c515c685be26e83ec5

Observation 30b93ae1-88e9-4b1c-9097-577eed10007a · inbound

Supporting System Testing with a Multi-Agent LLM-based Framework for Knowledge Graph Extraction: A Case Study with Ethernet Switch Systems cites this paper.

Supporting System Testing with a Multi-Agent LLM-based Framework for Knowledge Graph Extraction: A Case Study with Ethernet Switch Systems Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:33:08.873287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T08:32:39.748993Z digest=sha256:dbc614df362993da0d4564b27924e2354712c7c7f5f9ed331cd9c2b3f7d64680

Observation cf66b034-f40b-41cb-a939-c61f338d46b9 · inbound

A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models cites this paper.

A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:22:37.605994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T19:53:36.990878Z digest=sha256:a9c51a6c1224b56845e60525c2f07f5341764c631fd5281f9f14fc16e0d64c80

Observation 13a13b4e-1704-49b8-bffb-1735d23d7305 · inbound

TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation cites this paper.

TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 120

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:34.073545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:27:34.073545Z digest=sha256:ecd2aed557d394a919f097fcad4f7e58fa01c27a8c70ae8dcd48b53ceb1ffb80