Pith. sign in

Paper Citation Record · LEDGER

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models

As of 12 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2411.15320.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15320 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:29:54.284637Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T18:27:23.076544Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T18:31:44.524453Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6a8cb7d-5d4d-453d-af05-4bb025f95416 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.171039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.171039Z digest=sha256:73974be4cce5e3ddee4562422bcab0351ddae7fa38cdd019e43651b3746944ca

Observation af84ff0f-652b-464d-af39-d514e287583a · outbound

This paper cites write newline.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.176096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.176096Z digest=sha256:d38ca72a3848110e0f023c34d72ddb25ba2effd224204d3e340d3e742ec84e61

Observation 4f807f6a-add5-4a53-a64a-1e10b2e9877c · outbound

This paper cites Benchmarking Foundation Models with Language-Model-as-an-Examiner.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Benchmarking Foundation Models with Language-Model-as-an-Examiner

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.180727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.180727Z digest=sha256:aa38efe1d3ce29383ae5a8262298acad654f930cac10b95a4d63a85d12be20d4

Observation f97974cb-c8d3-44c4-a663-cad246c47a33 · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.604747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.185421Z digest=sha256:143f9b2173581c7ce061c9b524929ba64ff4f4d5826bb853a0cf20fafcfdb2f8

Observation 92ede92c-4fd9-40fe-afec-76154bfe11be · outbound

This paper cites J.; and Jurman, G.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models J.; and Jurman, G

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:29:54.591881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.189712Z digest=sha256:ccf2cd6964a7610e4a0b6d6bb33214d8049312078c6b1c4d362c2fd9b3ee62fb

Observation 95851375-d731-4b6a-9afb-76c012183f61 · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.580303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.193802Z digest=sha256:6a74bfe45ab405bf8986064b639fd91773bd1b0d1792f3037a28c7c709cdb0da

Observation 7f7695b5-5bea-4e40-a80e-11962a72be5f · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.569592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.197870Z digest=sha256:5e65e1cd242842680d93fe7c21a3539a9ca9840956bd34b69367a5b4049e1a1a

Observation 9862311d-7274-436e-a049-4d33d266bacb · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.557458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.201945Z digest=sha256:23703a21cee6a1bce73b9577b453162efd01e61e55531ce00fe5871bc5517b42

Observation 5736b260-f1f9-4fb1-a2ab-193f4afdaf42 · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.545583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.206207Z digest=sha256:83eeb1b74fd1b0c31e1c19cbc226e94153658cc9ca226e4422ea2ff4eab289c2

Observation eca368e3-d7ab-4f54-ace5-abb1cacc6d0b · outbound

This paper cites GPTScore: Evaluate as You Desire.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models GPTScore: Evaluate as You Desire

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.210042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.210042Z digest=sha256:b34fea841fa8b69769b0ecaa2d9ce2cb694ed2ba1d5a4b41d112d1e1f935ce6c

Observation 51f47c82-cf10-4bf2-98dd-215352389f4d · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.214139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.214139Z digest=sha256:7e0389f0be80bae21024bc0dd35b03e7b94f17d54fe9535f9b06d5937005b34b

Observation 6629d659-0bc5-404a-83cd-4b4a290c9877 · outbound

This paper cites Demystifying Prompts in Language Models via Perplexity Estimation.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Demystifying Prompts in Language Models via Perplexity Estimation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.218118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.218118Z digest=sha256:708c423b552cd3b23bc293e53a2280e50c153def5531d87f2bf72f02395f2cc9

Observation c6bcbf22-ed82-44d4-a177-9ab5992b46aa · outbound

This paper cites OLMES: A Standard for Language Model Evaluations.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models OLMES: A Standard for Language Model Evaluations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.222181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.222181Z digest=sha256:1a99a44c6b305496b76d99b9d3b18ed8682631d9a957806db51f567a76e35702

Observation 45e17965-2ffe-431c-bd3b-2ba97c179f1d · outbound

This paper cites Holistic Evaluation of Language Models.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Holistic Evaluation of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.226325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.226325Z digest=sha256:c59c0a372cd025dd58977dfff257af121876928060c1164486ef9c067f8f1351

Observation d0ec3e1e-495e-4972-8c14-7cd3f979194f · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.526086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.230138Z digest=sha256:7df2bcc24745a716401d46d1850fee84feefafe6d60d4e32614ec2f48cb04009

Observation 0b709059-0e66-46cb-adbc-02615707342e · outbound

This paper cites LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.233656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.233656Z digest=sha256:ef3bda04bfbc243af10d19570f3d382b919813ba2ec9121bcd384af88b0e6293

Observation e2fbde52-920a-424f-b224-4d129ac0a1a6 · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.514068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.237459Z digest=sha256:6c2c76848843e6d29bd1b52971d28f9523e8718018321bab510f6d5a21a1b758

Observation c37fd807-f80c-4f3c-a7b2-a4f2510030dd · outbound

This paper cites GPT-4 Technical Report.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.241101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.241101Z digest=sha256:618d4b3350896ae696f5473610b5db8fb3d1f4887d6a798ef9e704db0efc1bee

Observation b80bfecf-3dca-4998-bdb8-1e6f73568866 · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.502295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.245111Z digest=sha256:d9240ad71479d997e13cad7624579b9d580352a7845133645389b50e486236a0

Observation 68deed28-a79c-4e4a-9e8e-b625b5523a9a · outbound

This paper cites Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.248917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.248917Z digest=sha256:4ebba3afbf3b6111c634408d3a61ef0cfb750144e7025f31b3991b2ea25c0243

Observation 1c4e7523-582a-4ae9-a987-67840539f7b4 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.253765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.253765Z digest=sha256:981ee79ebdc151f5763599556514bb5b860246f0d5428b47df50a37acfc283dc

Observation afc954c1-e2ff-4f77-9be5-60d365c448c6 · outbound

This paper cites L.; and Parikh, D.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models L.; and Parikh, D

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:29:54.491310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.257663Z digest=sha256:0e2351c3e90c5e81b2485ae08d1ae72f84a3ed3787dd99d2a820734e5c8f4eea

Observation ec311917-7048-4f14-a72a-bea9dfee9684 · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.479863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.261221Z digest=sha256:d0bdebe5e67d0888996ae03eeb1bee65cd5cca86bce0b13b422274fc967270a6

Observation fd272bb9-6ebb-4a7b-aa4e-7e87dcc15d8e · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.467074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.264645Z digest=sha256:bd66123bf7804795d20ddaa225e6a0ae5e5cf93a5858658764aab6b4f324e02f

Observation 068c3815-3ce9-4658-9630-25aa6eacc72c · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.454739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.268688Z digest=sha256:6fa6999a823e9aa1b90da8c9bb9139c30f9ea381e6c47a6ffb41dc81efb9d09c

Observation 3aa9fed1-7440-45ca-b0bd-3b2d6da62464 · outbound

This paper cites A Systematic Evaluation of Large Language Models of Code.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models A Systematic Evaluation of Large Language Models of Code

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.272514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.272514Z digest=sha256:86b46c2ff095daa9e3c81055e8bfe219ae04e69d85861e2823f079a6da26fcf3

Observation 05efc8a0-510d-4f6a-b0ff-b90056466895 · outbound

This paper cites EvalAI: Towards Better Evaluation Systems for AI Agents.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models EvalAI: Towards Better Evaluation Systems for AI Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.276421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.276421Z digest=sha256:fe4d08d19e561c0d9244de84e87a626cb96034b7415eb233858728eecc798760

Observation 333d1227-68fb-4d2e-8c22-969d030b4255 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models BERTScore: Evaluating Text Generation with BERT

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.280278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.280278Z digest=sha256:f6be1c3af79e207f642b928bba510ff2977e3ce03e057da1f11afdce95ba8225

Observation f21dfe25-205d-4a58-be6d-cef4100443c4 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.284637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.284637Z digest=sha256:7153d6309dbbb4edd2011e65e3331d3c27cd2cf7432ffeaed943304069770351

Pith citing papers

Observation 59e10ad5-fa82-4eba-8981-577f18095a8d · inbound

Self-Aligned Reward: Towards Effective and Efficient Reasoners cites this paper.

Self-Aligned Reward: Towards Effective and Efficient Reasoners PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.527457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T18:27:23.076544Z digest=sha256:8e6ac7176596c54fff00cbdf5f69950ee80d4fa6999d88957e418a31e825ad0b