Pith. sign in

Paper Citation Record · LEDGER

State of What Art? A Call for Multi-Prompt LLM Evaluation

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2401.00595.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.00595 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:59:49.571352Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9c1c3d5e-eab5-40a0-b9de-b33195ebaef3 · inbound

Revisiting Sentiment Analysis for Software Engineering in the Era of Large Language Models cites this paper.

Revisiting Sentiment Analysis for Software Engineering in the Era of Large Language Models State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-24T06:13:59.976576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-24T06:12:38.139005Z digest=sha256:0872510c2d5c878d6773dc6206ddabd2b22abdbdf0a2175a9a6a0347ec15058e

Observation a1f4b398-d6c4-4504-8d34-9ade7a41a5a4 · inbound

CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution cites this paper.

CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:57:16.378396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T20:57:16.327963Z digest=sha256:9db4b25a79a4a62b644fb020e7ae8cac79fd009ddb3c2fe8be2531c35a477ec6

Observation 003918c0-7574-4de8-9d64-a19d8f777951 · inbound

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code cites this paper.

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:34:43.111481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T17:34:42.565806Z digest=sha256:fc8dd53e7028b4097273090777ce670c685b875bc2338735dbda3c3571fdc72e

Observation e7fe15f2-ee82-42fd-9a94-9449eb690554 · inbound

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models cites this paper.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:08:45.143108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:e7479c354c54f301fdc53218797bd946239268ff2231de6abf8778092ce8e5dd

Observation 004717d1-fbf6-4a2a-b95b-438caa3ec91d · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 137

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T18:44:49.776794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:af5530d5cfb4460615f9f299067de5358c2d845c8149fdcb8a86f11d78d5c3b2

Observation 67f2ded3-90f4-4e78-aa56-2b349ed0a4ee · inbound

JuStRank: Benchmarking LLM Judges for System Ranking cites this paper.

JuStRank: Benchmarking LLM Judges for System Ranking State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:49.571352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:49.571352Z digest=sha256:af1087e53b415ce1e84b9ad85882e6dc5c70d6b5392a72771a99046a2d910ae0

Observation 9b4d1b1f-c45e-4d0d-bf19-4c528f644464 · inbound

Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative cites this paper.

Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:47.279466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:47.279466Z digest=sha256:98d540e115da396a2c743123e49a23b2bf140deb6ffe6f821a64b40da88d49c6

Observation 810e7897-981a-4a76-afef-92a24701ff33 · inbound

LCTG Bench: LLM Controlled Text Generation Benchmark cites this paper.

LCTG Bench: LLM Controlled Text Generation Benchmark State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T13:53:45.078859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:53:45.078859Z digest=sha256:f4660490930d5ff7795e323a3afde7d1d6956e3cf0eb8b5e699fcefb0fb361aa

Observation ebf17e8a-803a-46f9-9a86-1f73a7e839db · inbound

Evalita-LLM: Benchmarking Large Language Models on Italian cites this paper.

Evalita-LLM: Benchmarking Large Language Models on Italian State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T12:43:45.343766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:43:45.343766Z digest=sha256:68dbcdf337622d03c992713915a960c080b2771bc71a48abc4a07eff2d0b5d33

Observation fff97b18-4589-4e30-a1df-0d7823c72029 · inbound

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation cites this paper.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.118860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.118860Z digest=sha256:fae1b36788137db95a6efaebde552d7be5fa7d515f7172cd430e432f51a1c8d7

Observation 8f9febc8-c580-4f34-857e-c867ebda099f · inbound

Personalizing Education through an Adaptive LMS with Integrated LLMs cites this paper.

Personalizing Education through an Adaptive LMS with Integrated LLMs State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:50:19.286087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:50:19.286087Z digest=sha256:ac7009fd546a7ad849da4604be63aaee8450745edc8efb3a4d4de30ff7561e3f

Observation 853fbbd7-28bc-4a13-adbd-0e86dd15e784 · inbound

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity cites this paper.

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T22:04:18.072955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-08T10:23:02.697982Z digest=sha256:715711cde2fa1381773699ca80e741eb73386991dcec685e3ff849746611efc6

Observation 7d7e74f0-daba-4236-9275-bd8ee43cafac · inbound

Predicting Performance of Symbolic and Prompt Programs with Examples cites this paper.

Predicting Performance of Symbolic and Prompt Programs with Examples State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T01:10:51.491532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T01:10:01.044650Z digest=sha256:92027b90f7264b4add98a8cc74b84b6db7285876426caf066bb204c4308ee164

Observation 6d86b257-40e5-4b84-81d0-4713d9f46cc5 · inbound

SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks cites this paper.

SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:54:00.986796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-29T22:50:48.263600Z digest=sha256:bba7136ffd70e734f584d8ae157ec498831709d31a488e4dd60ac05e630457a3

Observation 7d68ccd6-9171-44d5-b05f-9d322032e90c · inbound

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting cites this paper.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T02:07:33.257770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:360b9ae97c92a51846849fc1475d9fae4e312ed15f4e6161aacd8bedb3ed1a33

Observation 51592a10-a247-47d2-a4b1-a9a8a46924ea · inbound

Latent Confidence Alignment for LLM Self-Assessment cites this paper.

Latent Confidence Alignment for LLM Self-Assessment State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:39:41.782645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T11:21:33.744610Z digest=sha256:60c2e597400d6b49016d38aaf00facc5a5bf8e8d73b1c0d731e30725c53538b8