Pith. sign in

Paper Citation Record · LEDGER

Toward an Evaluation Science for Generative AI Systems

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2503.05336.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.05336 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:35.042357Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f44943f3-cb76-4bfa-aeeb-bf86bf08a589 · inbound

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation cites this paper.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Toward an Evaluation Science for Generative AI Systems

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.035975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.035975Z digest=sha256:a06322c7483b4ec57a069d644fea827e3f263eef703fe43734cd28645678b2c6

Observation 00bf41d9-f51f-42f5-b628-c437707f753c · inbound

Adultification Bias in LLMs and Text-to-Image Models cites this paper.

Adultification Bias in LLMs and Text-to-Image Models Toward an Evaluation Science for Generative AI Systems

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:53.536534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:53.536534Z digest=sha256:48c6728c7c016ebdbb1176a5f4e2e9e3947e8f250f416cc6b7f8c2fa43d42276

Observation 67010805-7da1-477f-b792-654e356efed1 · inbound

NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models cites this paper.

NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Toward an Evaluation Science for Generative AI Systems

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:24.091055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:24.091055Z digest=sha256:74533d430298ccdb7b69f0a7876e2a73db4e5db82f6abc369e4a990ea3fdb3c5

Observation b2f468dc-1180-48b3-b383-c073f710ffd0 · inbound

Correlated Errors in Large Language Models cites this paper.

Correlated Errors in Large Language Models Toward an Evaluation Science for Generative AI Systems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.341173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.341173Z digest=sha256:9eb2b040205f0d47648f91a023459769586f0b862651ca24ee16882db675ef84

Observation 7ea0c3d7-609e-468a-b979-7bd66ba0adf6 · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations Toward an Evaluation Science for Generative AI Systems

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:52.454150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:52.454150Z digest=sha256:66173ed992bbfc32ba0319677c6a236e18930d96f18c9423fc6330869d21ba40

Observation ec3598ac-b114-4790-9e78-9e1066a36c88 · inbound

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead cites this paper.

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead Toward an Evaluation Science for Generative AI Systems

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:12:55.745328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T02:12:48.586913Z digest=sha256:1d2356cd62cb6fc078ccbfccd7cb12714a62d08f9bfab4866a738a3da07a4bb9

Observation 828d7cf6-478b-4f3c-8fac-08bb24bb253c · inbound

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants cites this paper.

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Toward an Evaluation Science for Generative AI Systems

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:57.394750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:57.394750Z digest=sha256:db974b0e61366181623dc3a0faa983cd27b8f2707a8f710fa836880a2949b278

Observation 30b72dfc-149c-4062-9953-cf756c0ac8fa · inbound

Participatory AI: A Scandinavian Approach to Human-Centered AI cites this paper.

Participatory AI: A Scandinavian Approach to Human-Centered AI Toward an Evaluation Science for Generative AI Systems

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-04T16:37:30.927757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:37:30.927757Z digest=sha256:7233ea9720b7becd644047f17b37336821aabdca11b89ea93ba65c1bc95a216b

Observation d1c9d7d6-75c0-484d-90eb-1f5255204b8b · inbound

Participatory AI: A Scandinavian Approach to Human-Centered AI cites this paper.

Participatory AI: A Scandinavian Approach to Human-Centered AI Toward an Evaluation Science for Generative AI Systems

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T16:37:30.933003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:37:30.933003Z digest=sha256:0f648ea2769eb531a2c3dcd52bb03569ff4b2b479b6cef9d4cf5c5629528714d

Observation f6a775fc-c6c1-4af5-afe9-ea84f218b4db · inbound

An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications cites this paper.

An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications Toward an Evaluation Science for Generative AI Systems

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:12:39.805750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T14:12:08.776876Z digest=sha256:7df3c29cd71129ea28d469c58d93a47b31ef80a7bd2eeb760db0a9ba97ec9633

Observation 40b6930b-71b5-46b5-a009-952795f8cf87 · inbound

The Persistence of Cultural Memory: Investigating Multimodal Iconicity in Diffusion Models cites this paper.

The Persistence of Cultural Memory: Investigating Multimodal Iconicity in Diffusion Models Toward an Evaluation Science for Generative AI Systems

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:55:20.294670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T21:52:26.944212Z digest=sha256:719bcbeb242e7e40f23be4cf496c9ca1d31c20189c07a6768f06e704d68ac453

Observation eff0a74f-f146-4145-8d11-e8b26b292e0d · inbound

Making AI Evaluation Deployment Relevant Through Context Specification cites this paper.

Making AI Evaluation Deployment Relevant Through Context Specification Toward an Evaluation Science for Generative AI Systems

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:05.167848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T14:46:09.944168Z digest=sha256:052f35898c666ac4188bb4f5aaed7cb668855b217083325ceaaa49de2a19ad07

Observation 2e3f3e52-cd17-4f64-82e5-d5df6f151ba4 · inbound

Towards an Evaluation Methodology for AI in Second Language Education: Lessons Learned from Developing L2-Bench cites this paper.

Towards an Evaluation Methodology for AI in Second Language Education: Lessons Learned from Developing L2-Bench Toward an Evaluation Science for Generative AI Systems

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:25:23.882626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T06:20:32.174538Z digest=sha256:db50bfe969f346810623d23ce97a56eda96be9e1a03d584110499f373c7bb3f2

Observation ec26a659-b467-44f0-963d-666ff420317c · inbound

Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research cites this paper.

Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research Toward an Evaluation Science for Generative AI Systems

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:40:16.529102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T04:38:00.482850Z digest=sha256:75824aa87dd9f2a74e86f0a3b3144ae685ad9b59633e227674ca419a9922a0b5

Observation c2da2f11-fed6-4085-aeee-ee23ff3e76bf · inbound

Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing cites this paper.

Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing Toward an Evaluation Science for Generative AI Systems

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:53:23.366771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:52:17.310177Z digest=sha256:6f723509825d9a2c27b22cc65465a8a2cfd15bb0f2c41acf0097ab096f9223c7

Observation d5deee66-5baf-4f67-bdf2-f7cf1ddf2c66 · inbound

BenCSSmark: Making the Social Sciences Count in LLM Research cites this paper.

BenCSSmark: Making the Social Sciences Count in LLM Research Toward an Evaluation Science for Generative AI Systems

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:51:08.685249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:02:15.446874Z digest=sha256:3825d727af6ec8cd0f57d0538802ead0c84f1314f3b81c71cc5a25277f7490c4

Observation 63ac6204-afa9-429b-b0d0-b83fed439a25 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Toward an Evaluation Science for Generative AI Systems

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:28.016305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:34:55.538935Z digest=sha256:8356f0ff7d7b09c29c3e96a163b9fabe803b927ea54b5471bbe38f0cfd13e582

Observation 22ee5cdf-d680-43df-a261-32de1b528991 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Toward an Evaluation Science for Generative AI Systems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T14:22:26.753523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:22:26.753523Z digest=sha256:40cc231b51d771142c8e4668e6793715c754e44a9523782f89fae35659a94936

Observation 9d80f729-8b14-47c4-8441-e48c95297b47 · inbound

Unsteady Metrics and Benchmarking Cultures of AI Model Builders cites this paper.

Unsteady Metrics and Benchmarking Cultures of AI Model Builders Toward an Evaluation Science for Generative AI Systems

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.451740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T04:54:26.888562Z digest=sha256:0de7aa232564551fb4e715b6d26ac4fbc2b8c74c44805a5897e871defa8952c7

Observation 5fdd04f2-1eb7-4d1e-8b2a-1f4fae503f23 · inbound

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting cites this paper.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Toward an Evaluation Science for Generative AI Systems

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T02:07:33.265081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:e5b8127e63c1b57f0a2a9789080d8093119b344a25df42182021e8db52f3314f

Observation 6f756615-b9db-4248-ba42-eb1faca750f8 · inbound

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data cites this paper.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Toward an Evaluation Science for Generative AI Systems

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:54:44.494494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:965ce9b7e30ab9015e8114f44adbaefd2ee5d131681c69838dce46a4a6bd89c6

Observation af1ad875-3399-41f1-a4ac-2653c7c2e5cd · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Toward an Evaluation Science for Generative AI Systems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:03.614906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:03.614906Z digest=sha256:0230c2626ce70c9ff2bf96661ccb34b5dd47631469c275e57f907b111a2a2864

Observation d834dcae-40d0-4cbc-95ab-648cf68eb808 · inbound

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) cites this paper.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Toward an Evaluation Science for Generative AI Systems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.042357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.042357Z digest=sha256:1a5f9597a4e9a14f0a613594e1947020b178837e4bf5e6a64c1417f905b8c346