Pith. sign in

Paper Citation Record · LEDGER

Toward an Evaluation Science for Generative AI Systems

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2503.05336.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.05336 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:35.042357Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f44943f3-cb76-4bfa-aeeb-bf86bf08a589 · inbound

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation cites this paper.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Toward an Evaluation Science for Generative AI Systems

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.035975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.035975Z digest=sha256:0dd3cf2675611511357114b40285a91fca285740c409b26ce41204d8fa61344d

Observation 00bf41d9-f51f-42f5-b628-c437707f753c · inbound

Adultification Bias in LLMs and Text-to-Image Models cites this paper.

Adultification Bias in LLMs and Text-to-Image Models Toward an Evaluation Science for Generative AI Systems

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:53.536534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:53.536534Z digest=sha256:d880442a783e433f1eb0ab821971ed944d6a3f4825c6cc9f9f4a6706a5a4799c

Observation 67010805-7da1-477f-b792-654e356efed1 · inbound

NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models cites this paper.

NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Toward an Evaluation Science for Generative AI Systems

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:24.091055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:24.091055Z digest=sha256:70bae84e5184cdd87267bad27c22366b024de22a0a1dfe45bb71e983e548e24e

Observation b2f468dc-1180-48b3-b383-c073f710ffd0 · inbound

Correlated Errors in Large Language Models cites this paper.

Correlated Errors in Large Language Models Toward an Evaluation Science for Generative AI Systems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.341173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.341173Z digest=sha256:9a6d3c2dc4f21c6de17eb853bffec4ba375a7c9bfb4ed58ca942dfe73feca9f0

Observation 7ea0c3d7-609e-468a-b979-7bd66ba0adf6 · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations Toward an Evaluation Science for Generative AI Systems

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:52.454150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:52.454150Z digest=sha256:6f08da916cbf1ab918320c9757ac0e9fd4ae87dc70b98fee5cc5b6418e92d4c5

Observation ec3598ac-b114-4790-9e78-9e1066a36c88 · inbound

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead cites this paper.

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead Toward an Evaluation Science for Generative AI Systems

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:12:55.745328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-19T02:12:48.586913Z digest=sha256:7b598ae11e9b7fe4b0e1ef301191c3d5bb49bc4dee6802a3c825e4087bad833e

Observation 828d7cf6-478b-4f3c-8fac-08bb24bb253c · inbound

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants cites this paper.

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Toward an Evaluation Science for Generative AI Systems

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:57.394750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:57.394750Z digest=sha256:aca3fd5d5790c351ef60b5b913ba3e4e726c10a690bbe2ba1c8848428e058dd1

Observation 30b72dfc-149c-4062-9953-cf756c0ac8fa · inbound

Participatory AI: A Scandinavian Approach to Human-Centered AI cites this paper.

Participatory AI: A Scandinavian Approach to Human-Centered AI Toward an Evaluation Science for Generative AI Systems

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-04T16:37:30.927757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:37:30.927757Z digest=sha256:44e36bb5abc0783c322e83a14035e74173ec44fabbbaf8a6c08524c73937656c

Observation d1c9d7d6-75c0-484d-90eb-1f5255204b8b · inbound

Participatory AI: A Scandinavian Approach to Human-Centered AI cites this paper.

Participatory AI: A Scandinavian Approach to Human-Centered AI Toward an Evaluation Science for Generative AI Systems

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T16:37:30.933003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:37:30.933003Z digest=sha256:734550904a22092a4194b1a9d70e7d17b08a81d3680e33989c61fece3d9a645c

Observation f6a775fc-c6c1-4af5-afe9-ea84f218b4db · inbound

An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications cites this paper.

An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications Toward an Evaluation Science for Generative AI Systems

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:12:39.805750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T14:12:08.776876Z digest=sha256:95fd5c2e9ff05450b106629e5a39a036775144ea427883de15ddb0f302976520

Observation 40b6930b-71b5-46b5-a009-952795f8cf87 · inbound

The Persistence of Cultural Memory: Investigating Multimodal Iconicity in Diffusion Models cites this paper.

The Persistence of Cultural Memory: Investigating Multimodal Iconicity in Diffusion Models Toward an Evaluation Science for Generative AI Systems

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:55:20.294670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-17T21:52:26.944212Z digest=sha256:7de6a0055c41ff4d4b96c4827b45ca5a95384459a89f8ca3479d085ef66b126d

Observation eff0a74f-f146-4145-8d11-e8b26b292e0d · inbound

Making AI Evaluation Deployment Relevant Through Context Specification cites this paper.

Making AI Evaluation Deployment Relevant Through Context Specification Toward an Evaluation Science for Generative AI Systems

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:05.167848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T14:46:09.944168Z digest=sha256:33ac1c25e5156acf2d1d10e26651ecc4f63dba2e07b4f36a66d875c25cd5749a

Observation 2e3f3e52-cd17-4f64-82e5-d5df6f151ba4 · inbound

Towards an Evaluation Methodology for AI in Second Language Education: Lessons Learned from Developing L2-Bench cites this paper.

Towards an Evaluation Methodology for AI in Second Language Education: Lessons Learned from Developing L2-Bench Toward an Evaluation Science for Generative AI Systems

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:25:23.882626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-25T06:20:32.174538Z digest=sha256:71d5cf0405013316f61e2029b15214338752c9096e74b6acfc6134b1877f2561

Observation ec26a659-b467-44f0-963d-666ff420317c · inbound

Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research cites this paper.

Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research Toward an Evaluation Science for Generative AI Systems

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:40:16.529102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T04:38:00.482850Z digest=sha256:c68c9ef5df62354ae1c68c24e89f18f41efc4431a3e40be81008066edde6cc91

Observation c2da2f11-fed6-4085-aeee-ee23ff3e76bf · inbound

Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing cites this paper.

Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing Toward an Evaluation Science for Generative AI Systems

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:53:23.366771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T22:52:17.310177Z digest=sha256:e6f4fda712474457c526cc3c78c2bed5974ebe6e934a218b2fbfbdba02e2e1ad

Observation d5deee66-5baf-4f67-bdf2-f7cf1ddf2c66 · inbound

BenCSSmark: Making the Social Sciences Count in LLM Research cites this paper.

BenCSSmark: Making the Social Sciences Count in LLM Research Toward an Evaluation Science for Generative AI Systems

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:51:08.685249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T17:02:15.446874Z digest=sha256:4c68ca9deb47ce088d531743096bcb3d74b7cbb690f9aa1fb82b620edb921b8a

Observation 63ac6204-afa9-429b-b0d0-b83fed439a25 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Toward an Evaluation Science for Generative AI Systems

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:28.016305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T03:34:55.538935Z digest=sha256:dc245aee77eb6737924a0ac849b4599fdb61b18302f53cfcf9868e70e5f6b53e

Observation 22ee5cdf-d680-43df-a261-32de1b528991 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Toward an Evaluation Science for Generative AI Systems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T14:22:26.753523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:22:26.753523Z digest=sha256:a6c992023162f685eec5b3b50e3d3cf3e66d6307bfacb841d83d4e986a265f61

Observation 9d80f729-8b14-47c4-8441-e48c95297b47 · inbound

Unsteady Metrics and Benchmarking Cultures of AI Model Builders cites this paper.

Unsteady Metrics and Benchmarking Cultures of AI Model Builders Toward an Evaluation Science for Generative AI Systems

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.451740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T04:54:26.888562Z digest=sha256:f72a7dfb243936e7736565832f09dbc64c1ab56c9f78fa914cd31078c5a441df

Observation 5fdd04f2-1eb7-4d1e-8b2a-1f4fae503f23 · inbound

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting cites this paper.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Toward an Evaluation Science for Generative AI Systems

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T02:07:33.265081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:aad0e9e0546e674738a723e209d05aaf095f6c3a206bdc83ea200c5f3e844dd4

Observation 6f756615-b9db-4248-ba42-eb1faca750f8 · inbound

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data cites this paper.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Toward an Evaluation Science for Generative AI Systems

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:54:44.494494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:e0430bce828d3e6ee8125ff7540179fce51944c1580827a132c7fd594600790c

Observation af1ad875-3399-41f1-a4ac-2653c7c2e5cd · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Toward an Evaluation Science for Generative AI Systems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:03.614906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:03.614906Z digest=sha256:34aefcc09d918583456745b1d048bae5423b24cf8651d88c81bf57ec366e968a

Observation d834dcae-40d0-4cbc-95ab-648cf68eb808 · inbound

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) cites this paper.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Toward an Evaluation Science for Generative AI Systems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.042357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.042357Z digest=sha256:07bffac2e19226f94ad7779428639230f823422ac101bd843c3537c2a7afba66