Pith. sign in

Paper Citation Record · LEDGER

Evaluating Generative AI Systems is a Social Science Measurement Challenge

As of 22 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 11 inbound Pith citation observations for arXiv:2411.10939.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10939 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:12:55.209578Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:42:53.504979Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T03:41:00.103963Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c515841-4d64-40e6-9d02-41c64f387539 · outbound

This paper cites YouTube Hate Speech Policy.

Evaluating Generative AI Systems is a Social Science Measurement Challenge YouTube Hate Speech Policy

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-12T19:12:55.571058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T19:12:55.096303Z digest=sha256:4b8781655b5df4f68d03a8ebe1b384dd603f8ae9a7ed5bc9430470a3d8f60c8b

Observation 91d834e9-b16e-4393-9a37-040723445a03 · outbound

This paper cites Measurement validity: A shared standard for qualitative and quantitative research.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Measurement validity: A shared standard for qualitative and quantitative research

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.101173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.101173Z digest=sha256:da63d67571cff15a843d4cbf1de3060d1fd8e55dac623a83b277a58d04853f19

Observation b064fff7-1647-44f2-88fe-f573527e55fc · outbound

This paper cites Content analysis in communication research, 1952.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Content analysis in communication research, 1952

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.779337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T19:12:55.105802Z digest=sha256:5d61e3655fff2fb1c8c21208d9976362a06d47e07acca6a3ffe81f7e4851f98e

Observation 32835267-c04d-4a63-b5ff-2251fa531963 · outbound

This paper cites Making Intelligence: Ethical Values in IQ and ML Benchmarks.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Making Intelligence: Ethical Values in IQ and ML Benchmarks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.110397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.110397Z digest=sha256:036ffada15ae44b59f025b3fe216a1109eacc7d4a8d34bcacc2479a708fb85f9

Observation add1c81a-962e-4128-9e3f-b6108121cbff · outbound

This paper cites Sociolinguistically Driven Approaches for Just Natural Language Processing.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Sociolinguistically Driven Approaches for Just Natural Language Processing

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T19:12:55.763447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T19:12:55.114777Z digest=sha256:4372fda81c433b035c3381c41ee6d1ce6fa848ff1206a9efa2974b8638d30049

Observation 101e6a60-6366-44ab-8b14-386973d310db · outbound

This paper cites Language (technology) is power: A critical survey of ‘bias’ in nlp.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Language (technology) is power: A critical survey of ‘bias’ in nlp

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.748677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T19:12:55.120069Z digest=sha256:831bec1c0d62ad1e068fe6c7afddfeb71c613d95164a170366f905e06c1c4e57

Observation 1bc8f043-44c9-4829-a536-0649cd83b47f · outbound

This paper cites Stereotyp- ing norwegian salmon: An inventory of pitfalls in fairness benchmark datasets.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Stereotyp- ing norwegian salmon: An inventory of pitfalls in fairness benchmark datasets

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.733739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T19:12:55.125557Z digest=sha256:2ccc3716a0c8898fd3bf4c3fde246f91c15decee10eb1192fac788d29249aa8c

Observation 38afce52-aa80-413e-8b11-abab2ad3eda6 · outbound

This paper cites DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models.

Evaluating Generative AI Systems is a Social Science Measurement Challenge DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.131216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.131216Z digest=sha256:51a01e3de0f6cb06c9019f866c665a171cbc4057fffc5950dafba9729a0466f1

Observation 986f4050-05e4-4824-b4b9-a11795e3d20b · outbound

This paper cites Feder Cooper, Ellen Abrams, and NA NA.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Feder Cooper, Ellen Abrams, and NA NA

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.136919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.136919Z digest=sha256:fc7ed0af09af593f73bd99e4b76be2763269a9bef2576129ef0ca1f065a69b16

Observation 90b5c918-b4b2-44d1-852e-9292b1a017a6 · outbound

This paper cites Report of the 1st Workshop on Generative AI and Law.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Report of the 1st Workshop on Generative AI and Law

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.142295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.142295Z digest=sha256:dbdff0d5798bd18ed6d3a738be578784049554fbd7f8e3fcb233c8ce778e7fb8

Observation 8467d862-bdb5-40a0-9609-f968ebab6d99 · outbound

This paper cites Representational harms through the lens of speech act theory.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Representational harms through the lens of speech act theory

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.147508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.147508Z digest=sha256:33b9169e46e6c9a96e3a5af2c7b3192101181c52ff1a8fcdd61ae56dc74dbacb

Observation 75e2a104-d303-4c69-9809-96810cc6ed3b · outbound

This paper cites Construct validity in psychological tests.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Construct validity in psychological tests

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.707824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T19:12:55.151808Z digest=sha256:49d6343b6bb6be3e67ce92405ba3aeadec0934c6e59e87d17fa1cb0d03a472b9

Observation 5f8e446b-2916-46e3-947c-dbfd020ed0e0 · outbound

This paper cites Measurement and fairness.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Measurement and fairness

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.692482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T19:12:55.156530Z digest=sha256:0d363bc14af2a542d2e6edcf9844c336e43d147685faa08ea8c0e383d4f47340

Observation d8dd5ef3-fb54-4dd9-a304-c9c7657961a8 · outbound

This paper cites ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation.

Evaluating Generative AI Systems is a Social Science Measurement Challenge ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.160667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.160667Z digest=sha256:fbdedca4011bd53bf4075a67af5bcaacf5dceca00844239b5f37579240fc0fff

Observation 648da852-aae6-476c-b7ce-e6222b1123c2 · outbound

This paper cites Hate speech in public discourse: A pessimistic defense of counterspeech.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Hate speech in public discourse: A pessimistic defense of counterspeech

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.676896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T19:12:55.165191Z digest=sha256:2af436a87fb48a435079d1abbab5aa4290a26e851b2eb1aa083da7bdc01d025c

Observation 22e50b36-cf29-499a-b48c-59a449cd2b04 · outbound

This paper cites Vera Liao, Alexandra Olteanu, and Ziang Xiao.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Vera Liao, Alexandra Olteanu, and Ziang Xiao

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.662480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T19:12:55.169333Z digest=sha256:9952acbbf90381a3a55e94fb2c65aabbe966a998e24b640c53e2944da8169c55

Observation 2f209101-db34-418b-a817-9002c6ae1989 · outbound

This paper cites Validity and washback in language testing.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Validity and washback in language testing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.647368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T19:12:55.173553Z digest=sha256:ce78de01b6fef5217ae9836eebba82d26dc224c9e07142facc3b3d846c18db76

Observation 1fe3b165-cfb7-48bb-8116-eddcb5d620e2 · outbound

This paper cites Privacy is an essentially contested concept: a multi-dimensional analytic for mapping privacy.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Privacy is an essentially contested concept: a multi-dimensional analytic for mapping privacy

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.632568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T19:12:55.177711Z digest=sha256:4fe616d9394812fe2bf2481fad3870e5a5b07a96fe52dc1364349f35de1281d1

Observation a6ab75ab-7f1b-430b-8c02-5cd9e255dc96 · outbound

This paper cites Mulligan, Joshua A.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Mulligan, Joshua A

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.618323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T19:12:55.181927Z digest=sha256:a9337c1ddced7e77a23bed5623277aeb4365c28bd4fd2946819529896da57ff5

Observation 7ede09c7-c4b9-41b5-82bf-25ba3bad6834 · outbound

This paper cites StereoSet: Measuring stereotypical bias in pretrained language models.

Evaluating Generative AI Systems is a Social Science Measurement Challenge StereoSet: Measuring stereotypical bias in pretrained language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.185915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.185915Z digest=sha256:5384faf036cde70857d04f56e4e7db756ec42fc9676b8cb54471ab0284d494b8

Observation bf0a1e00-ce68-4a96-b888-926f9f26f5e2 · outbound

This paper cites CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models.

Evaluating Generative AI Systems is a Social Science Measurement Challenge CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.190662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.190662Z digest=sha256:549954da10b188fe157ef7b2b47f4c6b8a31f4dcb41831f5f8fb276f37789635

Observation 0736818e-a22b-4792-b0b0-ccbbdf97d5c3 · outbound

This paper cites Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, 2024.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.602679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T19:12:55.195061Z digest=sha256:72876aede69b7180c8d3e30482f75599384da5ad9cf471b734c2894ecb190aea

Observation fecc7a41-e79f-4c13-bd96-80764b4160f1 · outbound

This paper cites Red Teaming Language Models with Language Models.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Red Teaming Language Models with Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.199612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.199612Z digest=sha256:06e268b848479974b03b5b9131574fe16a54dc1b7a119074549ad78ca072156d

Observation fba9a1d0-413e-4320-abf7-c05e07483ebd · outbound

This paper cites Evaluating General-Purpose AI with Psychometrics.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Evaluating General-Purpose AI with Psychometrics

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.204761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.204761Z digest=sha256:a8c7a45240c4f234024aa6605985ad76e4ae066e5a5e0ff93e0297593909429d

Observation 51eb7304-e85e-4cce-940b-4126f761dd64 · outbound

This paper cites The nature and origins of mass opinion.

Evaluating Generative AI Systems is a Social Science Measurement Challenge The nature and origins of mass opinion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.587253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T19:12:55.209578Z digest=sha256:9adf2c3ae410db4f5e4788851c0895c745c926e88b72b3e497740a5b726de9a4

Pith citing papers

Observation d9a316a6-17c7-4f4c-8413-eac0b080776c · inbound

Adultification Bias in LLMs and Text-to-Image Models cites this paper.

Adultification Bias in LLMs and Text-to-Image Models Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:53.504979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:53.504979Z digest=sha256:1d8b3d6bf0c19ec787aece9a4cc2149c76ecca294f49ac0727833454f74267d0

Observation dc84554f-f5ac-46c7-9ab3-8b4ce16d6715 · inbound

Correlated Errors in Large Language Models cites this paper.

Correlated Errors in Large Language Models Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.330863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.330863Z digest=sha256:5f60c18e16427b22995df024c884f5208f5afe138a58a209ffee69fe1d7627cd

Observation 5d511600-3fcc-4607-90f9-c6c2637fb8cc · inbound

Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks cites this paper.

Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:11.241009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:11.241009Z digest=sha256:cbfa3500ac6f2925ffd229a26bc6aa85ac1259804296e0c26a20edee9d801cc0

Observation 71c16f38-00e1-4ca8-9015-6337f292eea6 · inbound

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges cites this paper.

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-05T16:40:07.216514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:40:07.216514Z digest=sha256:864a7bb83c8735b8f0f9a23718f4be5c563040c981152434abf7c7c4154c3e70

Observation 6f3764f2-2522-4499-89c9-7d0de8f765e4 · inbound

Understanding, Protecting, and Augmenting Human Cognition with Generative AI: A Synthesis of the CHI 2025 Tools for Thought Workshop cites this paper.

Understanding, Protecting, and Augmenting Human Cognition with Generative AI: A Synthesis of the CHI 2025 Tools for Thought Workshop Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-05T14:38:55.710426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:38:55.710426Z digest=sha256:4826333d14c4fa40e4fffecf0478a9df332fbc257f3af895268b586c6d24feb7

Observation ef6e2618-610c-446f-948e-752f28d4d8e0 · inbound

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants cites this paper.

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:57.153507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:57.153507Z digest=sha256:300ad94d6b9e1bbc9f0ab4a4ac384e10a0bc9e62c4a21c00c89c9b72036351b2

Observation ea55a5d3-78b6-4562-babf-f5bdb69bda7b · inbound

Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation cites this paper.

Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:51:21.097511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-07T16:09:27.944432Z digest=sha256:b020b7c7ded2988dbd3af79ed538f1ef20fde0da6eba8623712a2315fdda69d4

Observation 2456370c-6165-4ec4-9876-0f4431a0d06b · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:27.754116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T03:34:55.538935Z digest=sha256:2c110bc935391931b6abfe016490a1dad79d7fcbbd423d70c81135cdc88f20df

Observation e537999c-a49b-449d-b305-8ca37468295c · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T14:22:26.223129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:22:26.223129Z digest=sha256:c226e02f4b049cb01999c4f48fd61bcef299deeab1827de764947f9512790f69

Observation 48a4c397-c291-4af2-ab40-b86e1b1d9eab · inbound

Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory cites this paper.

Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:43:43.866041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T19:40:59.534448Z digest=sha256:b2bfafdaaa96a75fd22bbc554d45d35a469d9addefed57d7571b2d60c7db1986

Observation 36e9d93d-e554-4dd3-852f-f020103e3331 · inbound

Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions cites this paper.

Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:41:00.107255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T03:40:25.093487Z digest=sha256:6f4e6e84bc72be91f05fd2a6bbb3b33e0a67dc1ff382b403935495196f3cc92b