Pith. sign in

Paper Citation Record · LEDGER

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

As of 23 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 25 inbound Pith citation observations for arXiv:2502.00561.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00561 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:37:35.044086Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T10:55:22.643993Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact4
  • verified fuzzy21
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation af360f01-f194-42f0-a655-21296375d600 · outbound

This paper cites write newline.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.334756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.334756Z digest=sha256:afddfacf8f4f6683057b35de166aa9ea59e515b3a4123a2a509cf1711204773a

Observation c3a72428-0c5a-45db-9d35-367cb535f16f · outbound

This paper cites an unresolved cited work.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:37:36.313238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.366575Z digest=sha256:984d0b699cf10850214b684ea9f86404ec88d5ea8a7f3490102d58e19bcd0dce

Observation 0c520125-24c9-4a83-b453-0c65b3e0cf3b · outbound

This paper cites and Collier, D.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge and Collier, D

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.305529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.376133Z digest=sha256:3244eb88f6be4cb639e158249c13e8bdc2c9f41baa0a8ce84c20ddac18d78b48

Observation c61fc32b-79d0-4d41-9eba-95e3b2bce4f5 · outbound

This paper cites M., Blaauw, G.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge M., Blaauw, G

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.297384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.392459Z digest=sha256:8d72d461be21245cae16ebdfd38402f85317c14c1a172fff6eb352db37bd47b2

Observation 39527063-c915-45a4-8539-6c0a09fc0642 · outbound

This paper cites Content Analysis in Communication Research.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Content Analysis in Communication Research

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.289599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.414744Z digest=sha256:e19fac28a582a8da0fb7c7031d9a7ba52e7f7fc94c5626000482fe68da43792d

Observation ce85ba0e-1412-42af-9494-c3869d47e314 · outbound

This paper cites and Hancox-Li, L.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge and Hancox-Li, L

Reference 6

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T18:37:36.095316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.464749Z digest=sha256:577b7e462533c7f178efa60696a2f2e0bc599faded376fa3c2ab10837032d304

Observation 84ee5103-17e9-4a37-a368-65f16d3d0b4f · outbound

This paper cites an unresolved cited work.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.493955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.493955Z digest=sha256:72642af04a92efd4be64a3b0968bee1d267356b5510a35aaafec44b720019170

Observation 961ef34c-c694-4c8d-9628-df68d30a99b9 · outbound

This paper cites L., Barocas, S., Daum \'e III, H., and Wallach, H.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge L., Barocas, S., Daum \'e III, H., and Wallach, H

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.554747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.554747Z digest=sha256:9680c00ae87712e3edc5024c826a87b141ef1ccde4e2e033f1a05847810b8289

Observation 826776d4-29a0-4e1e-90e3-8894b7e9db1e · outbound

This paper cites L., Lopez, G., Olteanu, A., Sim, R., and Wallach, H.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge L., Lopez, G., Olteanu, A., Sim, R., and Wallach, H

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.280942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.588034Z digest=sha256:473d11f34645bc6e0b19151a004d2a190491c3fdcb18c2f05fefcb45ab2e2499

Observation b02377aa-91dc-49f3-abb1-f9f2bfeea3a1 · outbound

This paper cites Quantifying Memorization Across Neural Language Models.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Quantifying Memorization Across Neural Language Models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.273013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.624753Z digest=sha256:3c35386c63899ba2f6287df0c320e00facc363a76fdf2f91fb4aebc29b1fe3e4

Observation 6337746e-a493-41dc-894d-2499a7b1b4f3 · outbound

This paper cites A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.652479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.652479Z digest=sha256:db48c985a71ea3f7c041ef9e3e0d8b7350d93be56112b1db55f8b02874f6c843

Observation ba1dd32b-79b3-45eb-a9e2-76fd3688f3ca · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Training Verifiers to Solve Math Word Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.671148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.671148Z digest=sha256:4678471b43c841b7060f6790e5f1c33c3e54f6d2439453fcc9efc810d45323a1

Observation 99c724f1-2cae-449c-b891-88601cc55c1c · outbound

This paper cites The Files are in the Computer: On Copyright, Memorization, and Generative AI.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge The Files are in the Computer: On Copyright, Memorization, and Generative AI

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.714751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.714751Z digest=sha256:db0294a05a45f18b1ca4cf53c463a8e2230725f0ee29a03aeb522476fb63722c

Observation 38e4f769-2792-4351-bd31-64e837d6975f · outbound

This paper cites F., Abrams, E., and NA, N.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge F., Abrams, E., and NA, N

Reference 14

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T18:37:35.910047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.741941Z digest=sha256:8b20af088e97e5f042bd72b970a5fc7f5c6fdd8179c9e111603b68702437dffa

Observation 3f3a0b43-c517-40a4-bcfc-43eb0bd17c61 · outbound

This paper cites Report of the 1st Workshop on Generative AI and Law.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Report of the 1st Workshop on Generative AI and Law

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.744625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.744625Z digest=sha256:1ec70c23c4e7db3b8979b645de566c8a4fba994976a3b816ac80521ec8031b2f

Observation 658e6011-7732-4d6a-acc7-e889b657d432 · outbound

This paper cites F., Choquette-Choo, C.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge F., Choquette-Choo, C

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.764744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.764744Z digest=sha256:166973b4f043f9c29943c49b31e3355f54d3e4340b8ae5b0ef9d297fdd3b21e8

Observation 2de9b6f7-da44-4db5-83b0-9ed3461d381b · outbound

This paper cites Representational Harms through the Lens of Speech Act Theory , 2024.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Representational Harms through the Lens of Speech Act Theory , 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.265664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.795120Z digest=sha256:244963e173dde1bd7491c08c9f6793e982265b344e2c4104caf733b7a06e1059

Observation 575ef757-7849-446b-8724-3d0ff34abbef · outbound

This paper cites an unresolved cited work.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.813953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.813953Z digest=sha256:3773d3528088b85d9f6d66e6da1270012a72bb9d6b92f76cbc7dd8cb21a2d001

Observation 404b3271-5f3e-4a39-a6c8-9ae5f194edb7 · outbound

This paper cites an unresolved cited work.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:37:36.258551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.834743Z digest=sha256:1fd8cebee58b26452555b89324552d5c9e191c9ce508d92fefa3c78341fad63d

Observation 02db8153-78d2-4c3d-a026-77c3dbe5cb33 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.857548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.857548Z digest=sha256:479f9df72c13b7f94833348a706df3fe6968c7b7d17366fe40b653cca2ab409b

Observation dd7a6bd3-4f26-4e26-9241-cc0068f82a2f · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.251091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.904753Z digest=sha256:9c2466ea915f53b2212ccccacb3f0e0ad2ede2df987436e940d94bcd1ac6ad67

Observation e10e8f16-f5ad-4c24-ac75-6ffade724942 · outbound

This paper cites Evaluation gaps in machine learning practice.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Evaluation gaps in machine learning practice

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.243144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.916078Z digest=sha256:42f53f6ffc6630fbcfab715ae49efb15120ce86dbef8db8afa88697a8bb45d4e

Observation 18deee14-b6ec-4f9e-bb1f-69152923759e · outbound

This paper cites an unresolved cited work.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:37:36.235017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.919196Z digest=sha256:7a0abd7d8db3da57e328f141406be3db6f812325670f94343f134015c12df1e0

Observation ed758256-a1eb-451a-ac4f-6e6e0cab1395 · outbound

This paper cites and Kieran, C.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge and Kieran, C

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.227355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.922662Z digest=sha256:9c091bdd8668eaba92efdc227f734d7fa97aed42e0058188516869a9e42a722b

Observation 24f98484-80c8-4eea-871a-b0108718f39a · outbound

This paper cites an unresolved cited work.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:37:36.219671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.934758Z digest=sha256:1ade94507f5f11dacfa757d4749789d8e96707dff51e1a2ce08b12cdb12d039d

Observation 21ac5ce4-37d8-4860-a76a-22adfac29cd3 · outbound

This paper cites Algorithm = Logic + Control.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Algorithm = Logic + Control

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.211802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.938084Z digest=sha256:e5b7ccf1ebe739ccee97b34c61acb0f7fac8ffb09168696560c9d151020d9d64

Observation c0610033-aafd-407e-91e4-d5e702fcea9a · outbound

This paper cites Talkin' 'Bout AI Generation: Copyright and the Generative-AI Supply Chain.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Talkin' 'Bout AI Generation: Copyright and the Generative-AI Supply Chain

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.941234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.941234Z digest=sha256:5548c829db9ace16a87bef5d58e276e2fbc78618ec8f0b3dc3e0118192ba919a

Observation 7dbb41d7-a389-4598-a98c-075fd6eec3c5 · outbound

This paper cites IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.204138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.944139Z digest=sha256:162e97ff37fc4a83ee683fd0b17a7a3974f5fa658ac50e9cacc53abb8142db14

Observation 03b7621a-3feb-4b76-b4f6-415f0c21f577 · outbound

This paper cites L., Blodgett, S.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge L., Blodgett, S

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.196239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.955891Z digest=sha256:a4c9ef5e63dbbd7b2151ad9ca598b151a6d0837bd3bb1eafb3c93371651741ad

Observation 244fdae9-00bb-40ad-b6dc-2520ae733768 · outbound

This paper cites C., Shoham, Y., Wald, R., and Clark, J.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge C., Shoham, Y., Wald, R., and Clark, J

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.188290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.958473Z digest=sha256:945e5dd74972aa3065277a571a0d5a0eb0f5e6dd6076049cd7826c1ff7a65437

Observation ffd0c4f9-3f0f-473f-b074-081f813becc6 · outbound

This paper cites Validity.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Validity

Reference 31

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T18:37:35.561608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.960900Z digest=sha256:9dc315972852b911bf00affce322aea497358b5dbc8e08147061c97f3983ea8d

Observation 721c0e54-0bf0-480a-a5aa-31a32d060307 · outbound

This paper cites Validity and washback in language testing.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Validity and washback in language testing

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.180240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.966776Z digest=sha256:9dbb3c154640306d14a6294cec6a4c94665563dacb89d57bcf90a9f14910f9b6

Observation 0bc8f29e-77e0-4a01-8b5f-43967daa8284 · outbound

This paper cites K., Koopman, C., and Doty, N.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge K., Koopman, C., and Doty, N

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.171951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.969708Z digest=sha256:f5755053ae417fa391aa2c8ebb440b6ba0ce2e4a403cb184fe477523b6a3e88b

Observation 344c7471-839e-4ea2-8c4f-1c15538d3559 · outbound

This paper cites K., Kroll, J.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge K., Kroll, J

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.164417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.972247Z digest=sha256:66d72e04ce9b0f998c231ada7f614fe37f5ab4aec3a41de654c4572c0c68d5f9

Observation 83e9b15e-a9a3-47a9-ab88-503bfa2821e0 · outbound

This paper cites S tereo S et: Measuring stereotypical bias in pretrained language models.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge S tereo S et: Measuring stereotypical bias in pretrained language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.974821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.974821Z digest=sha256:da744020084e23e59fdf4656d93b5d87aa3523309f83261d9af661b574f46d05

Observation 10219736-ab27-4493-9b88-dfea96510fa0 · outbound

This paper cites an unresolved cited work.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.979147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.979147Z digest=sha256:5df280fb2d34ce237c37131273a6b9a7c6736b21410d69456bb5b89999a45732

Observation 148f9c7c-d07f-469d-a0b8-d5a20ca2c6dd · outbound

This paper cites Scalable Extraction of Training Data from (Production) Language Models.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Scalable Extraction of Training Data from (Production) Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.981842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.981842Z digest=sha256:5cf725499504334ebdf7b56bd393d0873461c215bfa96ae264ecafefc6030b15

Observation 7310aeb2-4b35-4655-9145-f52c6d08b9ba · outbound

This paper cites Artificial intelligence Risk Management Framework: Generative Artificial Intelligence Profile , 2024.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Artificial intelligence Risk Management Framework: Generative Artificial Intelligence Profile , 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.985743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.985743Z digest=sha256:ea6e6243f90b4382a3410d0ec298522eafda15b7449d36ebc94003afc5de7da3

Observation 225cc190-4abc-4f3f-ab94-291330f67773 · outbound

This paper cites Learning to Reason with LLM s , 2024.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Learning to Reason with LLM s , 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.156633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.989173Z digest=sha256:c4bb4e7e9190f4eed5f073eadef235c485abd2cdbea50a5150b7df3532a98e06

Observation c886a326-5968-44f6-ba75-5e2112710404 · outbound

This paper cites Red Teaming Language Models with Language Models.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Red Teaming Language Models with Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:34.992794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:34.992794Z digest=sha256:dc1b13c3e2a09357c5b113cc1b0cf94ca2c6af5144e4515ef57c1819452918dc

Observation a17f79a7-4018-43c4-84ee-18b6a3e75a8f · outbound

This paper cites an unresolved cited work.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:37:36.148281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:34.998973Z digest=sha256:97c7b28dc9c1c557d67b1a7b8fe1ad50d2a95d2fbd8b6be18b18f345975aa3d1

Observation 903acbaa-441f-4a98-984a-0248a8204015 · outbound

This paper cites Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:35.001891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:35.001891Z digest=sha256:e2f94c9c31be73583973abe188536079405f15d5b6ce7cb16d630aa9332fe25a

Observation 0434809c-7011-4eef-8db9-fb274e6d6e80 · outbound

This paper cites D., Bender, E.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge D., Bender, E

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.139712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:35.014742Z digest=sha256:e9cc831ba43d6f1548bf8f00e47136cf07df0e35f164ee98dacc162f73099f1e

Observation 72d475a8-8b21-47d9-8acf-91b8957dc12e · outbound

This paper cites A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., Isaac, W., and Weidinger, L.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., Isaac, W., and Weidinger, L

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.131519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:35.017831Z digest=sha256:21ead202870c1140cb6fbed77e5f30ca49c2b5a45df8e500a2ba453cf9848574

Observation cf79d6bf-9d3d-46c6-b054-33008fe256b9 · outbound

This paper cites L., Gardner, M., and Singh, S.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge L., Gardner, M., and Singh, S

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:35.020661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:35.020661Z digest=sha256:9f8dc329dac044f13c4022f36be7f528c26337de0d1e3af0936ddd38f2fc2f0a

Observation 89f36e89-4f16-4c1d-99e6-413403b33433 · outbound

This paper cites an unresolved cited work.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:37:36.122583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:35.024293Z digest=sha256:a2e9ed4d91cabd300cef0d4feb0517838907759f49dc2bc36c73f6bc27018ecc

Observation 7ae7e0de-d88a-4109-8241-eddd2975649b · outbound

This paper cites H., Reed, D.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge H., Reed, D

Reference 47

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T18:37:35.391628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:35.027380Z digest=sha256:ae4a3f7c9672445353ce5290eb9c3dfa5820eb75dee5f35c0cb8cf651708dda8

Observation 98797b12-7942-4d4c-bfed-81d607f38ec6 · outbound

This paper cites Evaluating General-Purpose AI with Psychometrics.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Evaluating General-Purpose AI with Psychometrics

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:35.029933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:35.029933Z digest=sha256:39bb9b0cb70230aa8175d45405aef985818d3c5a5c29c6d511e0025216e3b0d8

Observation ff236a4b-d7c0-4b55-93af-de8f18809acf · outbound

This paper cites Sociotechnical Safety Evaluation of Generative AI Systems.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Sociotechnical Safety Evaluation of Generative AI Systems

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:35.032673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:35.032673Z digest=sha256:1cc842b314ab358e143144c9a5f926c65c2d206bf2846ad050bc95650961a2af

Observation 4c7fdeb2-04d4-4476-8641-764a7d9fc553 · outbound

This paper cites Evaluating Mathematical Reasoning Beyond Accuracy.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge Evaluating Mathematical Reasoning Beyond Accuracy

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.114024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:35.038652Z digest=sha256:e8bd0e3d717eda12e2eb7aa01be0f989771ba4ea059367eee6e1c12b53c1292d

Observation b04728b6-67b8-49c0-b678-3d8a12626f31 · outbound

This paper cites The nature and origins of mass opinion.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge The nature and origins of mass opinion

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:37:36.104441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T18:37:35.041282Z digest=sha256:d70040d401609c8dcea4312f2b7b17e03ce39fff74d5d1d77fae7c2f226b8e6c

Observation 4ac29549-4631-46cc-9990-bbdd755c6763 · outbound

This paper cites HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions.

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T18:37:35.044086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:37:35.044086Z digest=sha256:393aeabab7cec4c7f0023dae21ee2997238518cb37cafa0d3d4ec0d447c096ae

Pith citing papers

Observation 217ffa83-ae6a-410d-9b76-3623634a1317 · inbound

AI, Jobs, and the Automation Trap: Where Is HCI? cites this paper.

AI, Jobs, and the Automation Trap: Where Is HCI? Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T21:53:35.156634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:53:35.156634Z digest=sha256:09062c4b01ceb94230394ebb631f473f1dc76e3a64c5e71842a99266d0fedc47

Observation 340fdeae-2f50-444b-a453-d2f97bd837fc · inbound

Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs cites this paper.

Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-09T14:03:44.621750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:03:44.621750Z digest=sha256:7269a52573889339277ad75601feee966a3a1159ca7742d4594f2e1616c7ca9f

Observation 3a680a5c-61f8-4fb4-bc0a-890748ccd3c8 · inbound

Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects cites this paper.

Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:35.180337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:35.180337Z digest=sha256:2fe051374c74825cf3035949bff40c157c504c29582364717954b097018c38d6

Observation 8757af9d-2f7e-40ff-80ce-9462a93d6221 · inbound

Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems cites this paper.

Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:03.085054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:03.085054Z digest=sha256:03944bfe4c0335029bbb4bc6e765b3b05d8725acab6f9196ead7433467d82b1f

Observation cab93fa3-9b77-49b4-947e-f7162d8d95a8 · inbound

Against 'softmaxing' culture cites this paper.

Against 'softmaxing' culture Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:56:03.057445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:56:03.057445Z digest=sha256:01b854b71c348e7fb13ab5e18ac7efc5e2106a100787e3bf41aa92c18d689bc7

Observation 2c97c0cc-c696-4c0b-b594-a0669ea2943a · inbound

Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances cites this paper.

Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:49.516194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:47:49.516194Z digest=sha256:858d2442a110fd85c78304b0292a6c71c9f9b4de9e411307476ecdcf753ac2a7

Observation c69343e9-7ea8-4ff1-ba37-df93a101e6d9 · inbound

Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents cites this paper.

Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:56:38.062223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T16:55:47.922639Z digest=sha256:a89817685f485b17d120cbf383620ac99604fd5ebdcdd43e07fa6e9bb3cdba37

Observation 3a33c585-df94-4a14-a6c2-f96c21ae8e7d · inbound

Responsible Evaluation of AI for Mental Health cites this paper.

Responsible Evaluation of AI for Mental Health Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:50:54.931535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T12:50:33.070425Z digest=sha256:1399099f8473e341ba3c504bad02c8751922ff86eead163be5e05ffaad3c0851

Observation 729587ef-4445-41e6-a3ae-e208359ae64a · inbound

Making AI Evaluation Deployment Relevant Through Context Specification cites this paper.

Making AI Evaluation Deployment Relevant Through Context Specification Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:50:05.159501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T14:46:09.944168Z digest=sha256:f5b0623425a3e5912d5af0539795e8d022c1c5a0e012b11af8b17a8a5e1a5518

Observation 70749d48-bbd7-4928-bf76-ad85e797bbcd · inbound

Grounded Chess Reasoning in Language Models via Master Distillation cites this paper.

Grounded Chess Reasoning in Language Models via Master Distillation Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T21:26:43.149095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:26:43.149095Z digest=sha256:cf4f194e76c3884c8023985bfc009495c725318c761211e2df2379e8d243bb4c

Observation 0191584d-4385-4212-ae50-19f6e3560e0f · inbound

RLHF May Not Reflect Genuine Preferences cites this paper.

RLHF May Not Reflect Genuine Preferences Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:40:46.318423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T08:39:08.880486Z digest=sha256:c99630aab6f49b633d23bd2fa503eb3f250ace1e2f03da2346feb308a05197e4

Observation e1116046-03f9-40e2-9144-9a0d98caa63a · inbound

From Ground Truth to Measurement: A Statistical Framework for Human Labeling cites this paper.

From Ground Truth to Measurement: A Statistical Framework for Human Labeling Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:30:58.092181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:09:43.892161Z digest=sha256:48923722a20e027ac58c62fbd7c6385c79897dd152c34fb06658b4cf8e043f5b

Observation 88589ab5-5cf6-4cf8-a6b4-3a499773718a · inbound

"I Just Don't Want My Work Being Fed Into The AI Blender": Queer Artists on Refusing and Resisting Generative AI cites this paper.

"I Just Don't Want My Work Being Fed Into The AI Blender": Queer Artists on Refusing and Resisting Generative AI Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:25:22.668278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T12:21:20.533180Z digest=sha256:e5546dee9fe101ef8be420c6b96adadc6f56876073710b92ed3657f7330d6c4c

Observation fd24ef72-532d-402a-bfae-f5c653595ac9 · inbound

Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios cites this paper.

Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:55:53.066086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T02:54:52.984994Z digest=sha256:5d09921fa2fe2a0cd41562b2cdce724b22e64e7ff4d60b8d2d252ec6257debe5

Observation a1357ba7-1d60-473f-bc96-e9f837330683 · inbound

AI as a Tool for Simulation-Based Experiments in Literary Studies cites this paper.

AI as a Tool for Simulation-Based Experiments in Literary Studies Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T15:02:18.979706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T14:54:14.909995Z digest=sha256:d0dcbfac1d2e21a39448e2c9d85474419b0f952e47c660aed33f4bea7fc367ce

Observation 9562a4fc-d8f1-4cda-ba93-c52ece092f06 · inbound

Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer cites this paper.

Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:27:40.352337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T13:13:19.690367Z digest=sha256:efaa62b2effbab4f4dfac017d12bea5dd2a057cb82fba57198bae66d98478a4f

Observation 7948401f-e3a0-43b7-9be0-a9cd98d0453b · inbound

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering cites this paper.

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T22:08:58.949519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T23:48:58.497927Z digest=sha256:c01e2b150f8b6e6977355977838877842ded690664fe6ed36a4c6a8787fd3d36

Observation 2f29de4b-0b44-496a-8c61-ed0646ca3c24 · inbound

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering cites this paper.

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T11:06:03.560871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:06:03.560871Z digest=sha256:e981210a6895b9fb6b92d85fad34b2bf8d0cb6be1d3cd8f33c159805fbe51f22

Observation 13e8e9b3-298f-419d-a664-93ccfc27f6f8 · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:03.248723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:03.248723Z digest=sha256:a4bc0cb2942d472f7d5b15d311eb7c507aed2dd5f896aaeb4fda2b0f4ac1b2e3

Observation 077493da-6a1e-4c2a-a038-729646409e34 · inbound

On the Convergent Validity of Offline Evaluation Designs for Recommender Systems cites this paper.

On the Convergent Validity of Offline Evaluation Designs for Recommender Systems Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T01:29:53.636089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:29:53.636089Z digest=sha256:a440075507c9a18d04fdf7e30a668cb82e350a0ef36686adf53cae91b0d58eab

Observation eab46b33-0a67-4492-931c-1842b1c368b1 · inbound

Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions cites this paper.

Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-04T06:17:17.719091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:17:17.719091Z digest=sha256:b9cf2171e568702460f87be4d1a1a3b18ad9276462837fc667ab7d0bd783bc23

Observation 432b5be4-cf3e-45a9-a420-d93609dc875e · inbound

Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions cites this paper.

Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-07T00:11:20.211326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:11:20.211326Z digest=sha256:44fe4f20afef333662b79a32d86aab5bff01b2d64962388cc373152bc22311bb

Observation 5b18ed79-9a45-4921-9b95-d578a6f61c52 · inbound

Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation cites this paper.

Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T01:01:58.715300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T01:01:58.715300Z digest=sha256:b2270d5fe7331dcbeefd3ef5ac016e04ae0ffa992e36f5a32ea3e1f25d3fc0c0

Observation 90fdb2a7-d6a1-4bdf-a364-b147057f1c14 · inbound

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) cites this paper.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.707630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.707630Z digest=sha256:08ef71601887b505b922604c644cb10a3344c82496d0cb239a060c7b270407de

Observation 679ee3a9-e744-4422-9a33-857f86bfc0f3 · inbound

Algorithmic Gender Prediction Is Illegitimate, But Gender Imputation Can Yield Valid Measurements cites this paper.

Algorithmic Gender Prediction Is Illegitimate, But Gender Imputation Can Yield Valid Measurements Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-14T10:55:22.643993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:55:22.643993Z digest=sha256:b59aac8ab7fb5aa028c943ae338b1e879a8a058e7cc617af53be4c6124baae63