Pith. sign in

Paper Citation Record · LEDGER

A Framework for Evaluating LLMs Under Task Indeterminacy

As of 14 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2411.13760.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13760 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:58:47.648449Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved24
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c077c67-c093-4ebf-88d3-ecb29673ac03 · outbound

This paper cites DICES Dataset: Diversity in Conversational AI Evaluation for Safety.

A Framework for Evaluating LLMs Under Task Indeterminacy DICES Dataset: Diversity in Conversational AI Evaluation for Safety

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.839258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.410920Z digest=sha256:8acb8c39e99a3739d072195d4523d4b83c8b8f31b3d3b7532d21d1035f96ce5a

Observation a57fd76d-a59b-4874-9ec1-70b47786c546 · outbound

This paper cites Stop Measuring Calibration When Humans Disagree.

A Framework for Evaluating LLMs Under Task Indeterminacy Stop Measuring Calibration When Humans Disagree

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.416548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.416548Z digest=sha256:cfc01313cefa99b6d2955c99273208d249ae200157234c4e53af5d43fb3803ee

Observation 2b3da26a-e236-43ed-9e86-cac277deb601 · outbound

This paper cites It’s the End of the Gold Standard as we Know it.

A Framework for Evaluating LLMs Under Task Indeterminacy It’s the End of the Gold Standard as we Know it

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.821568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.422176Z digest=sha256:ebae164e7b6e91f146cbedcd63213385de03095e8702ced583f67ba8a7cb6d57

Observation 8d37a0ce-5096-4610-a17a-d9fba6cdd5c2 · outbound

This paper cites Like trainer, like bot? Inheritance of bias in algorithmic content moderation.

A Framework for Evaluating LLMs Under Task Indeterminacy Like trainer, like bot? Inheritance of bias in algorithmic content moderation

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T15:58:48.520699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.427557Z digest=sha256:fee6b90e6c7f4ad173c47c401b8036174a7a4a6b6e5c85eb5cf3428fa426e8f4

Observation 28c48695-e4b3-47dc-8a24-ac5d9e0ce8b6 · outbound

This paper cites Holistic evaluation of language models.

A Framework for Evaluating LLMs Under Task Indeterminacy Holistic evaluation of language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.802595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.433122Z digest=sha256:6421177beac31fc92e6ca47c9374a1bd6f8342a743939c0efecaf7ae6138278f

Observation 28671047-f62c-4e76-bf92-a379645b173b · outbound

This paper cites an unresolved cited work.

A Framework for Evaluating LLMs Under Task Indeterminacy Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.438380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.438380Z digest=sha256:f880bbf95459e6e82c2960cc360f710c111ff65ce65d6c2491b8efe69e7482bd

Observation 00d71071-c431-4e74-8027-1c47e9925021 · outbound

This paper cites Revolt: Collaborative Crowdsourcing for Labeling Machine Learning Datasets.

A Framework for Evaluating LLMs Under Task Indeterminacy Revolt: Collaborative Crowdsourcing for Labeling Machine Learning Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.447529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.447529Z digest=sha256:767cb99164d5be902d74bcd3744796040562b24b17a9d0ca8bace4340d17da0b

Observation c6af6ccc-9894-468b-b70a-b43f38dc4034 · outbound

This paper cites Judgment sieve: Reducing uncertainty in group judgments through interventions targeting ambiguity versus disagreement.

A Framework for Evaluating LLMs Under Task Indeterminacy Judgment sieve: Reducing uncertainty in group judgments through interventions targeting ambiguity versus disagreement

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.786308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.453452Z digest=sha256:30749362f901b7cab52ecc13f1620736f774c97a9ae6ee9da927ad67a66e54ad

Observation 9a458d73-d10a-4790-80a3-159d996761a7 · outbound

This paper cites Judgment Sieve: Reducing Uncertainty in Group Judgments through Interventions Targeting Ambiguity versus Disagreement.

A Framework for Evaluating LLMs Under Task Indeterminacy Judgment Sieve: Reducing Uncertainty in Group Judgments through Interventions Targeting Ambiguity versus Disagreement

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:58:48.356776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.459514Z digest=sha256:4ca9b368b29354a4b9c75e5b4d32e6f465f165874f905fc107db063bfd3c261d

Observation 9b8ed86a-925b-4376-959a-e23aace9ef7d · outbound

This paper cites Dealing with Disagree- ments: Looking Beyond the Majority V ote in Subjective Annotations.

A Framework for Evaluating LLMs Under Task Indeterminacy Dealing with Disagree- ments: Looking Beyond the Majority V ote in Subjective Annotations

Reference 10

Resolution
malformed identifier
doi_truncated, observed 2026-08-12T15:58:47.793619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.465135Z digest=sha256:05fd536057ea962a27674e48feea305e9e9a711d329bcd797418faa0e8f29edd

Observation ff854f86-582a-476c-b805-f332c7290525 · outbound

This paper cites D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation.

A Framework for Evaluating LLMs Under Task Indeterminacy D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.470355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.470355Z digest=sha256:d5f79f4e1c07efdd295080cd1aedf8de38640a3d8b337532af10e8d2cb89bba3

Observation 3ea58abb-d693-4056-aac1-58b585a5ad0a · outbound

This paper cites Red-Teaming for Generative AI: Silver Bullet or Security Theater?.

A Framework for Evaluating LLMs Under Task Indeterminacy Red-Teaming for Generative AI: Silver Bullet or Security Theater?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.475898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.475898Z digest=sha256:2c9d3cfd469c85692df9a7bc5d2a0eba07d09f6aa45939118ab18a298bc5eb48

Observation 8aed3702-9d7b-445e-aefd-781624d01b0f · outbound

This paper cites Efficient Conformal Prediction via Cascaded Inference with Expanded Admission.

A Framework for Evaluating LLMs Under Task Indeterminacy Efficient Conformal Prediction via Cascaded Inference with Expanded Admission

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.481439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.481439Z digest=sha256:62ae0fe125eef6120df695116a3e04a748838263570c09bc12f2996e89700381

Observation 0ea16206-678b-402d-8ee8-379a7cd8505f · outbound

This paper cites The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels.

A Framework for Evaluating LLMs Under Task Indeterminacy The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.486664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.486664Z digest=sha256:248bc5369c8845705d74141bc1aeb135813eb060b52d71162913a42d9a949221

Observation cb126527-a912-4f1a-a3e6-558a7fc321a7 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

A Framework for Evaluating LLMs Under Task Indeterminacy Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.491762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.491762Z digest=sha256:9ccd339ff7e545323b3a05777b9a3a6e2000b2ee4e934bea4c3d77be64f0fd00

Observation 68160aa2-5850-4122-90bc-0f1ae0396ae7 · outbound

This paper cites Deep Label Distribution Learning With Label Ambiguity.

A Framework for Evaluating LLMs Under Task Indeterminacy Deep Label Distribution Learning With Label Ambiguity

Reference 16

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T15:58:48.243981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.496224Z digest=sha256:4f0930de1bb68e46d87dfe3602050707afdbeb193a554f175e856bba1e49ffa3

Observation 4809e15c-34b0-4d78-b633-a8e1ed9dc364 · outbound

This paper cites Label Distribution Learning.

A Framework for Evaluating LLMs Under Task Indeterminacy Label Distribution Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.500370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.500370Z digest=sha256:dc9efebf77643d62f7cd3b6f4fd993905f9ab4d6e698dc87b3b93ed17b74461d

Observation 210bc53a-7606-4252-9fbc-38252d34451d · outbound

This paper cites The disagreement deconvolution: Bringing machine learning performance metrics in line with reality.

A Framework for Evaluating LLMs Under Task Indeterminacy The disagreement deconvolution: Bringing machine learning performance metrics in line with reality

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.769904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.504721Z digest=sha256:752195c43b9b714e74b5a173bd28a9564f1b910bae6ec718c83af6fc82508c5c

Observation a7959e7b-177f-4f5d-a252-6e4a117e8c02 · outbound

This paper cites Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S.

A Framework for Evaluating LLMs Under Task Indeterminacy Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.508867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.508867Z digest=sha256:d5c23bfa1f0412c7a5f633fae26de2bfd9fc13c03b4a71f34b910a7ddb05a41a

Observation a7d6d6bd-d8b1-45c2-9a75-a31e0516dfad · outbound

This paper cites Jury learning: Integrating dissenting voices into machine learning models.

A Framework for Evaluating LLMs Under Task Indeterminacy Jury learning: Integrating dissenting voices into machine learning models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.752362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.513299Z digest=sha256:361113515e2e9874e560e6dcfd2282f1d0e0342d0e3cca94ee4f975765ebe5ab

Observation 9cbcbbe0-069c-473b-9568-f8cd5de417a1 · outbound

This paper cites Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation.

A Framework for Evaluating LLMs Under Task Indeterminacy Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.734640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.517490Z digest=sha256:57c66848b225521ea642d5f38c93e453cb6d18c0cd6a24657ab1b9a94e1298b8

Observation 50d372cd-4d44-4a5d-aa16-ca30d7b7bb7b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

A Framework for Evaluating LLMs Under Task Indeterminacy Measuring Massive Multitask Language Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.522165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.522165Z digest=sha256:7ff3e457dde0c6d395a17c06710024525f1a088cd0539161f54ff35d638161e6

Observation 28b35d4a-d867-48bc-9e37-4ffe9a43f1d8 · outbound

This paper cites Intersectionality in ai safety: Using multilevel models to understand diverse perceptions of safety in conversational ai.

A Framework for Evaluating LLMs Under Task Indeterminacy Intersectionality in ai safety: Using multilevel models to understand diverse perceptions of safety in conversational ai

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.717943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.527335Z digest=sha256:c288216c3e7ae871e508b1cb093d538d1fc9e6b159ce63fda61301a6c74d5319

Observation 1bf6244c-5633-479d-990b-713b16189533 · outbound

This paper cites Culturally Aware Natural Language Inference.

A Framework for Evaluating LLMs Under Task Indeterminacy Culturally Aware Natural Language Inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.531722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.531722Z digest=sha256:4a850e3cae49d5ee038f0830930be97e3db1cabeddcd7e59856b1d2a2aba615a

Observation 77dd0d86-8374-4649-9493-e5226e7d39dd · outbound

This paper cites Chaithanya Manam, Dwarakanath Jampani, Mariam Zaim, Meng-Han Wu, and Alexander J.

A Framework for Evaluating LLMs Under Task Indeterminacy Chaithanya Manam, Dwarakanath Jampani, Mariam Zaim, Meng-Han Wu, and Alexander J

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.699453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.536697Z digest=sha256:6a796b10149e80c42a04059d1d5a7b8c595c597eb77a20b0e44e0f371536627b

Observation a874b2e7-e4c3-487c-88ed-ec1d04b97ef8 · outbound

This paper cites Annotation Error Detection: Analyzing the Past and Present for a More Coherent Future.

A Framework for Evaluating LLMs Under Task Indeterminacy Annotation Error Detection: Analyzing the Past and Present for a More Coherent Future

Reference 26

Resolution
malformed identifier
no resolver link, observed 2026-08-12T15:58:47.541649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.541649Z digest=sha256:51cdd19a685314064c90dea7ab949478b49b6994541ba2322d62696df71cf097

Observation d383fbc1-768b-4684-90d8-75a83dca527e · outbound

This paper cites A Bayesian Framework for Modeling Human Evaluations.

A Framework for Evaluating LLMs Under Task Indeterminacy A Bayesian Framework for Modeling Human Evaluations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.546651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.546651Z digest=sha256:af46c163fac9a8ea60ebe0c20e63d833f53c5fdd258695babcf5bd756f73bcb0

Observation d3ef8992-3d96-46b0-ad4e-d464487125d0 · outbound

This paper cites Reconsidering Annotator Disagreement about Racist Language: Noise or Signal?.

A Framework for Evaluating LLMs Under Task Indeterminacy Reconsidering Annotator Disagreement about Racist Language: Noise or Signal?

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.679755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.552442Z digest=sha256:f9e3b16186731fc173fd4d88f793a7a8e496c94814f7a995f4361cedac083c94

Observation c16a2fd9-fa2f-488f-aada-166b24f23f4d · outbound

This paper cites Learning to predict population-level label distributions.

A Framework for Evaluating LLMs Under Task Indeterminacy Learning to predict population-level label distributions

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.661299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.557193Z digest=sha256:5825591b5cd78f9aa0ad81204d7174635484259c9a72c30729bd8de3b163ee5a

Observation f5c68c08-3d94-4613-8b4f-72dae82cb0c8 · outbound

This paper cites A Safe Harbor for AI Evaluation and Red Teaming.

A Framework for Evaluating LLMs Under Task Indeterminacy A Safe Harbor for AI Evaluation and Red Teaming

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.561742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.561742Z digest=sha256:5c8308c49345b47e22fa3d1ecf36b0fc3ec1aadd14799ad2a5d3984c42b12efd

Observation b032cac8-7e8e-4f0f-ac4b-c349810d2c66 · outbound

This paper cites A Framework for Automated Measurement of Responsible AI Harms in Generative AI Applications.

A Framework for Evaluating LLMs Under Task Indeterminacy A Framework for Automated Measurement of Responsible AI Harms in Generative AI Applications

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.566997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.566997Z digest=sha256:42a1a5b35c09776a9ff799817042cfcb78fc9aea9f903d8d8b9685da2f18f31d

Observation e9a420e2-695e-46aa-97db-9bd5713d93c7 · outbound

This paper cites an unresolved cited work.

A Framework for Evaluating LLMs Under Task Indeterminacy Unresolved cited work

Reference 32

Resolution
verified exact
doi, observed 2026-08-12T15:58:47.716100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.572311Z digest=sha256:5552403d2c9b48a5ba195e41e353763190e75e366cd05e674574d9166146928b

Observation 630a4f9b-ddfe-4f13-bd99-180ec7bc36f4 · outbound

This paper cites StereoSet: Measuring stereotypical bias in pretrained language models.

A Framework for Evaluating LLMs Under Task Indeterminacy StereoSet: Measuring stereotypical bias in pretrained language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.577191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.577191Z digest=sha256:d4ba9e9dfca108696ba75351605029dd32e7e2934ad849eebc3c818eb86eb54a

Observation 6657a046-e9a1-4318-a733-e6ca4ee0f37e · outbound

This paper cites Diversity-aware annotation for conversational ai safety.

A Framework for Evaluating LLMs Under Task Indeterminacy Diversity-aware annotation for conversational ai safety

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.645028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.581750Z digest=sha256:d375835d75f51acd7bbea41610cd310a597e9ad113a8263e89d8dafd0efcd866

Observation a6e9d212-c5aa-486e-ac2b-6bb52305a3f3 · outbound

This paper cites Inherent disagreements in human textual inferences.

A Framework for Evaluating LLMs Under Task Indeterminacy Inherent disagreements in human textual inferences

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.627575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.586220Z digest=sha256:e3d21904d1e4b97dd6b019c45bb3d4f1d8ab778a91ad1192448dfb7bb1362d82

Observation 4540f88d-e957-4faf-8bf6-61f13e4b7019 · outbound

This paper cites Inherent Disagreements in Human Textual Inferences.

A Framework for Evaluating LLMs Under Task Indeterminacy Inherent Disagreements in Human Textual Inferences

Reference 36

Resolution
malformed identifier
no resolver link, observed 2026-08-12T15:58:47.591125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.591125Z digest=sha256:e40a9102191fc882b708b04f9c1fe87ac4e9974db6b67e82a520ef07619f8ab7

Observation ea6b370c-3186-416d-83f1-dc46594621fe · outbound

This paper cites Human uncertainty makes classification more robust.

A Framework for Evaluating LLMs Under Task Indeterminacy Human uncertainty makes classification more robust

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.610984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.595776Z digest=sha256:46b4bb4b9bb410202b6844fdf89232e3e8ce66ed75f02be332ca705d13bb6a2b

Observation 8c67e5e5-81c0-44dc-87a4-d46a433706d4 · outbound

This paper cites The 'Problem' of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation.

A Framework for Evaluating LLMs Under Task Indeterminacy The 'Problem' of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.600485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.600485Z digest=sha256:f2d46c2b4260d703419a1a4a4a085598235142ac1a1cad4b71b89f268d8a57f8

Observation 272987e6-d92a-44e6-a2a4-c316535e7c9a · outbound

This paper cites On Releasing Annotator-Level Labels and Information in Datasets.

A Framework for Evaluating LLMs Under Task Indeterminacy On Releasing Annotator-Level Labels and Information in Datasets

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.605264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.605264Z digest=sha256:453d55cba39de54e71e723e90e203dbb26ca1a719df3e9329f2a37d52fe2fa96

Observation f1e7394c-d3f0-43e8-b506-d192bc39d250 · outbound

This paper cites Principled Evaluation with Human Labels: One Rater at a Time and Rater Equivalence.

A Framework for Evaluating LLMs Under Task Indeterminacy Principled Evaluation with Human Labels: One Rater at a Time and Rater Equivalence

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T15:58:47.871195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.610350Z digest=sha256:da9057a30b26f4d38b46cc0bedaa98a144d040137ca5c09bb5037f60ae131fcf

Observation 602c5eaf-4359-4f96-8eaa-e3f5f29ae945 · outbound

This paper cites Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection.

A Framework for Evaluating LLMs Under Task Indeterminacy Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.615404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.615404Z digest=sha256:76d1c9d101632fa52c6291975336adbf220986cd8b82918b96d0da93217222e4

Observation 09088a7d-bbd7-4f5d-ac12-f8a906a710b3 · outbound

This paper cites Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences.

A Framework for Evaluating LLMs Under Task Indeterminacy Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.620627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.620627Z digest=sha256:8e3baff8561866ae45a88ec94c4e9feb1af990f00ba724d2fe08b5866c5ded42

Observation 83efe431-4f75-4d3f-a57a-9139d956405b · outbound

This paper cites A case for soft loss functions.

A Framework for Evaluating LLMs Under Task Indeterminacy A case for soft loss functions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.625408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.625408Z digest=sha256:e2e2abc4b4550a3c16c2a64b6c1e82ea3a0ea9387e377241e8f226a5ff1f11cc

Observation e5c52f95-88a5-46ce-9d23-698238ed1636 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

A Framework for Evaluating LLMs Under Task Indeterminacy GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.630144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.630144Z digest=sha256:6d728cfd5db6c5d7052014851645621850f2cd1b1c2c54516e4d870023a91540

Observation affa9a7a-e8f2-410a-a2c9-bcd63e415c00 · outbound

This paper cites gold data.

A Framework for Evaluating LLMs Under Task Indeterminacy gold data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.582824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.634597Z digest=sha256:5adaba8a58339837088997070690b41a660b97cd63848c958381ebb6c1abe85e

Observation 993a6c94-0d98-48a6-ac30-d1adeb6b00b7 · outbound

This paper cites Super- naturalinstructions:generalization via declarative instructions on 1600+ tasks.

A Framework for Evaluating LLMs Under Task Indeterminacy Super- naturalinstructions:generalization via declarative instructions on 1600+ tasks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.638882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.638882Z digest=sha256:edacd695e36b4295a46d4d73a5b09b742c8fa6de94ea2a22c9fde630c0e1b95a

Observation 86907fa7-da6a-4695-adfa-7f5c27498499 · outbound

This paper cites Disagreement Matters: Preserving Label Diversity by Jointly Modeling Item and Anno- tator Label Distributions with DisCo.

A Framework for Evaluating LLMs Under Task Indeterminacy Disagreement Matters: Preserving Label Diversity by Jointly Modeling Item and Anno- tator Label Distributions with DisCo

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:47.643852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:47.643852Z digest=sha256:34c281316bade8361f41c9ea9b05986b58f03bc3d289b0e5e3451bdcb8140d97

Observation 35c8acda-de24-46ca-a0e7-f120e657f2f7 · outbound

This paper cites Many islands, many problems: An empirical examination of online safety behaviors in the caribbean.

A Framework for Evaluating LLMs Under Task Indeterminacy Many islands, many problems: An empirical examination of online safety behaviors in the caribbean

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:48.555782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:58:47.648449Z digest=sha256:0116f0afdf6d95500284584b89a59bd93d8f9857abf870fbc9f650bca399f724

Pith citing papers

No inbound Pith citation observations are available.