Pith. sign in

Paper Citation Record · LEDGER

Dynabench: Rethinking Benchmarking in NLP

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2104.14337.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.14337 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:21:52.595165Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 39a34301-2c01-4c36-8d7d-52060d5d47d8 · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam Dynabench: Rethinking Benchmarking in NLP

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.375124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:df80518d2afbb7600a4d8e0e3efdaae63a8224ec6b42b310e2f8b8ef41149fa9

Observation 3b929376-432b-4282-9211-7a2ad3be4b8f · inbound

Thinking beyond the anthropomorphic paradigm benefits LLM research cites this paper.

Thinking beyond the anthropomorphic paradigm benefits LLM research Dynabench: Rethinking Benchmarking in NLP

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T22:21:52.595165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:21:52.595165Z digest=sha256:7f1ad63508da0077f5898694495568aa64754036ef67ab0dd2ada18c1d4d9b68

Observation dc4387f9-e218-4af0-a674-4f10df6c376e · inbound

LLM Performance for Code Generation on Noisy Tasks cites this paper.

LLM Performance for Code Generation on Noisy Tasks Dynabench: Rethinking Benchmarking in NLP

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:41.067226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:41.067226Z digest=sha256:5eabd2e7b35c22c38a0425b7b6243fafd01435d06c30d9fec3806af25204ddd1

Observation 6c072be9-1d61-4f5e-97e5-5ac8b700aab8 · inbound

Potemkin Understanding in Large Language Models cites this paper.

Potemkin Understanding in Large Language Models Dynabench: Rethinking Benchmarking in NLP

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:30.708684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:30.708684Z digest=sha256:880c835307e66321b76a55a39bdc43f12c2e9024ed3cbb3c33299d3a8cc50e75

Observation 011ad500-ec58-4f4c-8ec8-45ea96369d38 · inbound

Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models cites this paper.

Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models Dynabench: Rethinking Benchmarking in NLP

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:15.386951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:15.386951Z digest=sha256:27d85da442524d54637efdc3a084ef54138d678f6b5590b5e67775a59c57e810

Observation 2282e2fc-22ea-45e8-86ae-724b13ad9049 · inbound

Agentic Web: Weaving the Next Web with AI Agents cites this paper.

Agentic Web: Weaving the Next Web with AI Agents Dynabench: Rethinking Benchmarking in NLP

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-06T13:05:37.910504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:05:37.910504Z digest=sha256:6ed11d1a664289187f487bac5f6a656499fa496d20de2626bba0bd8a7076f89c

Observation 6cf568cf-ac71-422f-8d4f-e25c6af4ebed · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Dynabench: Rethinking Benchmarking in NLP

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.364886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:ffb5174fc4522ef186f86b5f1339a52a6bb366f62acd5376db906ae5e295a762

Observation b6ad4775-1b20-42e1-9058-1c72ab173879 · inbound

Private, Verifiable, and Auditable AI Systems cites this paper.

Private, Verifiable, and Auditable AI Systems Dynabench: Rethinking Benchmarking in NLP

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-05T15:43:59.083960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:43:59.083960Z digest=sha256:fd6c70f9072f8512e55c09d88b7c4189243df43e0e4ca6750f4ffa2bac45092c

Observation 91451675-098c-4626-8d54-ab661566828f · inbound

Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M cites this paper.

Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M Dynabench: Rethinking Benchmarking in NLP

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T19:48:22.922729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:48:22.922729Z digest=sha256:10427547e92f2cd630f6665a9920f80e3f7d23881d84d6cdd97401f688dfb276

Observation 0c5adf30-311c-4f75-b382-796f0467f70f · inbound

Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation cites this paper.

Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation Dynabench: Rethinking Benchmarking in NLP

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:12:29.823956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T08:12:02.449352Z digest=sha256:2d5df60d19aca47551b566a3ed1d2d50eed0ee379add587790db2e09597d849b

Observation 12d933b8-0880-4d12-86f6-c66322eed828 · inbound

EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning cites this paper.

EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning Dynabench: Rethinking Benchmarking in NLP

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T12:21:38.373016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:21:38.373016Z digest=sha256:30af164bcb62eaf79fe3b399e71c874e19940769b7a57a37fdebdb4d2e37e01b

Observation 7fa4cbc4-7498-4013-9463-71b8fa7e0024 · inbound

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI cites this paper.

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI Dynabench: Rethinking Benchmarking in NLP

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:23.268582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T10:19:56.003219Z digest=sha256:aeb736a2cf932f6c0daf129872581d82fa54ca74ea474080c902d5d6fb20bb22

Observation 142def31-9698-490a-a6c7-a17862a62603 · inbound

RoboPlayground: Democratizing Robotic Evaluation through Structured Physical Domains cites this paper.

RoboPlayground: Democratizing Robotic Evaluation through Structured Physical Domains Dynabench: Rethinking Benchmarking in NLP

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:52.133066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:46:08.897540Z digest=sha256:dc60ef6f4e18b55a51346cb770a50d2c01580dfd84b32e2b6340196d6a987eb3

Observation 04341ae4-8797-4e6b-9c8b-15d01ebd9d6c · inbound

Too long; didn't solve cites this paper.

Too long; didn't solve Dynabench: Rethinking Benchmarking in NLP

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:42.876444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:29:25.830962Z digest=sha256:c114c58a852d131fd54421ad7f5f7dcb0ee462013ed44942d32b9f9ebed03f42

Observation 1382ce5a-86b0-4d76-9a13-c56ed3d8f611 · inbound

Too long; didn't solve cites this paper.

Too long; didn't solve Dynabench: Rethinking Benchmarking in NLP

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T08:26:49.097626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:26:49.097626Z digest=sha256:037e58dc85dcde8f1d53d9eca06eee810c5f75911de363fe1dc627e31d87c650

Observation 00e9179d-a40f-4902-a1e7-dfa5a8904107 · inbound

CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction cites this paper.

CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction Dynabench: Rethinking Benchmarking in NLP

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.232510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:02:50.603020Z digest=sha256:dd1101c6d098c22d08ab52dc3415e594ab146d71dc0fcde906c43bddbafecc0d

Observation 8cbb151b-15f8-4fd4-999b-69396e71f8ff · inbound

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks cites this paper.

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks Dynabench: Rethinking Benchmarking in NLP

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:29.611935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T04:27:11.735657Z digest=sha256:05f6bfe7a6756b255810de3d212c959e32c404ab412e077bff55ca6609e1a3c0

Observation b8ad867c-3b0a-4c9f-b96a-8b3ae646d7ba · inbound

TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation cites this paper.

TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation Dynabench: Rethinking Benchmarking in NLP

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-09T20:17:04.499790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T20:13:06.376037Z digest=sha256:31d442fd8bde0f63b1aca96ea41d2d554fabc48b9e90f85f9ce2ea0944289011

Observation 8c63ef9a-691e-4399-b0ed-37d88cdde605 · inbound

Analysis and Explainability of LLMs Via Evolutionary Methods cites this paper.

Analysis and Explainability of LLMs Via Evolutionary Methods Dynabench: Rethinking Benchmarking in NLP

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:11:05.214722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T20:37:40.932811Z digest=sha256:0be1ddbb4120d8ae7a3e4db7a3cc53632b43bccef0c28333752164407713c266

Observation 10663faa-a7ed-4d7c-b0e8-5359b4e29507 · inbound

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games cites this paper.

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games Dynabench: Rethinking Benchmarking in NLP

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:17.250381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T17:06:32.814188Z digest=sha256:1cd836d6e08831565c2d5fdec6fe00c94196c104f35a4824c713d4a88ef20ec1

Observation 3f039523-3804-412f-a2bc-39b04018becd · inbound

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks cites this paper.

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks Dynabench: Rethinking Benchmarking in NLP

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:23.769739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T05:28:45.453455Z digest=sha256:d375b58b8f5e0d6f1a492ee61afc96cfe3a0c717e1983a87838a53d516209fe5

Observation f0e8d7ea-342f-4d28-bfcd-1ae36ec07607 · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science Dynabench: Rethinking Benchmarking in NLP

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:14.040649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:303df100e82dd2943514ba3240f265a76822d6f344061df45a313acab2a0ddb8

Observation e6026c7b-c1ff-4fa7-b33c-fd7ec7ec5ce6 · inbound

Open-World Evaluations for Measuring Frontier AI Capabilities cites this paper.

Open-World Evaluations for Measuring Frontier AI Capabilities Dynabench: Rethinking Benchmarking in NLP

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:39:43.734758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T06:38:51.427985Z digest=sha256:012a77fe0c9ba18c024fd5e870fb41f01315a8f05ddb04b0c56a1207c3f6a3dc

Observation 9137d3af-0e49-425f-865c-49b6bf62c79c · inbound

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks cites this paper.

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks Dynabench: Rethinking Benchmarking in NLP

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.468788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T10:46:24.554332Z digest=sha256:ee9424d30eaa03211dd6e80208eb0653c3ce982905bfcec3745b2f212ab642c8

Observation 3610d9eb-e494-4f55-87ed-ae935d5d4b59 · inbound

Meta-Benchmarks for Financial-Services LLM Evaluation cites this paper.

Meta-Benchmarks for Financial-Services LLM Evaluation Dynabench: Rethinking Benchmarking in NLP

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:18:22.686773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-03T14:08:20.432931Z digest=sha256:88bc798457c788cb4c7400961a8cf68f901fc6dd99e3858416d1b1862d4f1112

Observation d18b0e1a-4d83-4284-a968-2162598b044d · inbound

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation cites this paper.

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation Dynabench: Rethinking Benchmarking in NLP

Reference 306

Resolution
unresolved
no resolver link, observed 2026-08-03T00:24:07.971141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:24:07.971141Z digest=sha256:4540768c78c8d779c181e19f0149aebd1e0e482862d35d1cc82ac03cba68f575

Observation b805c310-4960-4f48-ab76-5432e72d93d6 · inbound

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks cites this paper.

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks Dynabench: Rethinking Benchmarking in NLP

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:33.194776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:45:33.194776Z digest=sha256:8807fa1b736f0b988fd2af4f4b3e00d5b878666e869967dbf75d0ea4720206fe