Pith. sign in

Paper Citation Record · LEDGER

Beyond Accuracy: Behavioral Testing of NLP models with CheckList

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2005.04118.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2005.04118 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:54:50.853489Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T02:49:24.847192Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5caf5a5a-1988-475f-8864-1b178e1e3610 · inbound

Jailbreaking Black Box Large Language Models in Twenty Queries cites this paper.

Jailbreaking Black Box Large Language Models in Twenty Queries Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:48:33.255295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T09:48:31.721745Z digest=sha256:9a3da10e6e6ed3972068b3e9fede60e0bd37eb5a988a7a892bbd4d2d89dc6f6c

Observation c6fbf12b-d000-483f-9986-45df04151659 · inbound

Measuring Diversity in Synthetic Datasets cites this paper.

Measuring Diversity in Synthetic Datasets Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:50.853489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:54:50.853489Z digest=sha256:1cdcd39596051c215645b195cd3a83e721bb1efefb1350d712baf411314723cd

Observation f1b3c440-7823-4e4c-b899-9957d5f700d2 · inbound

Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models cites this paper.

Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:10.675956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:14:10.675956Z digest=sha256:60dd80f6213e8fe12c468ffe27c6900604ae5e251ddd4c0070eeb79ffaabb9e7

Observation 44185dd5-4404-45c2-b8e8-5e3dcf2f7052 · inbound

GenFair: Systematic Test Generation for Fairness Fault Detection in Large Language Models cites this paper.

GenFair: Systematic Test Generation for Fairness Fault Detection in Large Language Models Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:42.214448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:42.214448Z digest=sha256:306f81cdd8669bd6d5e55d884f568689a61546fbeb88069150905f21fd81a4f4

Observation 1bbb4541-0b04-4f46-83a6-dcf80abbd733 · inbound

Potemkin Understanding in Large Language Models cites this paper.

Potemkin Understanding in Large Language Models Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.993054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.993054Z digest=sha256:e868ad56c30c722ca299c2694d2429fe7228d92c180521b7078b83f0fc8e3b56

Observation ff9351b1-e60b-4ae7-a2cd-11fd9d455bf6 · inbound

ASSURE: Metamorphic Testing for AI-powered Browser Extensions cites this paper.

ASSURE: Metamorphic Testing for AI-powered Browser Extensions Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:43:29.244466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:43:29.244466Z digest=sha256:01725857efe982344a1d7ad047d3f8d26f4d7f35cdd9ff999965936a8b457617

Observation d9db4d8b-1f4d-4a14-a274-8c4700c34de4 · inbound

Agentic Web: Weaving the Next Web with AI Agents cites this paper.

Agentic Web: Weaving the Next Web with AI Agents Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 186

Resolution
unresolved
no resolver link, observed 2026-08-06T13:05:39.875211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:05:39.875211Z digest=sha256:2307f20ebeab67c3f8c73282bf2be26555dbcc69386c3085774e5df915d6a8c9

Observation ef847eb1-bc82-4595-8096-466b8c60d63f · inbound

Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation cites this paper.

Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T22:50:43.780478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:50:43.780478Z digest=sha256:34364e56e78e92ca1bcb7373a1e4a919964028e4c941f3b194bf7c6efeeda7c8

Observation 997716ef-ddec-459c-8def-2bc06e2f7fb0 · inbound

SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation cites this paper.

SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T22:46:44.391762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:46:44.391762Z digest=sha256:1cdcb78a60f57432250279d89b6f88e5d13441c716d04f6bf338958aff81f188

Observation f4c8d66b-1d0a-43ca-ab53-e87408ad5479 · inbound

Evalet: Evaluating Large Language Models through Functional Fragmentation cites this paper.

Evalet: Evaluating Large Language Models through Functional Fragmentation Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:01:39.982194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T16:57:25.259866Z digest=sha256:e72801774566f656c0affcb8e1046fefd3119b4657e1e8e60e92bd7e195d937c

Observation 8c9f4954-7740-491d-b601-f9c219ce2473 · inbound

When to Trust the Answer: Question-Aligned Semantic Nearest Neighbor Entropy for Safer Surgical VQA cites this paper.

When to Trust the Answer: Question-Aligned Semantic Nearest Neighbor Entropy for Safer Surgical VQA Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:22:15.719282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T01:21:26.854636Z digest=sha256:6de41b09a4d1d9bc463131745da89f73b725be2eb9cfd02e6f5f00c38b728c6a

Observation 26e056f2-14f0-433a-a8d8-32a364ff2adf · inbound

Learning to Discover at Test Time cites this paper.

Learning to Discover at Test Time Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:16:04.194337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:16:04.001700Z digest=sha256:4c152e0f12cb736fa26a951f823dd8e98dd20fc4b4a76b9e0b680e59aa0a6e35

Observation 97b09954-a9ce-4b3f-864e-8c7ddecbcdf8 · inbound

Measuring Representation Robustness in Large Language Models for Geometry cites this paper.

Measuring Representation Robustness in Large Language Models for Geometry Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:38:10.520784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T19:35:32.531660Z digest=sha256:2baae3b8040406493b414a755b20ca26cf4c194930d41b87f54de61b836e782d

Observation 38f3e839-bca3-4ff1-a7b5-be4ac92820fc · inbound

Social Bias in LLM-Generated Code: Benchmark and Mitigation cites this paper.

Social Bias in LLM-Generated Code: Benchmark and Mitigation Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 153

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:36:08.937852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T19:34:51.433422Z digest=sha256:0cf451d091615c0b7b5195f200772cae79b3488c3fa54749b18243c97d760ce1

Observation 633d3a0d-5872-477f-a314-cacd2f379e40 · inbound

How to Interpret Agent Behavior cites this paper.

How to Interpret Agent Behavior Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:27:35.960649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T18:23:25.269217Z digest=sha256:e6f9281ae4d6779c041745bf5432221cfd1fe22f70f1a40d4b9a5dc027ebc1a8

Observation 93a4eb04-fcab-4d4a-8a79-1ffdac38e016 · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:14.003198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:5fcdc5fb8cf01e204171780952ad614f822f6b452fb6525a28d33a2a2d64f6ab

Observation f9bb0636-5f55-41d3-b279-05e048ec2fce · inbound

Trustworthy Recommendation in the Era of Large Language Models: Opportunities and Challenges cites this paper.

Trustworthy Recommendation in the Era of Large Language Models: Opportunities and Challenges Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 195

Resolution
verified exact
arxiv_id, observed 2026-06-28T18:42:29.189519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T18:41:06.636352Z digest=sha256:bec0202d717b3b1045c4717ee5da67830b3a612753e2d7d71f95cab06f94d6ea

Observation c6b290ae-5898-429c-baf2-7260f431fd4e · inbound

CRAFT: Cost-aware Refinement And Front-aware Tuning of Prompts cites this paper.

CRAFT: Cost-aware Refinement And Front-aware Tuning of Prompts Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:46:46.394764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T06:39:17.268337Z digest=sha256:96d361a1801f2865b3b882d4cb3c9a2e31fada07258ef7adf5e71702cdf95eeb

Observation 1d509cc0-ac8e-4408-8151-33cc6236c888 · inbound

What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks cites this paper.

What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:47:31.706084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T16:18:01.850874Z digest=sha256:38648e01a721449a53945ab3599594d40158b193d4da789f78a3f098c1cda4c6

Observation 1238b618-2488-4f13-bd1e-02eb26ae55dd · inbound

Iterating Toward Better Search: A Two-Agent Simulation Framework for Evaluating Agentic Search Architectures in E-Commerce cites this paper.

Iterating Toward Better Search: A Two-Agent Simulation Framework for Evaluating Agentic Search Architectures in E-Commerce Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:18:23.057855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T07:05:31.737761Z digest=sha256:4293c3514e1546ba98fa41e34963fd6ed4f433a871eb06928cf516153736a847

Observation 142e94c2-a96f-4559-95d7-a7ec100e14a1 · inbound

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning cites this paper.

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.849283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T19:04:45.062426Z digest=sha256:ebe85ecabf6a6ee20a3032748592a82359c0c62d61d58203ce016eb733113e40

Observation 2e78ebc0-63ab-4619-8c6f-7b5caf2a74da · inbound

Carolina Guide: A Multi-Agent RAG System with Institutional Guardrails for Academic Policy Assistance cites this paper.

Carolina Guide: A Multi-Agent RAG System with Institutional Guardrails for Academic Policy Assistance Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:14:37.507549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T11:13:23.740219Z digest=sha256:01d72a2663596a15b4381841d7ec3e7e997239a95abefec0f1fcbf165cde42be

Observation 7c711e76-babd-4a98-80a1-2ccccfe33175 · inbound

Testing Retrieval-Augmented Generation Systems with Chunk Coverage cites this paper.

Testing Retrieval-Augmented Generation Systems with Chunk Coverage Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T15:54:00.365263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:54:00.365263Z digest=sha256:e3196fd45bdc9cc39bf1de3ec90291874d9e1cd4de7fb5a3c4120628d31fae95

Observation dd85f4fe-5239-4949-89ba-8c79e8ae304e · inbound

Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development cites this paper.

Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T07:40:17.900025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:40:17.900025Z digest=sha256:a59ec6f439851681679a3ccf6ecc4b986641615a8ab1f991c5ddee0d8da26eaf

Observation 48282414-8bb1-4cee-aeb7-6aafdb4a4c1f · inbound

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy cites this paper.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.525133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.525133Z digest=sha256:90ef6b29193e77692a1791574eae9bf0c03b4b1b24d8152f77c858bf12d6fa7a