Pith. sign in

Paper Citation Record · LEDGER

Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2103.14749.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2103.14749 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:23:34.771117Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:49:41.632097Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ae4bee83-4dc3-49df-9e32-52238d2e61be · inbound

Wake Vision: A Tailored Dataset and Benchmark Suite for TinyML Computer Vision Applications cites this paper.

Wake Vision: A Tailored Dataset and Benchmark Suite for TinyML Computer Vision Applications Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:08:41.948583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T01:06:48.298874Z digest=sha256:d6478a1f286e0ca57d1aba81d9b5122c9a66d2b521dafc603f733dcaae0d52e7

Observation cb113683-04c4-4100-9fa5-7dcd2199b93e · inbound

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI cites this paper.

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:44:01.690041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-17T00:44:01.658214Z digest=sha256:d6b3f428cc0ec42aa1b2c6f760c54d97f0393b8f3373e6faf8a3ef4fedf3f11e

Observation a4cca4e6-213c-4971-9fa4-8b4f1d74cda1 · inbound

Fundamental Challenges in Evaluating Text2SQL Solutions and Detecting Their Limitations cites this paper.

Fundamental Challenges in Evaluating Text2SQL Solutions and Detecting Their Limitations Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T00:23:34.771117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:23:34.771117Z digest=sha256:53408e1d9b2b5072e9809df5f90560c7268314f7bf21ca3f6bb588a603ec1c8c

Observation 05fdc43a-0a68-41be-aaad-864ab2683443 · inbound

Do Large Language Model Benchmarks Test Reliability? cites this paper.

Do Large Language Model Benchmarks Test Reliability? Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T04:45:34.602333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:45:34.602333Z digest=sha256:fce7f362221a73948b064807ad5590d475c6b458212f2008009c4ea024397400

Observation 906ea888-465e-4af4-a526-fd00a2f4b67d · inbound

Network Intrusion Datasets: A Survey, Limitations, and Recommendations cites this paper.

Network Intrusion Datasets: A Survey, Limitations, and Recommendations Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 158

Resolution
unresolved
no resolver link, observed 2026-08-08T14:44:41.854110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:44:41.854110Z digest=sha256:c0c25bfe43807ce34c526738c0da1480bad3b769fd9c248160051918954775a0

Observation 8b001dc8-5ec3-459b-acf6-207481d05d37 · inbound

TMLC-Net: Transferable Meta Label Correction for Noisy Label Learning cites this paper.

TMLC-Net: Transferable Meta Label Correction for Noisy Label Learning Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T11:50:06.274969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:50:06.274969Z digest=sha256:88c4722a2f7dcc0bb3491e16020340cf8b84338b08a20f3d10ba42d1c30fe211

Observation 62238fca-fcb3-4dbc-b782-a8949f7f8a78 · inbound

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models cites this paper.

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T20:55:05.895061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:55:05.895061Z digest=sha256:d3732f2e99aa8519223a35db63d4687169d18af231a71dca4145fb03bdbd7135

Observation 6564c559-b69d-4101-bf50-bc41bfde0b53 · inbound

The Achilles Heel of AI: Fundamentals of Risk-Aware Training Data for High-Consequence Models cites this paper.

The Achilles Heel of AI: Fundamentals of Risk-Aware Training Data for High-Consequence Models Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:34.923297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:34.923297Z digest=sha256:a31acb50f1b93509e6baac1a398141643cb23b891a8650e755304bb983c576cd

Observation 44cd219f-29c8-47d5-bc64-9c0ee8395bae · inbound

When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification cites this paper.

When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:10.178439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:10.178439Z digest=sha256:21d2ad7a21039adf38ae0e976161b01cbba4b80f8957a15500b9dd9d4c82d4ab

Observation c983470c-c349-46e1-9e9a-ff4b00faa454 · inbound

Machine Unlearning for Robust DNNs: Attribution-Guided Partitioning and Neuron Pruning in Noisy Environments cites this paper.

Machine Unlearning for Robust DNNs: Attribution-Guided Partitioning and Neuron Pruning in Noisy Environments Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:10:33.552063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:10:33.552063Z digest=sha256:1ca296b4d849af67b65afdc1b042de369a048d5a94b435808511d93658df53b7

Observation d581c8ac-9371-4066-90de-ec9b09c9c99f · inbound

GFLC: Graph-based Fairness-aware Label Correction for Fair Classification cites this paper.

GFLC: Graph-based Fairness-aware Label Correction for Fair Classification Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 9984

Resolution
unresolved
no resolver link, observed 2026-08-07T00:00:21.325181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:00:21.325181Z digest=sha256:d6cad0df2e8a09fe9068904827ee21ec91e6f007a7dc1f8371ce32e44ac0c884

Observation e61dbde6-c28c-4c7c-8d89-c116a8b5d06c · inbound

First-of-its-kind AI model for bioacoustic detection using a lightweight associative memory Hopfield neural network cites this paper.

First-of-its-kind AI model for bioacoustic detection using a lightweight associative memory Hopfield neural network Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:19.810351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:40:19.810351Z digest=sha256:b5a78cbd0dd28d128d0dcb83cb39662ee8b080407982c69eb5aec86bccf07e31

Observation 735d93c6-57d2-4076-b1c0-99856465a8ed · inbound

Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning cites this paper.

Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:02.589586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:45:02.589586Z digest=sha256:d021f8730a76b2d8349f011d40134c460e42f1c773dbfc6a11184fb6856a30db

Observation 4b3f88df-eb48-4c24-9b32-10475d3c1ecf · inbound

Advancing Mental Disorder Detection: A Comparative Evaluation of Transformer and LSTM Architectures on Social Media cites this paper.

Advancing Mental Disorder Detection: A Comparative Evaluation of Transformer and LSTM Architectures on Social Media Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:46.776213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:46.776213Z digest=sha256:1021a1e1cde7537601346f3d11a5f308ff74454bd0876ee184e986b858a338af

Observation a1c85b76-e1e6-4e7f-b84f-3a570739fec6 · inbound

Image Recognition with Vision and Language Embeddings of VLMs cites this paper.

Image Recognition with Vision and Language Embeddings of VLMs Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.475663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.475663Z digest=sha256:a3be624c4c4fcb5bf45184e299bb32cbed7f92f8b2212f462bc881008580bb25

Observation cdd606df-3baf-4f78-9f09-0af27a748c93 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:20:34.418769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T01:18:44.523602Z digest=sha256:5d32fd4754641140297c291e6eed1238e71e74ed8140bc1ecbec3479ddff3d57

Observation 538af51a-99cf-4371-af49-aea48be821c7 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-04T00:11:51.702024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:11:51.702024Z digest=sha256:ea2a058c8fd139fb98af252d0944852cf9d20129eea78c343a1d573aa0373297

Observation 94067a78-7cdd-491b-b535-394938cf65af · inbound

PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic cites this paper.

PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:44.099592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:44.099592Z digest=sha256:05ce9102c01bca08218fb31621a321ab82c4d6acef9ad5d4a1de2ff6654d7592

Observation 7db6f62e-1de8-4d6d-b77a-8e2f1417505b · inbound

Representation Unlearning: Forgetting through Information Compression cites this paper.

Representation Unlearning: Forgetting through Information Compression Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T06:59:18.572514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:59:18.572514Z digest=sha256:47cdb604c3a58c294f9bbb8fb1d4f2674c08e3d271ff43624c2ab8920909d804

Observation 120e7793-8c65-47a7-8bf5-281f9c6479f3 · inbound

Noise is not always detrimental: the capacity of quantum batteries is enhanced in black holes cites this paper.

Noise is not always detrimental: the capacity of quantum batteries is enhanced in black holes Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T09:29:29.046109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:29:29.046109Z digest=sha256:0fd584f89108664320a6fc2acb5d6ac721e86a83e17166d305d53154499277d0

Observation 333fe779-501d-49ee-81fd-ba6230820886 · inbound

Semantic Trimming and Auxiliary Multi-step Prediction for Generative Recommendation cites this paper.

Semantic Trimming and Auxiliary Multi-step Prediction for Generative Recommendation Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:50.716942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:48:25.382343Z digest=sha256:5657f4f87b998d7c43c5f12390f15fc33f03b748ff7ec350f2314ad59bdbec34

Observation 04d406b2-0b06-452c-a1cb-0acbbca4b45b · inbound

Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks cites this paper.

Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:41:36.875706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T06:37:19.313543Z digest=sha256:a91dba595aaa5f4b0e0a10e5801891a90643098ca52410628a6e575c91c4f698

Observation a803ab18-8460-41a3-a06e-4bb99044d7c9 · inbound

Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing cites this paper.

Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:53:23.358641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T22:52:17.310177Z digest=sha256:878731fc0d768417babd86fdc125ed5f6a651f7203cc7830625d8409abe6ddca

Observation d085801d-28cc-4707-abe8-e8ab6d8f6c1c · inbound

Signal-to-Noise Ratio and Sample Size Govern Representational Alignment in Neural Networks cites this paper.

Signal-to-Noise Ratio and Sample Size Govern Representational Alignment in Neural Networks Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T16:33:39.671679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T15:46:31.066600Z digest=sha256:ee3f942eec2c62984cecde27851a1d876d360536288ed80789f0632b486a3e99

Observation ce9465aa-659b-456a-a37c-3c98149b61e3 · inbound

Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets cites this paper.

Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:16.473620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:34:18.117267Z digest=sha256:79e67734d94af4b76cc2018972083667cd346d45eb0551b9bee1402f914196e3

Observation b4ed5274-eeda-4b73-b64c-57ed3b2d434b · inbound

Learning to Annotate Delayed and False AEB Events: A Practical System for Extreme Class Imbalance and Asymmetric Label Noise cites this paper.

Learning to Annotate Delayed and False AEB Events: A Practical System for Extreme Class Imbalance and Asymmetric Label Noise Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:09:19.667230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T20:36:13.042553Z digest=sha256:e1915efddfa78ca4e9cca75a29ce5a39184cf4cba36b5fd7a24d3578aa9af28c

Observation 3796f0b2-e934-442e-a549-4c610d763e0f · inbound

MMGist: A Comprehensive Multimodal Benchmark for 2027 cites this paper.

MMGist: A Comprehensive Multimodal Benchmark for 2027 Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:41.633870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T11:05:14.573386Z digest=sha256:2d4cccc711fa761c138ba6b4e1bc85c0a91ab3850fc26c0829e440a472535296

Observation c8e8a57a-9880-4be8-9c03-8b7aa0cb71ca · inbound

A Novel Method to Evaluate Models on Unreliable, Noisy and Inconsistent Labels: Adaptive Resolution Label Aggregation (ARLA) cites this paper.

A Novel Method to Evaluate Models on Unreliable, Noisy and Inconsistent Labels: Adaptive Resolution Label Aggregation (ARLA) Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T06:14:56.071808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:14:56.071808Z digest=sha256:0d3efaffb91252565b66ee4f1c9c96385e1d6f7a574997e65dfd882eb082afd7

Observation 5d3000d0-13a7-4bfb-8455-5a40d5cdd5d4 · inbound

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges cites this paper.

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T10:02:37.001079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:02:37.001079Z digest=sha256:70afde45e4fd6d4ba39c22b3145264f917b60dd5e03afcff0357d67a547b3f50

Observation f2a424e3-bfad-492c-af98-d2b9780de8c4 · inbound

Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings cites this paper.

Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-01T06:18:10.049087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:18:10.049087Z digest=sha256:8461d12f823704cc165438e0a3cbe42187f442956e969a1388e7fe2c6d0e3341