Pith. sign in

Paper Citation Record · LEDGER

Good practices for evaluation of machine learning systems

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2412.03700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03700 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:47:35.137643Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:57:32.628243Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 61db57a4-e31f-47e7-8b9d-9aa1840eab81 · inbound

A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition cites this paper.

A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition Good practices for evaluation of machine learning systems

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-15T00:03:31.986628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T00:03:31.986628Z digest=sha256:dd3f15b7009707a3bf7c58427d8e2d3e8f43ad3a923ae595043717744103b1a3

Observation 9b10e715-edaf-4f83-8432-d3c49d9e84c8 · inbound

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting cites this paper.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Good practices for evaluation of machine learning systems

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.629761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:c294607cdf6905ddc35f6be378ac478b0cf1ab7cc1be759929ec8fce3e3899e5

Observation f8c9d90c-89ff-480e-b37f-19f7c1095783 · inbound

Comparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study cites this paper.

Comparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study Good practices for evaluation of machine learning systems

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:34:18.660841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:33:01.811696Z digest=sha256:8c9f1a5df5a91302e0af22b0b4f3b44e1bc66f184786cafa100c494434024aca

Observation c96feeed-11a1-4b43-a57c-815754327e04 · inbound

L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education cites this paper.

L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education Good practices for evaluation of machine learning systems

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-13T06:19:53.291826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T06:19:53.291826Z digest=sha256:6caf45e27a1ade4ee057f89b174ec4478eed3b004014c8ebf362d1bd8a3830e8

Observation eba6c760-4656-4a98-86c6-5602bab565ee · inbound

L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education cites this paper.

L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education Good practices for evaluation of machine learning systems

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T07:47:35.137643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:47:35.137643Z digest=sha256:3c595ac35fa6f8a023ffae39a80a388ec68d9d798938d372b24c2ca9598f70a2

Observation 7434e100-2f37-452f-a909-1c95257413d8 · inbound

Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results cites this paper.

Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results Good practices for evaluation of machine learning systems

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T13:37:34.417441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:37:34.417441Z digest=sha256:dce086216f996b9b81b96cf8c85a9ad7864907efa94a5ec7f0a3b5440ee759ae

Observation 5c6b848a-f267-4043-9705-f4d3ecfe1a31 · inbound

From Read Speech to Spoken Digits: A Task-Specific Evaluation of Speech Privacy With Informed Attackers cites this paper.

From Read Speech to Spoken Digits: A Task-Specific Evaluation of Speech Privacy With Informed Attackers Good practices for evaluation of machine learning systems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T07:37:18.768072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:37:18.768072Z digest=sha256:2aed9a21d30ddbb62a5922f51511cc8c642600bc2f5c3636efac428333c2e2e2