Pith. sign in

Paper Citation Record · LEDGER

Position: AI Evaluation Should Learn from How We Test Humans

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2306.10512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.10512 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:57:02.199615Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:00:41.508241Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8707592e-c62f-4339-83ba-6007f190d26f · inbound

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models cites this paper.

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models Position: AI Evaluation Should Learn from How We Test Humans

Reference 250

Resolution
unresolved
no resolver link, observed 2026-08-11T22:57:02.199615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:57:02.199615Z digest=sha256:4383a0a8d78bd3c61e7291ac57ad561d0f799007f2aa2a00647f3b00b48415b2

Observation 1e07450a-5866-43fd-8161-ac1d8c509204 · inbound

The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories cites this paper.

The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories Position: AI Evaluation Should Learn from How We Test Humans

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-10T17:01:28.302975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:01:28.302975Z digest=sha256:0c3dea9f551fabeb4222065bd5da2f900e540c946f8e7a5ed63ca5c5e776e0c7

Observation 912f4b53-cf7d-48e8-92eb-8b8aad5c6f93 · inbound

Examining Alignment of Large Language Models through Representative Heuristics: The Case of Political Stereotypes cites this paper.

Examining Alignment of Large Language Models through Representative Heuristics: The Case of Political Stereotypes Position: AI Evaluation Should Learn from How We Test Humans

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:19.226127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:18:19.226127Z digest=sha256:30ecf4aed15c4ba3d7f4f7accb29c840732cd5d9ae5a5e50cdc7909ef1390463

Observation 68778ea7-cdff-46c2-b429-b893d4500153 · inbound

Psychometric-Based Evaluation for Theorem Proving with Large Language Models cites this paper.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Position: AI Evaluation Should Learn from How We Test Humans

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.647330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.647330Z digest=sha256:b1576294fa31080fbbaced1ceabc7bac796fc7fcdf432986bc24d2dc47e2b296

Observation 05ed1da9-f499-4840-9920-0ff28eff1f3c · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Position: AI Evaluation Should Learn from How We Test Humans

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:25.866299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:25.866299Z digest=sha256:5e57303e67bf367fdb70fade8fe3e81c521acc120cf60136636efa3f4ab1f053

Observation 3fc827b1-b721-4472-8bcc-617ffe7d9033 · inbound

Fluid Language Model Benchmarking cites this paper.

Fluid Language Model Benchmarking Position: AI Evaluation Should Learn from How We Test Humans

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T17:10:43.454204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:10:43.454204Z digest=sha256:a72d51a96134f115a1fbb34f78e222bebaaf04e46797868e88c24c5cfe84d392

Observation 9ce79528-af9e-4bc9-a45c-9b7b9f3488d5 · inbound

Position: AI Evaluations Should be Grounded on a Theory of Capability cites this paper.

Position: AI Evaluations Should be Grounded on a Theory of Capability Position: AI Evaluation Should Learn from How We Test Humans

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:00:41.510890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-21T21:57:55.834632Z digest=sha256:5b75bc88b19e77d1f16df9c4382c778ea85c6a2409daf795d27c055dd632d546

Observation 3cbdb515-ecc6-42aa-84ee-abe526a966b9 · inbound

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks cites this paper.

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks Position: AI Evaluation Should Learn from How We Test Humans

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T08:10:10.570262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:10:10.570262Z digest=sha256:93234b9cfd8272087f1808420e3bfdc4dfe66249efbfea54c89f6239c4ac5197

Observation b116e399-813e-4881-a96a-e464b8e0b6a1 · inbound

FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition cites this paper.

FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition Position: AI Evaluation Should Learn from How We Test Humans

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:03.899427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T03:32:02.079410Z digest=sha256:187295775cae2c4d848e521a6f75470995bab31ec7820f9d7a4a386724a7a24b

Observation 5504ef60-7b19-44d4-9713-039aaca1b0b4 · inbound

Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation cites this paper.

Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation Position: AI Evaluation Should Learn from How We Test Humans

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:27.320015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T08:02:49.168872Z digest=sha256:14890a084fae84af629bec4e96cebdfb13f29b58d26cd46766cba4c6770fbc9c