Pith. sign in

Paper Citation Record · LEDGER

Human-Centric Evaluation for Foundation Models

As of 16 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 3 inbound Pith citation observations for arXiv:2506.01793.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01793 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:36:10.926833Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:59:53.097164Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:25:41.538596Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0925fa6d-86b5-4107-91d1-39b216dd577b · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Human-Centric Evaluation for Foundation Models ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.031077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.031077Z digest=sha256:b6e2baf6728154a4465c39f971b3903e1dc5296f04c86c2ec034e57cd23b2edc

Observation 6a0d0b08-ad7f-49d6-b64d-7f052af9b507 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.102859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.102859Z digest=sha256:8ab6d6ecf565c1ecb7e8933da2edc82936ba668b89e91ee038f923b0c600d0a9

Observation 10f8423d-b43e-4101-9f64-c10296dc7f76 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.224321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.224321Z digest=sha256:df259eece7fb198c01ea7c7b8b67142f6d5e00fc13460ca359f43afd22697677

Observation 7cb0088f-27c9-4efb-89af-fd601b966365 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.474041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.474041Z digest=sha256:ad5b8b20c459af41ffed90b0b90425e538fa4ea5e879081870c5c9561ce60e56

Observation cb08326d-a2b4-4860-bddf-5e5d0d273b09 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Human-Centric Evaluation for Foundation Models Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.593518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.593518Z digest=sha256:f733d60cdf00c641b5666a20335a10bcecbe79d2400ab6234726cc964fa3a869

Observation c9ef97ce-3767-40b3-b057-acc3948f46e0 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.706899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.706899Z digest=sha256:59d2bb00e681f20d9ea0780aa701de8696951a50c2af193c631f33e7a992ad9e

Observation 42158aae-ee92-4301-92d9-8ff88a539e89 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Human-Centric Evaluation for Foundation Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.792029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.792029Z digest=sha256:68e4b0700077a1f1e6641c0f8cada150a92caa468ba8da0fd7c26e8f0bf968d4

Observation 8ee8cdb2-dfa9-42af-9b33-144ce5b9461d · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:36:12.166425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:36:08.874604Z digest=sha256:02d322a9470d9f0727fb6a47830f45ddb28bbff1f9d06cd0863062b28410bb67

Observation a17b21af-4cf2-4fc7-a023-0c07623c5ad2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Human-Centric Evaluation for Foundation Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.961655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.961655Z digest=sha256:de36cb06678202e6e21a954bdece64f62d2861bc70ac794d1f9206a2aaaa8e58

Observation 8b0f499c-efa2-4a89-b27c-39659bdb72cf · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

Human-Centric Evaluation for Foundation Models Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.124209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.124209Z digest=sha256:7c27f43751090ed3f68f62fe2f3dbfe5e117ec4e13e8edb6860f47d81941428a

Observation d97275f3-aeb7-4259-beb4-4f5b339d4c49 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Human-Centric Evaluation for Foundation Models Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.278576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.278576Z digest=sha256:aea2e8fa94d486043a219768b987037ae9722ef209e8713551efde17533feaea

Observation 5b1ff2b2-0971-427a-8e64-d414fca976d8 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:36:12.015503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:36:09.435590Z digest=sha256:22a079f57e61f64fc2830224e1132ae77e783e3fa26cfa6968fbf03135bf0e3f

Observation 25ade072-4507-4075-b16b-4c97191acf3a · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Human-Centric Evaluation for Foundation Models SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.544344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.544344Z digest=sha256:1001b0ac19f8aa084c192c9ba8d33ee343bd7be2026bc1b505ae5a3a17134799

Observation 151c695c-ddd2-49fe-9a92-79b3978cd901 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.708758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.708758Z digest=sha256:f5435b8297efcd23de3206e87dfade8fc611d164a47b876ab0957828be380026

Observation d031f4f1-005b-4baf-8be1-021ff3e3a4fa · outbound

This paper cites Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena.

Human-Centric Evaluation for Foundation Models Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.030772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.030772Z digest=sha256:fd3bb9f156538f8abfe9486e6c57d44acbb300b15b0473c4f6305fed9dfb3c4f

Observation 919e515e-177e-4bd9-aeb6-75efb7ded9e5 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:36:11.856502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:36:10.154302Z digest=sha256:dd4c6e549d6e5d1133723edc34266db1528269bc516798ec8ddcace3bba03920

Observation 260f860c-a932-449e-8e9a-da62a1004ff0 · outbound

This paper cites Humanity's Last Exam.

Human-Centric Evaluation for Foundation Models Humanity's Last Exam

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.296684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.296684Z digest=sha256:bf89dac229210b5498c14ed8b594026d5d4a8249b2330aadfde8ccc312c49581

Observation b1bfe8b1-1829-4482-91fa-ee62fedbcd31 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.369987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.369987Z digest=sha256:faeb6e26070bf8e8dc7cf0aa64071920c40d2da8e31ba706b810e9457d2814ed

Observation 0c8d298b-56b9-4d13-bb50-b23deab3d747 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.453562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.453562Z digest=sha256:c8e5754a794bcc23a4af1030299c313b43f84bc120f18157a2e25daefc124589

Observation 12fd8df5-1cc1-44d1-8fb9-895e0f64c910 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Human-Centric Evaluation for Foundation Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.624298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.624298Z digest=sha256:e380242a01cbc99187f6a9b502421fb8e17e131ac73ddc6cd8f9fa7fb0bf3282

Observation 34896e7b-0bd2-4806-bb20-04bc90db23c2 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:36:11.406036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:36:10.706120Z digest=sha256:88e38845eaabd719c80456423d53ff6f312a2e4e21c48a8ebbdf186922e47520

Observation e1633281-b1ea-4bbf-b403-de3c5dbe2028 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Human-Centric Evaluation for Foundation Models LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.812719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.812719Z digest=sha256:3c4e4832af3bc286f9fbf2089bb811513ed8908a0926cdcd9b43f45aa37d3604

Observation 6efeae68-bae3-4bbb-98a0-2fe772516df6 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:36:11.186918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:36:10.926833Z digest=sha256:d6e5bfd15191f239fb07a0dfab1ef622598fb34715d1fc2ca17600e4ae2b4881

Observation dd24c201-aef2-4daa-9cec-1c0213776a6f · outbound

This paper cites FinQA: A Dataset of Numerical Reasoning over Financial Data.

Human-Centric Evaluation for Foundation Models FinQA: A Dataset of Numerical Reasoning over Financial Data

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.350199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.350199Z digest=sha256:a069ffaf3d6da906d431957987ae26751fe123b58d6e3553a883247d5a4de38e

Observation fcb0f193-1d3f-4d51-bf99-f699c382f548 · outbound

This paper cites Holistic Evaluation of Language Models.

Human-Centric Evaluation for Foundation Models Holistic Evaluation of Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.881359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.881359Z digest=sha256:e4355e106e6df9afb71da9adca1b0e3050ea951b1afa945ce2f5336a3c3cd7ec

Observation b7d07220-01ed-4edf-855e-d4ed56d2bb6a · outbound

This paper cites Nature Machine Intelligence 5, 1 (2023), 46–57.

Human-Centric Evaluation for Foundation Models Nature Machine Intelligence 5, 1 (2023), 46–57

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:11.617747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T11:36:10.530656Z digest=sha256:214029a80522a67861c3ead09060bf68b2f941d08eb62e527b1ca1d07bd83ddb

Pith citing papers

Observation 9e614170-5bc0-409a-a2c2-1da423bb0e56 · inbound

Affordance Benchmark for MLLMs cites this paper.

Affordance Benchmark for MLLMs Human-Centric Evaluation for Foundation Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:53.097164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:53.097164Z digest=sha256:865a90e48dd4b33b476286f94e7fdc895903764ffd0d4efae1e4e8edae11fde3

Observation 6a1415e8-4082-45cf-a4ea-ad53bda88d4d · inbound

First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows cites this paper.

First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows Human-Centric Evaluation for Foundation Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:21:06.094847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T03:47:54.377400Z digest=sha256:7d267013bc5eb3a14f9f637cc3afff4fa37bf04fa12399ef16fc2698c3fa0f24

Observation 548330fb-bdcc-40f8-b0eb-a81748eef3eb · inbound

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning cites this paper.

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning Human-Centric Evaluation for Foundation Models

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:25:41.540358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-01T05:33:52.027771Z digest=sha256:534bdc0d0e2433288e6981aa4ff699097592a3dc1dc1cb1aed25f3629fd55f2f