Pith. sign in

Paper Citation Record · LEDGER

ChatBench: From Static Benchmarks to Human-AI Evaluation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2504.07114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.07114 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:21:52.487178Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T05:05:54.821162Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 526e8617-104d-40f6-ab93-e9d62e1244be · inbound

Thinking beyond the anthropomorphic paradigm benefits LLM research cites this paper.

Thinking beyond the anthropomorphic paradigm benefits LLM research ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T22:21:52.487178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:21:52.487178Z digest=sha256:6b8d7e76030428e327258bcc4082eabb5b43d059d8c6364bb91754975132e025

Observation 2790fd59-30e4-41f4-81cc-bf53fcf0615d · inbound

LLMs Get Lost In Multi-Turn Conversation cites this paper.

LLMs Get Lost In Multi-Turn Conversation ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:11:09.390293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T00:57:10.262350Z digest=sha256:086668cf513292ac0016f9a8056a13c95938de8d26bc8b71aed19e367101b9c5

Observation 5989d867-c302-43ac-a69b-f03d7820edff · inbound

ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations cites this paper.

ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:38.646723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:23:38.646723Z digest=sha256:4c49363806e4d6175dc34a0e900b4b254ef5421ee7ce4c49e27e4831c9c11973

Observation 570370e0-e76e-4430-9284-c1068694b47e · inbound

Potemkin Understanding in Large Language Models cites this paper.

Potemkin Understanding in Large Language Models ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:29.181666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:29.181666Z digest=sha256:6d5eba1d354b86c3d0a0c961d294d0832c1df341326833b053170280d2452015

Observation e4479443-39dc-4cf1-96c9-f96965848b79 · inbound

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models cites this paper.

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:07.543314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:16:07.543314Z digest=sha256:06bf81697098376669080bf5d4f7e89583ec519c50b087d0d31a4dcd9ba3ca43

Observation f6526d21-e822-4690-ae2d-1722ca13d9c4 · inbound

Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards cites this paper.

Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:05:54.829589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:04:10.226166Z digest=sha256:e57b8577f63a97408004f2f255b5c20b8e36419967433f4a9cba1c948043be84

Observation 37b2e37a-518e-46b9-b839-f64d1398afe8 · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T11:49:31.595873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:49:31.595873Z digest=sha256:b830adf9dafaf9a083f3ff6b8207f294f03c9b60a749fc4c27b23a94d9639ce6

Observation 52ea65f1-f84f-45ff-b0fc-ca81773892be · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T07:20:34.591235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:20:34.591235Z digest=sha256:c9d10d9cd2edbb8008db21a1ca24bf980f18261461350f3e2a9048306d444a99