Pith. sign in

Paper Citation Record · LEDGER

ChatBench: From Static Benchmarks to Human-AI Evaluation

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2504.07114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.07114 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:21:52.487178Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T05:05:54.821162Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 526e8617-104d-40f6-ab93-e9d62e1244be · inbound

Thinking beyond the anthropomorphic paradigm benefits LLM research cites this paper.

Thinking beyond the anthropomorphic paradigm benefits LLM research ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T22:21:52.487178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:21:52.487178Z digest=sha256:dab6b720401234695c06c5fb812d131272e0fd028b5a65a73d1be774608e9d5b

Observation 2790fd59-30e4-41f4-81cc-bf53fcf0615d · inbound

LLMs Get Lost In Multi-Turn Conversation cites this paper.

LLMs Get Lost In Multi-Turn Conversation ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:11:09.390293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T00:57:10.262350Z digest=sha256:9bf2cbf5e824735387ff84ccc4ded1217fad05caca5691505acb95d701b96e54

Observation 5989d867-c302-43ac-a69b-f03d7820edff · inbound

ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations cites this paper.

ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:38.646723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:23:38.646723Z digest=sha256:7c9a3c0bf251aac8accd1b1ed0ae9101229ab73bcecbee956abc711944ca7e32

Observation 570370e0-e76e-4430-9284-c1068694b47e · inbound

Potemkin Understanding in Large Language Models cites this paper.

Potemkin Understanding in Large Language Models ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:29.181666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:29.181666Z digest=sha256:1388186f03a6065dd89e2939fa3a733b01dfa3b324b65901c6c76b997d7108c7

Observation e4479443-39dc-4cf1-96c9-f96965848b79 · inbound

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models cites this paper.

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:07.543314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:16:07.543314Z digest=sha256:57627d9b6cb46674e17daf383e30be0ed6b45fc1e08b507578aa051faf99f6b0

Observation f6526d21-e822-4690-ae2d-1722ca13d9c4 · inbound

Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards cites this paper.

Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:05:54.829589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T05:04:10.226166Z digest=sha256:8724bea975b34d41631283584a783c8809b5095890b1f1b3759d81291f421253

Observation 37b2e37a-518e-46b9-b839-f64d1398afe8 · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T11:49:31.595873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:49:31.595873Z digest=sha256:32cfbc84ba2afe8076b76e9eab17ffd57cdce882ad9411a8cc603b87a1cd504a

Observation 52ea65f1-f84f-45ff-b0fc-ca81773892be · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T07:20:34.591235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:20:34.591235Z digest=sha256:a8ccff3d421bace235fb1e3c995493bbf3f3305c814dfd45642ce2c34404aee1