Pith. sign in

Paper Citation Record · LEDGER

Challenges in Trustworthy Human Evaluation of Chatbots

As of 13 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 3 inbound Pith citation observations for arXiv:2412.04363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04363 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:35:10.178460Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:44:33.037670Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T13:00:21.940022Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4bc876d-4547-4da1-9aa7-62bde285d456 · outbound

This paper cites online" 'onlinestring :=.

Challenges in Trustworthy Human Evaluation of Chatbots online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.017736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.017736Z digest=sha256:6886579e999a827c8aeea69839b6c6a2a01f103b21458e955de39eb0e7e81500

Observation 7477e201-baa7-43a6-a740-a8563303c002 · outbound

This paper cites write newline.

Challenges in Trustworthy Human Evaluation of Chatbots write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.023896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.023896Z digest=sha256:35d03b8f3b94a814ed579c66e93a4a4acae2cf095c7c717878ed8c831d6e5f0b

Observation 4a3c6876-39e4-48f6-a107-a8bab6e4c898 · outbound

This paper cites Thomas Adler and Luca de Alfaro.

Challenges in Trustworthy Human Evaluation of Chatbots Thomas Adler and Luca de Alfaro

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:35:10.657392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T21:35:10.029664Z digest=sha256:c310b91c569eb098e91b2de6af7b8aacfad2a74400da96f396afa6de94ccc440

Observation 627bd83c-8148-42d8-b296-061a80484f19 · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.036132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.036132Z digest=sha256:0eb28e0f804c27949bf473d1cb7a66fd1df083d9926b9923310cd79ecbf47d1d

Observation d1026ae3-a743-491e-9326-0721f2ae74ab · outbound

This paper cites Evaluation of Text Generation: A Survey.

Challenges in Trustworthy Human Evaluation of Chatbots Evaluation of Text Generation: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.041377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.041377Z digest=sha256:5cd9bb9c66f14b9fd45254bf5b888699c124e69589a873376cf0ab0dcc479158

Observation 30e822cc-5a28-4b9f-b210-8cd64862ee81 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Challenges in Trustworthy Human Evaluation of Chatbots Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.046855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.046855Z digest=sha256:b084dd0cad32f3f7348cc56637672f88da637a26a1471b8211ebf0d12c0d824e

Observation 86ef0ec2-be88-4443-9601-babaf1ef591a · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:35:10.630005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T21:35:10.052150Z digest=sha256:5f3777ae6743c060557ba041541ffba5596e37161cb2520ff4ede3ff536707da

Observation 4fe1221b-39ae-439d-8616-2ab5a147d26b · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:35:10.612211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T21:35:10.057205Z digest=sha256:420bb20350f75450b49e986bf37fdf3e5be6f51736d9432db843d92854109a5b

Observation 8ab24dbb-d804-4966-8dd8-a31c9d0e06f0 · outbound

This paper cites A case for better evaluation standards in nlg.

Challenges in Trustworthy Human Evaluation of Chatbots A case for better evaluation standards in nlg

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:35:10.595554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T21:35:10.062830Z digest=sha256:c8101f6c2077b5e476a17f360a53c40e419043110f4df133990c532ad0c07698

Observation cdc81f2e-1c9f-4158-a7d0-8d617f9092fe · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.067838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.067838Z digest=sha256:1f3673b27687908d77cfedb5441a9f89b54ffed9a5646a22630dcaa34a808e2c

Observation eb7018e9-b305-435f-bd3d-50a0a7a42a9b · outbound

This paper cites News Summarization and Evaluation in the Era of GPT-3.

Challenges in Trustworthy Human Evaluation of Chatbots News Summarization and Evaluation in the Era of GPT-3

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.072571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.072571Z digest=sha256:17ee4b71cc89607dabb0b88b373ebd3e52571c68d397bf2e4a2bf5dbc7fb9d61

Observation ac50e686-be95-4588-bd0e-472f70a48937 · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:35:10.567852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T21:35:10.078423Z digest=sha256:561124b2716ab821aad62c7d714e4cddf5eda84b1043989c73aebc72ba894878

Observation 14882b1e-8046-4cf9-b64a-d4c89529d12a · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:35:10.551984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T21:35:10.083979Z digest=sha256:ee09f0ed1243135092acf4935c3784d3db784a8b373e9bfb560d0b6a07a4ef1a

Observation 749e2448-a9c7-4a9d-8be8-9724388c6623 · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:35:10.534217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T21:35:10.089048Z digest=sha256:e1addbfeb63d812ccd38ec5a81f236d186c4f606b6ef066bca97ac459df35919

Observation eb18da14-56c3-4dce-9a85-be0658210046 · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.093790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.093790Z digest=sha256:e47effaa6fb13037e537ce8ea375fd431ac08828279feeb40f195c53a80e46ea

Observation 147fd015-165a-405e-968b-7c56670da86e · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.098641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.098641Z digest=sha256:8e5ec0cba1e239a53051ecd21a658599b9d06d9c6871e5e778c88fdedbf5e21d

Observation 568daef6-fd59-497d-a797-f7de8a6ec442 · outbound

This paper cites LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization.

Challenges in Trustworthy Human Evaluation of Chatbots LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.103400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.103400Z digest=sha256:ec05ea2ae1a9ff5134cbcad90a95b57a9d52cb0940bdf54826d69f4a53d5bd22

Observation 1fc75ec8-d3e3-40cf-a70b-ac6f6b376cbf · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:35:10.496786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T21:35:10.108712Z digest=sha256:c312b036bf756b4aa902efd2940aa9acd089633158e51a94e30e544e0a29f5c4

Observation 1a626e91-b9fa-47c9-a2e9-a1460d957bae · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Challenges in Trustworthy Human Evaluation of Chatbots RewardBench: Evaluating Reward Models for Language Modeling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.113495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.113495Z digest=sha256:589b903180c506654d93ab1cc690042ba8087e48fd5f188fc7acadba8d2d9b43

Observation c4c05b23-1f7f-4313-930c-32e1663f9e9d · outbound

This paper cites Gonzalez, and Ion Stoica.

Challenges in Trustworthy Human Evaluation of Chatbots Gonzalez, and Ion Stoica

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:35:10.480018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T21:35:10.118860Z digest=sha256:c6959435b3a3215b1b43b29fc922fae83e5d270401734b5ba7cb45098c85bfe1

Observation 56aaa448-f7b0-4eca-84af-f25409091567 · outbound

This paper cites Hashimoto.

Challenges in Trustworthy Human Evaluation of Chatbots Hashimoto

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.123626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.123626Z digest=sha256:c6ebbc19786b5bac9db2fba5e92dd741611c79a755c41f6a3d00eb919084bdeb

Observation c94a94d0-8d79-4acf-bc1e-693d7a76a0f6 · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

Challenges in Trustworthy Human Evaluation of Chatbots WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.128299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.128299Z digest=sha256:9900e3491b4695f1aea0e7ce116b627e84eb7d77220254c7cea555107e4d2e33

Observation 7e4f43f4-92d3-4ab2-90a8-eaf9f3b87713 · outbound

This paper cites WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences.

Challenges in Trustworthy Human Evaluation of Chatbots WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.133432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.133432Z digest=sha256:04e63094fcac7cae310bbcf85be0730409c4f217cd75310c7309b486994f23a7

Observation e4a7079c-a578-4769-b67e-71545909be96 · outbound

This paper cites Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries.

Challenges in Trustworthy Human Evaluation of Chatbots Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.138711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.138711Z digest=sha256:31adf096e7ca78ec064776fd3984e9f2e2c24435fabf3ef5206f76ff075ab5a9

Observation b3c05f4b-b4d7-4e9d-92b6-3a6c95a97e8b · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:35:10.452181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T21:35:10.143834Z digest=sha256:c7cc336600df94d677f0106e6861418ae08a2943fd96a6d183d281a088a0d982

Observation 5bb5e92d-9583-499e-abeb-39b9c576a79d · outbound

This paper cites MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures.

Challenges in Trustworthy Human Evaluation of Chatbots MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.148411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.148411Z digest=sha256:caaecd683817935c25a6da4f998b9f08617a58655144c5bbea8a692f0f74a0ba

Observation e03030ea-289d-4050-bf43-410f4e551efd · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.153553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.153553Z digest=sha256:5095770bd0eeec06707cfd77701acfa1f086ac7692f787e9397f07a09c161689

Observation 28e57caa-4b32-46c7-8932-adda8675dea2 · outbound

This paper cites Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents.

Challenges in Trustworthy Human Evaluation of Chatbots Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.158763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.158763Z digest=sha256:aefa44952698501e98b7e984c1828e54bf7aaee4f1348cadbaa0d163cf03fd74

Observation 3f935ff2-3f31-407f-a3f8-6a99a9847034 · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:35:10.425182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T21:35:10.164095Z digest=sha256:8b73e8c629297d5f06e96de6e9327238d21c7ccab206d51182549e78d3e88cdf

Observation 423f78c5-04c0-4345-a01b-a13d532e3ae8 · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:35:10.408342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T21:35:10.169004Z digest=sha256:31a2f574e59f99df0ec5a89dcdf82af89c8c4a6864a2e414de021b5fbe0ccb22

Observation 3c18d4c3-5d27-4f57-afcc-30d57baadddf · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:35:10.390132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T21:35:10.173836Z digest=sha256:f301081dee853f90ca58ba62cf4f31e264331c9777cbf9ea13951978a1384c5c

Observation ce16abef-d296-4c55-8bb9-df3245f1f573 · outbound

This paper cites an unresolved cited work.

Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:10.178460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:35:10.178460Z digest=sha256:230b8480a4b44385c5af88e9d0344382c07359abc650c33261448cb335f079f2

Pith citing papers

Observation 6cd424cc-46e3-4707-91d2-eabc9c3d152b · inbound

Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards cites this paper.

Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards Challenges in Trustworthy Human Evaluation of Chatbots

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:44:33.037670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:44:33.037670Z digest=sha256:2e750fcd44496555b6e4f253381988e054a6c6857cb1735acac6e9b7778684e3

Observation 876afd0a-3cc5-4ed7-9913-991092bfe31e · inbound

Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild cites this paper.

Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild Challenges in Trustworthy Human Evaluation of Chatbots

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:18.124381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:18.124381Z digest=sha256:96992c03bf131c8876732b30d2e406464884d796285bbdbf2defd0a83fd2fae0

Observation db0840cf-73eb-4d3a-80db-cf3cfba17788 · inbound

Curiosity by Design: An LLM-based Coding Assistant Asking Clarification Questions cites this paper.

Curiosity by Design: An LLM-based Coding Assistant Asking Clarification Questions Challenges in Trustworthy Human Evaluation of Chatbots

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:00:21.953135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T13:00:21.827159Z digest=sha256:9bc289d02e6471da9a796df3655c023dbe10f724e15e9662619cd8f72c1c0e2a