Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:35:10.178460Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 3 inbound Pith citation observations for arXiv:2412.04363.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:35:10.178460Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:44:33.037670Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T13:00:21.940022Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c4bc876d-4547-4da1-9aa7-62bde285d456 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7477e201-baa7-43a6-a740-a8563303c002 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a3c6876-39e4-48f6-a107-a8bab6e4c898 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Thomas Adler and Luca de Alfaro
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 627bd83c-8148-42d8-b296-061a80484f19 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1026ae3-a743-491e-9326-0721f2ae74ab · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Evaluation of Text Generation: A Survey
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30e822cc-5a28-4b9f-b210-8cd64862ee81 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86ef0ec2-be88-4443-9601-babaf1ef591a · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4fe1221b-39ae-439d-8616-2ab5a147d26b · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8ab24dbb-d804-4966-8dd8-a31c9d0e06f0 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots A case for better evaluation standards in nlg
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cdc81f2e-1c9f-4158-a7d0-8d617f9092fe · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb7018e9-b305-435f-bd3d-50a0a7a42a9b · outbound
Challenges in Trustworthy Human Evaluation of Chatbots News Summarization and Evaluation in the Era of GPT-3
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac50e686-be95-4588-bd0e-472f70a48937 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 14882b1e-8046-4cf9-b64a-d4c89529d12a · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 749e2448-a9c7-4a9d-8be8-9724388c6623 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eb18da14-56c3-4dce-9a85-be0658210046 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 147fd015-165a-405e-968b-7c56670da86e · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 568daef6-fd59-497d-a797-f7de8a6ec442 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fc75ec8-d3e3-40cf-a70b-ac6f6b376cbf · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1a626e91-b9fa-47c9-a2e9-a1460d957bae · outbound
Challenges in Trustworthy Human Evaluation of Chatbots RewardBench: Evaluating Reward Models for Language Modeling
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4c05b23-1f7f-4313-930c-32e1663f9e9d · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Gonzalez, and Ion Stoica
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 56aaa448-f7b0-4eca-84af-f25409091567 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Hashimoto
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c94a94d0-8d79-4acf-bc1e-693d7a76a0f6 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e4f43f4-92d3-4ab2-90a8-eaf9f3b87713 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4a7079c-a578-4769-b67e-71545909be96 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c05f4b-b4d7-4e9d-92b6-3a6c95a97e8b · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5bb5e92d-9583-499e-abeb-39b9c576a79d · outbound
Challenges in Trustworthy Human Evaluation of Chatbots MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e03030ea-289d-4050-bf43-410f4e551efd · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28e57caa-4b32-46c7-8932-adda8675dea2 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f935ff2-3f31-407f-a3f8-6a99a9847034 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 423f78c5-04c0-4345-a01b-a13d532e3ae8 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3c18d4c3-5d27-4f57-afcc-30d57baadddf · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ce16abef-d296-4c55-8bb9-df3245f1f573 · outbound
Challenges in Trustworthy Human Evaluation of Chatbots Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cd424cc-46e3-4707-91d2-eabc9c3d152b · inbound
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards Challenges in Trustworthy Human Evaluation of Chatbots
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 876afd0a-3cc5-4ed7-9913-991092bfe31e · inbound
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild Challenges in Trustworthy Human Evaluation of Chatbots
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db0840cf-73eb-4d3a-80db-cf3cfba17788 · inbound
Curiosity by Design: An LLM-based Coding Assistant Asking Clarification Questions Challenges in Trustworthy Human Evaluation of Chatbots
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.