Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:34:42.445714Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2504.15219.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:34:42.445714Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:28:02.231013Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T13:28:02.879377Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dc80ca97-e0d2-4389-88bf-08dd822a47f8 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Art or Artifice? Large Language Models and the False Promise of Creativity
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8bf1413d-b26a-4b87-867f-348f66932095 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69546b1e-1f75-4132-827f-40495f4dd5a5 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Complex Claim Verification with Evidence Retrieved in the Wild
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d06759a-6ef0-478d-a0c1-1fdeb0a93276 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3985c6c8-bb01-497a-b7bf-2406a0c5be53 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c962e89d-71ce-44e6-829c-f7db689d7c1b · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Gonzalez
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 52971542-9699-47ac-9a56-320483853afa · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adba927f-252c-4890-b8ab-a746d39eef5b · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c348ca0-2464-4ca0-913a-9373a1a9281b · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99eacea3-4d3c-4d4c-8e35-5aac3f5fdfda · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Linguistically-Informed Specificity and Semantic Plausibility for Dialogue Generation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 12ccbb68-75cb-41a7-afaf-d1e0d590d972 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Large Language Models Are State-of-the-Art Evaluators of Translation Quality
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7342278c-4d72-44ab-8575-f2f90ce1aacc · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Hashimoto
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 90a8c738-2023-4469-90ea-8bf2bb399a19 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Wildbench: Benchmarking LLM s with challenging tasks from real users in the wild
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2b6920b-4535-4644-b467-0f4012614289 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Self-Refine: Iterative Refinement with Self-Feedback
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5d9e88f1-b382-4fac-ba89-99bf33df1689 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web ExpertQA: Expert-Curated Questions and Attributed Answers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70159c1a-3ed7-4fe4-b530-5500477f9a25 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Dolomites: Domain-Specific Long-Form Methodical Tasks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f429166-e12e-4055-bc4d-b0f351b5269b · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c71e040-0d1a-411e-9b3d-05ecea94e92f · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao, and Armando Solar-Lezama
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3c210ad4-14a1-4870-89f0-2571ad844cf3 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web InFoBench: Evaluating Instruction Following Ability in Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1ff29d4-0a29-4ec6-9522-9d03deb9ac29 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Do LLMs Plan Like Human Writers? Comparing Journalist Coverage of Press Releases with LLMs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b088437-5a3f-4471-9a81-7898f9261f3b · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0386a8b-126c-45c7-ba52-a148b8c25294 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Using natural language explanations to rescale human judgments
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 266acaca-ff34-41b3-84b9-c5aedf4f34f5 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Learning to Refine with Fine-Grained Natural Language Feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61b8e228-8783-4e71-88c8-65c7bee1ed19 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web WritingBench: A Comprehensive Benchmark for Generative Writing
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f90d70-e7da-4f5f-9053-a5fe3b0fe9b3 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation be34f127-a47e-4eab-a3cd-84903821029b · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Hanjie, Runzhe Yang, and Karthik R Narasimhan
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4f45d2a8-3615-4530-a97f-2505229d2333 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Self-Rewarding Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e16b302-6e24-4253-b436-5ebc178351e8 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 82d1d276-15e5-4e96-becb-8f4aca36dc77 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Learning to Control the Specificity in Neural Response Generation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 22ccd00a-76de-4e38-90c6-d54f38d16858 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 841621a3-eabb-478c-bcb5-bca8d6fbebdc · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Instruction-Following Evaluation for Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b70ff4d5-af34-457f-9fad-2c108ef8d972 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web write newline
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab15910-6e9c-4742-8b63-95b0ea2be823 · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web @esa (Ref
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b725652b-dc17-47ff-9468-f79dd0acbfac · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 749d0ce5-4339-44c9-bda6-780be7e3582a · outbound
EvalAgent: Discovering Implicit Evaluation Criteria from the Web HF f? -`U w DZYH+ RX AI 0 6g:B0ʁEڑD @Q Vx 6y!
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dece3487-360c-4aff-851e-49ebc34f9eb9 · inbound
BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum EvalAgent: Discovering Implicit Evaluation Criteria from the Web
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ce7d7a5e-20c1-46ec-be8d-f92522bbf4ee · inbound
Self-Improvements in Modern Agentic Systems: A Survey EvalAgent: Discovering Implicit Evaluation Criteria from the Web
Reference 1948
Source-reported events for the cited work
Unavailable: canonical work link unavailable.