Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:13:12.902147Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 4 inbound Pith citation observations for arXiv:2507.00769.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:13:12.902147Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T07:22:25.349398Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
27 of 27 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation ef8ed2bf-a7c3-4489-99eb-6f312d11a0c4 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 160a1d43-6ae7-4e53-bd73-e4231c4001a8 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Michele Elam
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3734a281-009a-463b-92e3-f5f32ad260de · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0f362a63-d5b8-4501-b82a-401b6dac941a · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Hierarchical Neural Story Generation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 790eb472-0233-4070-a125-d8f79dc8d3ce · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c53d829a-8867-473e-965c-79136e8a2a20 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Measuring Mathematical Problem Solving With the MATH Dataset
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5454d203-8d80-4b60-8f56-e155b6c3a38d · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e702214-1cbc-444d-b801-93ab3005e6b5 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 540ace0e-2b2d-453c-87f7-81efab574bbd · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Danielle S McNamara, Scott A Crossley, and Philip M McCarthy
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e2a442eb-6d4a-4234-a29d-d7e774f546db · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing URL https://store.macmillanlearning.com/us/product/ The-Practice-of-Creative-Writing/p/1319215955
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7f42e472-9288-48a4-b21e-52497005cb66 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Large Language Models are not Fair Evaluators
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da944267-1197-4876-a579-307bef239571 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb23fb3a-4549-4e08-ba92-dce7597f81b9 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a6a7bc0-6122-4e84-b9a1-4aade24a303d · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8d50100-4b50-41ba-87b6-88cd5fd5e405 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd4f0ce7-ed4d-46c5-a31b-24a4de1bb202 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Ralph Allan Bradley and Milton E Terry
Reference 1960
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 396c0b59-ba62-47b4-8706-57861229aab1 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Heather Sellers
Reference 1961
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4e8f375f-9296-4f6d-b565-707d3a1291c8 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Sher Badshah and Hassan Sajjad
Reference 1982
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cdb43df-3c7c-47d1-b6df-57627e9f6525 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Reference 2002
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c73d6e15-03c8-42fd-b98f-58feffac5554 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Danial Alihosseini, Ehsan Montahaei, and Mahdieh Soleymani Baghshah
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8c121b04-ca26-4562-a6da-517cc74fe131 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing doi: 10.18653/v1/W19-2311
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3dd84b21-46ad-4476-be1d-040e1a3b7021 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Learning to summarize from human feedback
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1f6de1b-dcb4-4290-ba53-2d2310d64a66 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97d114a7-7b7c-4d69-9b6d-911abf566ec4 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Training language models to follow instructions with human feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f0f115e-15bc-4373-8a33-6c11120d099c · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing URL https: //doi.org/10.1111/exsy.13292
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e6d6e5d1-a24b-4a5b-8604-3191bdde2e67 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d53d8de-684d-42e3-a7ef-c7511b2986a6 · outbound
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Modifying Large Language Model Post-Training for Diverse Creative Writing
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce6660c-0d80-466f-b9e7-3c2317bc2b09 · inbound
StoryAlign: Evaluating and Training Reward Models for Story Generation LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 483e377e-1c38-40ac-b50f-427e716ce30b · inbound
Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bb475fb5-907b-4599-8794-87e0c302ca8e · inbound
Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e18507b1-8146-4282-a1eb-620798cc0b66 · inbound
HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.