Pith. sign in

Paper Citation Record · LEDGER

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

As of 9 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 4 inbound Pith citation observations for arXiv:2507.00769.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00769 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:13:12.902147Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T07:22:25.349398Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact5
  • verified fuzzy4
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ef8ed2bf-a7c3-4489-99eb-6f312d11a0c4 · outbound

This paper cites Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:44.575215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:44.575215Z digest=sha256:d0d0c1ca263de235833e388a20af962ea2c6d77df9cde0bf86624cf81000b657

Observation 160a1d43-6ae7-4e53-bd73-e4231c4001a8 · outbound

This paper cites Michele Elam.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Michele Elam

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:13:13.988740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:12:44.647758Z digest=sha256:7ea0d09c9fa1f83e434cb5f0272d6e203559241270ea3770564950dcc20291cb

Observation 3734a281-009a-463b-92e3-f5f32ad260de · outbound

This paper cites Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:13:13.716456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:12:44.712552Z digest=sha256:30178eb447649185f35cdc68541f69fb55485985eeec16267db4580155a36440

Observation 0f362a63-d5b8-4501-b82a-401b6dac941a · outbound

This paper cites Hierarchical Neural Story Generation.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Hierarchical Neural Story Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:44.783846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:44.783846Z digest=sha256:9aaa9dc14edb727de00bea2fb8bfa51a79f226bd9a16297e072c56b0d555268a

Observation 790eb472-0233-4070-a125-d8f79dc8d3ce · outbound

This paper cites Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:44.870655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:44.870655Z digest=sha256:f4f95975651591b0f0b26637e0738b3aec0ffd951388e34d67bc5c1ed6b2e01f

Observation c53d829a-8867-473e-965c-79136e8a2a20 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Measuring Mathematical Problem Solving With the MATH Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:44.939856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:44.939856Z digest=sha256:7c26eee3886818553a8445be60059bfde538af2315f9065ace9fc12bd29cdbbc

Observation 5454d203-8d80-4b60-8f56-e155b6c3a38d · outbound

This paper cites Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.724054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.724054Z digest=sha256:dc94da674a48c66cac0c7b48a7488ca7305b52472954f00de1cde0c255194aa8

Observation 7e702214-1cbc-444d-b801-93ab3005e6b5 · outbound

This paper cites G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.744670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.744670Z digest=sha256:105c17973a089ffa420e291e8d03ace56b2d6e4599c15480e2953f83ab9a8a58

Observation 540ace0e-2b2d-453c-87f7-81efab574bbd · outbound

This paper cites Danielle S McNamara, Scott A Crossley, and Philip M McCarthy.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Danielle S McNamara, Scott A Crossley, and Philip M McCarthy

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:13:13.916229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:13:12.763796Z digest=sha256:983880a67da251c2a3dbe97b5fed9d6da8d41b9e0f1ffc34a5e4418b0523ffc2

Observation e2a442eb-6d4a-4234-a29d-d7e774f546db · outbound

This paper cites URL https://store.macmillanlearning.com/us/product/ The-Practice-of-Creative-Writing/p/1319215955.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing URL https://store.macmillanlearning.com/us/product/ The-Practice-of-Creative-Writing/p/1319215955

Reference 20

Resolution
verified exact
raw_fallback, observed 2026-08-06T21:13:13.352387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:13:12.796800Z digest=sha256:c5b5ef9395bcbca4862f035082810e3b627bf22c18c899d80e4761c7914aec9a

Observation 7f42e472-9288-48a4-b21e-52497005cb66 · outbound

This paper cites Large Language Models are not Fair Evaluators.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Large Language Models are not Fair Evaluators

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.830159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.830159Z digest=sha256:d37c76bcf7039fa75c25271a1ce05026ddfcb2e2d08fa81790221e0d6c3172f1

Observation da944267-1197-4876-a579-307bef239571 · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.843092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.843092Z digest=sha256:19ec483ab1de13c310501950105a550f28944203323997ceb7b315f3a9ed15d3

Observation cb23fb3a-4549-4e08-ba92-dce7597f81b9 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.854546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.854546Z digest=sha256:a86131548fca3602cced2043628319a803f8ca8a1c5224e8ed264cf9983f465a

Observation 0a6a7bc0-6122-4e84-b9a1-4aade24a303d · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.874087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.874087Z digest=sha256:ba74e7d9de9374d05798217a6f3fc7445dbeaf0efc700665ee6ebd0d17af2a9a

Observation e8d50100-4b50-41ba-87b6-88cd5fd5e405 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.902147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.902147Z digest=sha256:2a11c5157b17b5f11d783d4c9512906318095cb15d21d70de4f34a5507ed14bd

Observation bd4f0ce7-ed4d-46c5-a31b-24a4de1bb202 · outbound

This paper cites Ralph Allan Bradley and Milton E Terry.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Ralph Allan Bradley and Milton E Terry

Reference 1960

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:13:14.072352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:12:44.302548Z digest=sha256:f961e0c3a86ec8915b15e88100c1f4c5d72f4fd7ac250b88a7b5ca1a938e1ddf

Observation 396c0b59-ba62-47b4-8706-57861229aab1 · outbound

This paper cites Heather Sellers.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Heather Sellers

Reference 1961

Resolution
verified exact
raw_fallback, observed 2026-08-06T21:13:13.511638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:13:12.782574Z digest=sha256:4b5f13382936f26bf41ba563b832938f10928a32a360ef86e2432e5be7db4827

Observation 4e8f375f-9296-4f6d-b565-707d3a1291c8 · outbound

This paper cites Sher Badshah and Hassan Sajjad.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Sher Badshah and Hassan Sajjad

Reference 1982

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:44.121415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:44.121415Z digest=sha256:9899dffa1c9c3a015bad1927313383c293d70cda53a1def537b0e5c2076e4ebe

Observation 6cdb43df-3c7c-47d1-b6df-57627e9f6525 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.864529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.864529Z digest=sha256:93cd1537cb9a1d67fb7d6b2102b1ba3f78012600400332cb6c87d3d1f6c4ea51

Observation c73d6e15-03c8-42fd-b98f-58feffac5554 · outbound

This paper cites Danial Alihosseini, Ehsan Montahaei, and Mahdieh Soleymani Baghshah.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Danial Alihosseini, Ehsan Montahaei, and Mahdieh Soleymani Baghshah

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:13:14.145774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:12:44.004545Z digest=sha256:6ee0ccea4205b227ce16cb61cffd522bc82ecafecda496af8d17833f195d1ed2

Observation 8c121b04-ca26-4562-a6da-517cc74fe131 · outbound

This paper cites doi: 10.18653/v1/W19-2311.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing doi: 10.18653/v1/W19-2311

Reference 2019

Resolution
verified exact
doi, observed 2026-08-06T21:13:13.021589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:12:44.053218Z digest=sha256:d40d5e7e754b80e88fb7f45214849dc0d6278dbe57a163436e6e22596cf25d92

Observation 3dd84b21-46ad-4476-be1d-040e1a3b7021 · outbound

This paper cites Learning to summarize from human feedback.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Learning to summarize from human feedback

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.819935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.819935Z digest=sha256:23f096f764ce461f728d5ebab40f23635b37fb85eabc84d044d9c5f972543783

Observation e1f6de1b-dcb4-4290-ba53-2d2310d64a66 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:45.023262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:45.023262Z digest=sha256:b0a869864ee54ca71fa0bb6b62972601f4eb7c4c3ae86366ddb08d6c3dd5b030

Observation 97d114a7-7b7c-4d69-9b6d-911abf566ec4 · outbound

This paper cites Training language models to follow instructions with human feedback.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.773984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.773984Z digest=sha256:3242e66453dcd753b0c2af537ee2c795801034b8c9f8910a56b3b4cfaa68755a

Observation 3f0f115e-15bc-4373-8a33-6c11120d099c · outbound

This paper cites URL https: //doi.org/10.1111/exsy.13292.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing URL https: //doi.org/10.1111/exsy.13292

Reference 2023

Resolution
verified exact
doi, observed 2026-08-06T21:13:12.961108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:12:44.379126Z digest=sha256:f163e70c1f741f9742e5b6874cb011491613b54a13f7bcf9f3fa465ba3717b9c

Observation e6d6e5d1-a24b-4a5b-8604-3191bdde2e67 · outbound

This paper cites Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:44.221989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:44.221989Z digest=sha256:4b2237e880cebc136ea11e445fecd20191bc5c036715a35cbeed7394b6f02195

Observation 4d53d8de-684d-42e3-a7ef-c7511b2986a6 · outbound

This paper cites Modifying Large Language Model Post-Training for Diverse Creative Writing.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Modifying Large Language Model Post-Training for Diverse Creative Writing

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:44.473350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:44.473350Z digest=sha256:3f85d7039d7b8503796dda9677bcaf088b52902d4a31b6ef19b263e5a30e4cdc

Pith citing papers

Observation 1ce6660c-0d80-466f-b9e7-3c2317bc2b09 · inbound

StoryAlign: Evaluating and Training Reward Models for Story Generation cites this paper.

StoryAlign: Evaluating and Training Reward Models for Story Generation LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:31:07.707466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:29:13.549559Z digest=sha256:2e3d0c0195a41b4b1b4a898317f4806d31ef361c0b73deb79ae28433567be004

Observation 483e377e-1c38-40ac-b50f-427e716ce30b · inbound

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches cites this paper.

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:42:25.073554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:35:01.285534Z digest=sha256:2226e87c503b2130448645d96a8de13a0795eea1feba3561b2a63884a1fbe6fc

Observation bb475fb5-907b-4599-8794-87e0c302ca8e · inbound

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches cites this paper.

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:25:28.228212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:22:25.349398Z digest=sha256:6b5af55e13c4b2a3f6d1cc1b33e09127a0bcb9b1a0590b84c477e4746b9ecf17

Observation e18507b1-8146-4282-a1eb-620798cc0b66 · inbound

HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice cites this paper.

HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T18:00:01.338538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T23:08:27.995672Z digest=sha256:e1317d7caf15366f4777987764a6bfc033c27c56104780002cf753c6e0170191