Pith. sign in

Paper Citation Record · LEDGER

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 4 inbound Pith citation observations for arXiv:2507.00769.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00769 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:13:12.902147Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T07:22:25.349398Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact5
  • verified fuzzy4
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ef8ed2bf-a7c3-4489-99eb-6f312d11a0c4 · outbound

This paper cites Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:44.575215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:44.575215Z digest=sha256:209b7bfced64e387ce89ddde2dc9666b836a4714ae7576bf5c919fda3f888e6e

Observation 160a1d43-6ae7-4e53-bd73-e4231c4001a8 · outbound

This paper cites Michele Elam.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Michele Elam

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:13:13.988740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:12:44.647758Z digest=sha256:c611673c0167685fd2b1dc6981cf0a0e62f7e19e244c1048e27e53696caecfdd

Observation 3734a281-009a-463b-92e3-f5f32ad260de · outbound

This paper cites Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:13:13.716456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:12:44.712552Z digest=sha256:2fd664be86a7a595cc13995082c81815bc61e3a9972368e26b7308834c97fb3f

Observation 0f362a63-d5b8-4501-b82a-401b6dac941a · outbound

This paper cites Hierarchical Neural Story Generation.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Hierarchical Neural Story Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:44.783846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:44.783846Z digest=sha256:91a12483626055cb2aba5793aa1a12c293b88ceba0822544ebbcb9333575bc8f

Observation 790eb472-0233-4070-a125-d8f79dc8d3ce · outbound

This paper cites Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:44.870655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:44.870655Z digest=sha256:eb904d9a013e40e2a44f801e39661754653765fe679534e544251d264f674a76

Observation c53d829a-8867-473e-965c-79136e8a2a20 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Measuring Mathematical Problem Solving With the MATH Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:44.939856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:44.939856Z digest=sha256:b011cb2be97ed9ad817521289e5af22feb3973cdb13c6d1b5c70df8b57c0ee51

Observation 5454d203-8d80-4b60-8f56-e155b6c3a38d · outbound

This paper cites Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.724054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.724054Z digest=sha256:6de193fa799057a8405dbde78573a9a1d1460172fedffe26dc1ff524ed5babc1

Observation 7e702214-1cbc-444d-b801-93ab3005e6b5 · outbound

This paper cites G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.744670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.744670Z digest=sha256:a57cead3aeeb2d9534992d9c02a8ced990683ca75a6ff15136d8535b49d2953b

Observation 540ace0e-2b2d-453c-87f7-81efab574bbd · outbound

This paper cites Danielle S McNamara, Scott A Crossley, and Philip M McCarthy.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Danielle S McNamara, Scott A Crossley, and Philip M McCarthy

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:13:13.916229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:13:12.763796Z digest=sha256:08dae011ca01829ea7eee437126e201ff90b00742570a2d24bb26cfcdec503cf

Observation e2a442eb-6d4a-4234-a29d-d7e774f546db · outbound

This paper cites URL https://store.macmillanlearning.com/us/product/ The-Practice-of-Creative-Writing/p/1319215955.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing URL https://store.macmillanlearning.com/us/product/ The-Practice-of-Creative-Writing/p/1319215955

Reference 20

Resolution
verified exact
raw_fallback, observed 2026-08-06T21:13:13.352387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:13:12.796800Z digest=sha256:3c88f2a71e6978bb2ddfd0012344deacd1df1ef455a907762123fbd82473ae5e

Observation 7f42e472-9288-48a4-b21e-52497005cb66 · outbound

This paper cites Large Language Models are not Fair Evaluators.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Large Language Models are not Fair Evaluators

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.830159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.830159Z digest=sha256:3c0cf927cf7b40fc12341e89214dc26468c28d7e5fd8828518da34946f0f06c4

Observation da944267-1197-4876-a579-307bef239571 · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.843092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.843092Z digest=sha256:e2be8ce1798b573ae384c2fb8232fb01bb644fd53d820b0aedbf1a2726f2283c

Observation cb23fb3a-4549-4e08-ba92-dce7597f81b9 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.854546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.854546Z digest=sha256:f9381f23cec8c5d03fea879d920ee324e2835817c313f5b5cf758b023871b6b7

Observation 0a6a7bc0-6122-4e84-b9a1-4aade24a303d · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.874087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.874087Z digest=sha256:ecbc71c927484ac29465c00c683de6060b775b0f48c803070ea3c0089cb74cc8

Observation e8d50100-4b50-41ba-87b6-88cd5fd5e405 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.902147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.902147Z digest=sha256:4b58136998dad1cf2ec5ed96073da92c624c768bc9056118b301d851acb8b6de

Observation bd4f0ce7-ed4d-46c5-a31b-24a4de1bb202 · outbound

This paper cites Ralph Allan Bradley and Milton E Terry.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Ralph Allan Bradley and Milton E Terry

Reference 1960

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:13:14.072352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:12:44.302548Z digest=sha256:6441756a319102074823a9a6bb3be1e20aafdbb0f391f8200e453a1568726515

Observation 396c0b59-ba62-47b4-8706-57861229aab1 · outbound

This paper cites Heather Sellers.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Heather Sellers

Reference 1961

Resolution
verified exact
raw_fallback, observed 2026-08-06T21:13:13.511638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:13:12.782574Z digest=sha256:d0d2545aa61043a61db298bc32ef2566d7a24d3956aa6fe96fa660d40cb46469

Observation 4e8f375f-9296-4f6d-b565-707d3a1291c8 · outbound

This paper cites Sher Badshah and Hassan Sajjad.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Sher Badshah and Hassan Sajjad

Reference 1982

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:44.121415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:44.121415Z digest=sha256:c9ae88cda8ccecd4896d80b56952d90530d4d5d3acc5a905a1114ca9cf546363

Observation 6cdb43df-3c7c-47d1-b6df-57627e9f6525 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.864529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.864529Z digest=sha256:5d3591d9ef6c12851ad977b2901d0e26bd6c778b700096c4a2cf88e597d34e4f

Observation c73d6e15-03c8-42fd-b98f-58feffac5554 · outbound

This paper cites Danial Alihosseini, Ehsan Montahaei, and Mahdieh Soleymani Baghshah.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Danial Alihosseini, Ehsan Montahaei, and Mahdieh Soleymani Baghshah

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:13:14.145774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:12:44.004545Z digest=sha256:749d532de2ac11046e28dd3683276e61ca4d66bf56ca05622912434181e858e0

Observation 8c121b04-ca26-4562-a6da-517cc74fe131 · outbound

This paper cites doi: 10.18653/v1/W19-2311.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing doi: 10.18653/v1/W19-2311

Reference 2019

Resolution
verified exact
doi, observed 2026-08-06T21:13:13.021589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:12:44.053218Z digest=sha256:0bf35a558388c2fcdd076eb30f54f709d98c7f3a03aa00df5805c2f9cfc61089

Observation 3dd84b21-46ad-4476-be1d-040e1a3b7021 · outbound

This paper cites Learning to summarize from human feedback.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Learning to summarize from human feedback

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.819935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.819935Z digest=sha256:599245cc6edc07a98c02e32cf9ff1384e38cfd08071174cf9d386eb5f12cb70d

Observation e1f6de1b-dcb4-4290-ba53-2d2310d64a66 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:45.023262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:45.023262Z digest=sha256:2d83112cd14fcad762ed583f6c205db56ffd44f7219ea6059ed030241201e072

Observation 97d114a7-7b7c-4d69-9b6d-911abf566ec4 · outbound

This paper cites Training language models to follow instructions with human feedback.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:12.773984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:13:12.773984Z digest=sha256:392d6d6cf6ebd64804f650d35c274b91dadb44dad50bad453b133948bb844b02

Observation 3f0f115e-15bc-4373-8a33-6c11120d099c · outbound

This paper cites URL https: //doi.org/10.1111/exsy.13292.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing URL https: //doi.org/10.1111/exsy.13292

Reference 2023

Resolution
verified exact
doi, observed 2026-08-06T21:13:12.961108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:12:44.379126Z digest=sha256:e1bf669e9d9093b4deda659824f49d539a90b73453793ef7d3551e8be4444e49

Observation e6d6e5d1-a24b-4a5b-8604-3191bdde2e67 · outbound

This paper cites Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:44.221989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:44.221989Z digest=sha256:6ad6eb01d5dc039ec91d146003c52692cfd20b62d4b93d2d7fa8e194d3eac8ee

Observation 4d53d8de-684d-42e3-a7ef-c7511b2986a6 · outbound

This paper cites Modifying Large Language Model Post-Training for Diverse Creative Writing.

LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing Modifying Large Language Model Post-Training for Diverse Creative Writing

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:44.473350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:44.473350Z digest=sha256:562504253b8ffaa43abd387ff522c5ac55dd93aa27ec5e0cbdc2caaff870efc1

Pith citing papers

Observation 1ce6660c-0d80-466f-b9e7-3c2317bc2b09 · inbound

StoryAlign: Evaluating and Training Reward Models for Story Generation cites this paper.

StoryAlign: Evaluating and Training Reward Models for Story Generation LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:31:07.707466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:29:13.549559Z digest=sha256:96a6345d22b2dbff57b6487c7c2d690d5ddd75619d35baba35f7c06b85914ceb

Observation 483e377e-1c38-40ac-b50f-427e716ce30b · inbound

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches cites this paper.

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:42:25.073554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:35:01.285534Z digest=sha256:155f2ad8467096c0b3c449169cec4b689a4050ebb525357e731b83bd6d924587

Observation bb475fb5-907b-4599-8794-87e0c302ca8e · inbound

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches cites this paper.

Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:25:28.228212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:22:25.349398Z digest=sha256:10e1591624efd629a88f71e2c0ba9d7a967270de3fcda018f727300639a94c54

Observation e18507b1-8146-4282-a1eb-620798cc0b66 · inbound

HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice cites this paper.

HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T18:00:01.338538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T23:08:27.995672Z digest=sha256:7347e043b1af8f417c80db6e50661d4234b2a63e939fba6c9caae4a6fc8178fa