Pith. sign in

Paper Citation Record · LEDGER

NoveltyBench: Evaluating Language Models for Humanlike Diversity

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2504.05228.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.05228 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:15:04.585634Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0ed3dc4b-ab78-4ac8-bf22-3eb0be8c7ff9 · inbound

Avoidance Decoding for Diverse Multi-Branch Story Generation cites this paper.

Avoidance Decoding for Diverse Multi-Branch Story Generation NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:55.182171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:56:55.182171Z digest=sha256:47008f4de7949cb458d8777535132729b1903c0e48b15d6cdd41007e7f72868c

Observation 41e0bbe0-176c-4be0-b935-3b4d5c574488 · inbound

Polychromic Objectives for Reinforcement Learning cites this paper.

Polychromic Objectives for Reinforcement Learning NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:19.942452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:6512fd5d0784b5472a768eeeda8477a37c47d26ae800d0d8ad8468134ad39971

Observation b8cfa137-d169-4d5c-ae94-65c69850c69c · inbound

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models cites this paper.

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:16.890777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T20:29:31.354743Z digest=sha256:592b09e1d07e27193520714f6993334cab775294f0bf532bb1e10a8e589b8756

Observation 905e3790-6f3c-4f3f-88a2-5549622611e8 · inbound

Beyond One Output: Visualizing and Comparing Distributions of Language Model Generations cites this paper.

Beyond One Output: Visualizing and Comparing Distributions of Language Model Generations NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:09:57.103042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T04:05:37.447023Z digest=sha256:c682cfe864bc4b29206e93236f5b61ad8adc41ca122900c02dc852b6fbe14a3e

Observation 1e0680c2-bf83-4c14-a7ad-e05373e37085 · inbound

Beyond One Output: Visualizing and Comparing Distributions of Language Model Generations cites this paper.

Beyond One Output: Visualizing and Comparing Distributions of Language Model Generations NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T05:28:01.297854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:28:01.297854Z digest=sha256:444106370c904450d0dda7ee981a6932469fd7fcdbc0543e7a43a99d82943091

Observation 2b10123e-694f-49a5-9aea-e07e89e48ffd · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T20:37:32.132735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-09T20:32:37.788283Z digest=sha256:fbdb955c572f8654ed7129577fa8cec568a003c8e54544d5ccb232269da4f28b

Observation 6ac08124-4c92-4e4c-bfe6-56a8e2779122 · inbound

Diversity in Large Language Models under Supervised Fine-Tuning cites this paper.

Diversity in Large Language Models under Supervised Fine-Tuning NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:11:18.186726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T03:10:22.314719Z digest=sha256:8f9d921be28ad79762fe5fd591be1b689bdbe5d22851681c7e65399bec4833eb

Observation f7225068-720f-42d0-8397-4cc7d60f5805 · inbound

Ex Ante Evaluation of AI-Induced Idea Diversity Collapse cites this paper.

Ex Ante Evaluation of AI-Induced Idea Diversity Collapse NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:09.484270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T09:49:57.122084Z digest=sha256:deb20ea3683e0ba4dd464d2a0f57b7cf07ece898225330ebc0555941c42669f4

Observation 6f5d2ecf-c661-4bff-ba64-02597be182f6 · inbound

Annotations Mitigate Post-Training Mode Collapse cites this paper.

Annotations Mitigate Post-Training Mode Collapse NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:46:51.414237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T03:58:11.179607Z digest=sha256:ae083a8412d0dc8321106358774e2c951a8f9db37ce68c1693bbae515744f67d

Observation 2704aad1-20cd-4a77-b328-56c8806bcb01 · inbound

When to Ask a Question: Understanding Communication Strategies in Generative AI Tools cites this paper.

When to Ask a Question: Understanding Communication Strategies in Generative AI Tools NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:17:02.601609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T01:12:50.314892Z digest=sha256:4d8b676107bcdf46f2e0be61c0edd573ebba6bd7d70021669e4dada86f9f30e1

Observation 50aef6d6-b0ae-472f-99e7-7ec3dcea4d36 · inbound

Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers cites this paper.

Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:12:50.830173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T19:11:07.615149Z digest=sha256:3be2441a9f040990d3aa65b665d0aeae2aaae0cccad4a98d1ee9db9e3ed626b7

Observation a7a8196a-df54-4be0-a7b5-c5cf303262b8 · inbound

"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise cites this paper.

"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.791428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T14:51:06.448678Z digest=sha256:20f8998d6daf34cce7f1fb8bbe35c42dd562d5a18260b158485faf9b335c9822

Observation 02ca02fd-5f87-4ae0-bf3c-53235ec9b36f · inbound

AI as a Tool for Simulation-Based Experiments in Literary Studies cites this paper.

AI as a Tool for Simulation-Based Experiments in Literary Studies NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-06-28T15:02:19.020822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T14:54:14.909995Z digest=sha256:fae2f122804228cf6aa4b7aa6e0d83b2e13d042c682714a35ec4b57039e5d7e3

Observation 01148e74-d189-469f-8dcb-32bf37b814ea · inbound

Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems cites this paper.

Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-28T02:01:28.966868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T01:56:56.127785Z digest=sha256:ef8db39b5384c3c035fa4caa7fb98022f2267675a1a15dcd7f4bf91b49bbd0ff

Observation f05f7d33-f232-4777-a65f-774d4bb87a30 · inbound

Recalling Too Well: Sycophancy Evaluation and Mitigation in Memory-Augmented Models cites this paper.

Recalling Too Well: Sycophancy Evaluation and Mitigation in Memory-Augmented Models NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:57:41.835632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T12:59:17.696736Z digest=sha256:9575c9607e18d9f3cdccad5420f929a9f00cb9c8d363a93e5287fd5671bae707

Observation a8e0c06b-6bce-4c85-abfe-37b259b4dd44 · inbound

AI Coding Agents in Social Science: Methodologically Diverse, Empirically Consistent, Interpretively Vulnerable cites this paper.

AI Coding Agents in Social Science: Methodologically Diverse, Empirically Consistent, Interpretively Vulnerable NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.538163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T13:04:40.640910Z digest=sha256:5d8082de68f64b9c93577a3c0c45568e75078104fdfa22eb26a0c876b2368d09

Observation b5411bf4-8fe2-41b2-9662-c41298b9e127 · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:00:08.478945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:d2a5d3085f5947aaf4b2ec11421792e09fc60b85873ecafeeaf239025ea4c9a1

Observation 3b594db5-f635-4e9f-9f5f-9657aff90f0f · inbound

Improving LLMs via Validator-to-Generator Alignment cites this paper.

Improving LLMs via Validator-to-Generator Alignment NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-12T07:52:00.900202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T07:52:00.900202Z digest=sha256:25876f01c5f69c99fba11473ea5e09f6b0111c17866bb44421a5428f09f2ed60

Observation 936bca94-5995-49a8-9678-86debf1c3588 · inbound

The One-Word Census: Answer-Choice Conformity Across 44 Language Models cites this paper.

The One-Word Census: Answer-Choice Conformity Across 44 Language Models NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T06:23:48.955749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:23:48.955749Z digest=sha256:efc1a6645eace8c0190324add157da6e1929e992bf5cf5eb2c5377defbf715cd

Observation dc70fbbe-a51f-4f8e-8210-1de44999d0fb · inbound

Structured Output Collapses Answer Diversity Across 44 Language Models cites this paper.

Structured Output Collapses Answer Diversity Across 44 Language Models NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T15:23:45.409510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:23:45.409510Z digest=sha256:40da860fa133bac2ec57be4c86054ba51a8036ff240d3dadcebeb46ae3a55b1a

Observation c524adc3-b61f-4a28-ad18-499a4ec70cbd · inbound

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL cites this paper.

One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL NoveltyBench: Evaluating Language Models for Humanlike Diversity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:15:04.585634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:15:04.585634Z digest=sha256:0060ba0bdeb9f0e21df98a0725fbbe03a1552efef978f56e13f0280cb9a04c1a