Pith. sign in

Paper Citation Record · LEDGER

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play

As of 9 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2607.19523.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19523 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T12:36:12.843668Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation afce5e30-3810-43fd-b616-0e2742d6513d · outbound

This paper cites GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:10.267086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:10.267086Z digest=sha256:ac37586a72e50511249176c9fb5117b6c768bf480e6f78e8de848132e1b5e524

Observation 0e6af7b0-3c11-4fed-be2a-8c01d6caec69 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:10.939373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:10.939373Z digest=sha256:9ea8f60220a792c0ec2fe463a069dbf8d64c2e785b3b9b96e13a02c4d1e56618

Observation f89f6fef-cf17-449d-b882-b044d01dd1b0 · outbound

This paper cites MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:11.078426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:11.078426Z digest=sha256:885f1fe9891a189da314fdf95c024588856bf4059a9a1543db22144e4cf402e4

Observation c829a6af-1648-4c32-b739-84b78005ebf6 · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:11.245333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:11.245333Z digest=sha256:bf926ad8175bf835f5c1e5dce468d0cb03cbd2bda6a5b9c38405a5da7db2b46c

Observation 16923aae-33af-4e0a-b388-7133ba93204c · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:11.470843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:11.470843Z digest=sha256:1dfc7ea566b06f8fa553c22614b792ed7cd1fb6eaeab0172ae4a01341654fa81

Observation f0ece43e-d7be-4773-9af5-a08a25aaca6a · outbound

This paper cites One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:11.887984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:11.887984Z digest=sha256:87e9662ad35d9aef33685befa126feabe72f538bb38db41ec88835c96fa436c0

Observation fd6c8850-31ea-49cf-8869-08b1b3f52da3 · outbound

This paper cites Do llm agents have regret? a case study in online learning and games.arXiv preprint arXiv:2403.16843,.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Do llm agents have regret? a case study in online learning and games.arXiv preprint arXiv:2403.16843,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:12.113398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:12.113398Z digest=sha256:f80cb5e8d7ab44f091ce38bc918ccf8ed9c96aaf184ec0b7490f27bdafae2cee

Observation 62ae9cd7-9a10-4c82-ba49-9080037c237c · outbound

This paper cites Offline Learning of Controllable Diverse Behaviors.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Offline Learning of Controllable Diverse Behaviors

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:12.177892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:12.177892Z digest=sha256:3adb8fb8a2a48c4be79a989f533765640d094f3fdb1b1aa6ebe548ce29fb6120

Observation 347ec6d5-84a2-47d0-b02c-4a838aa54c0a · outbound

This paper cites Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:12.257522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:12.257522Z digest=sha256:d66e247271f690e7fae36de1bc50a500de995c05ded2fb355c38fefa27679776

Observation 362703cf-7d95-4f44-b663-1f445fe8c7fa · outbound

This paper cites Chessqa: Evaluating large language models for chess understanding.arXiv preprint arXiv:2510.23948,.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Chessqa: Evaluating large language models for chess understanding.arXiv preprint arXiv:2510.23948,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:12.357853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:12.357853Z digest=sha256:a7fd702d158750712f8daa6e6a1edd7150ec73a01aa4a71804d4fad7e143bb8a

Observation 1934517f-cac9-4a64-9ac3-74851e6dce68 · outbound

This paper cites AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:12.420962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:12.420962Z digest=sha256:7db67e1a1cdd7eda7e9d75f8e695b2d26d67c51ca1927557efa265ff40f885f5

Observation f5e532c6-4ab3-43d1-9178-076df8ed43f8 · outbound

This paper cites Qwen3 Technical Report.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Qwen3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:12.503627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:12.503627Z digest=sha256:2ec4a5dc6abbffbc7350a789bb69b01a000d3e03b97efc8e31e9dbf4f06dd38c

Observation 843f753f-d335-4f27-b7c2-eeeffabd27cc · outbound

This paper cites How to Leverage Diverse Demonstrations in Offline Imitation Learning.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play How to Leverage Diverse Demonstrations in Offline Imitation Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:12.597503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:12.597503Z digest=sha256:afba1b18f047dba981f97b2121ee7d50d3945617cfbf2aa97dab67fceff15c3e

Observation 546ef691-ef62-4252-9512-6728fa6add49 · outbound

This paper cites The Price of Format: Diversity Collapse in LLMs.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play The Price of Format: Diversity Collapse in LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:12.681644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:12.681644Z digest=sha256:55e33ccab80ef189ce7ac16b75566ea3b8d10faf5d29f2cc7a6ae813a93178ab

Observation dd8517d9-118c-4693-b629-f7ae8f4bfd2a · outbound

This paper cites Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:12.755535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:12.755535Z digest=sha256:4839f02260e003bbd97f0dc774230669bc709cd3de20e992e5217609fdfdaa93

Observation 2e02c846-5d54-4d8e-a633-59e6e566deef · outbound

This paper cites move": <action_label>}</action> The JSON key must be.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play move": <action_label>}</action> The JSON key must be

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:12.843668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:12.843668Z digest=sha256:43b01f39a8fe5dfdc548db5f66118d50e751ebc77aefe094dd95a5856b8119e6

Observation e8cb85e9-1b8f-4505-bafe-81e03d552cd5 · outbound

This paper cites RvS: What is Essential for Offline RL via Supervised Learning?.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play RvS: What is Essential for Offline RL via Supervised Learning?

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:10.784085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:10.784085Z digest=sha256:44298fefb90245e1823b2d701e821c532f5c13d80b30724b278be0db6315eee8

Observation f85fdf8e-bbee-47d3-9028-4bcd64ee1471 · outbound

This paper cites Preserving Diversity in Supervised Fine-Tuning of Large Language Models.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Preserving Diversity in Supervised Fine-Tuning of Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:11.617634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:11.617634Z digest=sha256:b9806b7f32b389b4dbc95104fb01ee3fb3cf94eb44b1823d48af7e8a708a2829

Observation d3c6f4a9-5b75-4101-be29-bcf99b26ebca · outbound

This paper cites Sed-sft: Selectively encouraging diversity in supervised fine-tuning.arXiv preprint arXiv:2602.07464,.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Sed-sft: Selectively encouraging diversity in supervised fine-tuning.arXiv preprint arXiv:2602.07464,

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:10.487813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:10.487813Z digest=sha256:e1fe79eef4bc85a6223fe3fdc07330b978bb39570bda5dd57c6cfab1702b8fc3

Observation 57a024b7-7231-4e20-8eba-43a32a9c4749 · outbound

This paper cites Attributing mode collapse in the fine-tuning of large language models.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Attributing mode collapse in the fine-tuning of large language models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:12.023276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:12.023276Z digest=sha256:cc5a02e71c293aed896c3b0b00ecc77da8642aeb507472798064824e2e4b381e

Observation cc0bde1d-6aab-476d-853a-41364ac8003d · outbound

This paper cites Llm chess: Benchmarking reasoning and instruction- following in llms through chess.arXiv preprint arXiv:2512.01992,.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Llm chess: Benchmarking reasoning and instruction- following in llms through chess.arXiv preprint arXiv:2512.01992,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:11.356713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:11.356713Z digest=sha256:710b3d12417f2deb9957003835a9106f4492f265dc9affe29ced9a42e34b1385

Observation 8d203cf8-07a4-4a66-a48a-7c24f94d052e · outbound

This paper cites TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:11.742193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:11.742193Z digest=sha256:88b64926d7998d9e13394cca4a6e762c7b9880cea260fc06d7cabf84fab9025b

Observation e6b0c72c-0b3c-4188-9c25-0e11b169e95f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:10.373328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:10.373328Z digest=sha256:b42bd23d4f65bb45b5f70aa1b0aaaea3e9b927dfe942662416bfcf49eb7e8cc9

Observation 3c5fc3b4-ae13-44a5-94e6-d153a273df12 · outbound

This paper cites Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:10.643992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:10.643992Z digest=sha256:428490b46f231b6535b476b3d0136e9c764ab279971f33369e953c83224b0309

Pith citing papers

No inbound Pith citation observations are available.