Pith. sign in

Paper Citation Record · LEDGER

Revisiting LLM Reasoning via Information Bottleneck

As of 18 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 5 inbound Pith citation observations for arXiv:2507.18391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18391 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:21:08.555405Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T14:45:37.426688Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T18:31:44.611173Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2a7529d4-b8f5-4a22-b79f-ba64f9d1f1d2 · outbound

This paper cites Training language models to follow instructions with human feedback.

Revisiting LLM Reasoning via Information Bottleneck Training language models to follow instructions with human feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.320500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.320500Z digest=sha256:b2e2910aee9fbba71ccf0d4ec6c22c2fab664b055f22c899c04cec760e7a54c7

Observation 1611a904-9368-4fa7-8ee1-9674d51be878 · outbound

This paper cites OpenAI o1 System Card.

Revisiting LLM Reasoning via Information Bottleneck OpenAI o1 System Card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.327634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.327634Z digest=sha256:7dabde193c0f851a9c9ce2ee98d1a2e0f1c1a18c92c67ab95428f842c343dcef

Observation 193a10d9-63cd-400c-804b-181e67687cc6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Revisiting LLM Reasoning via Information Bottleneck DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.335428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.335428Z digest=sha256:8ead0cbcb7550cd0caf4ca526700917bfd6c74d9d6ff623704e4269431364734

Observation 0cec214d-853b-41c7-91a3-5b196632a34d · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Revisiting LLM Reasoning via Information Bottleneck The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.342231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.342231Z digest=sha256:878600bf91a835d2c63d27e279d51d24630272635849333a31795f42eb1ee9be

Observation cc130e31-d274-4e06-9d93-05cb6a128d2c · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Revisiting LLM Reasoning via Information Bottleneck Reasoning with Exploration: An Entropy Perspective

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.349102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.349102Z digest=sha256:b5773786284d688681ac6d7d99acc9d0a71a359d456ae03418b99c9f57c4fdc2

Observation 016890d6-cc52-4f80-82b0-13d438ca4d5a · outbound

This paper cites Diversity-aware policy optimization for large language model reasoning.

Revisiting LLM Reasoning via Information Bottleneck Diversity-aware policy optimization for large language model reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.358014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.358014Z digest=sha256:2ad3a95b263f950bccee8e8f5ecac2da7911e19d65e3acbbcbd715216b3d5236

Observation c495eb06-2334-428c-9726-44a73613d688 · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

Revisiting LLM Reasoning via Information Bottleneck The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.364352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.364352Z digest=sha256:652c1cc829624bdc20b09a96522228d12cbce6ac584fce69db51b8ce19ed5705

Observation 48ee6a67-a6b9-4c1a-802f-d87700e28cc0 · outbound

This paper cites One-shot Entropy Minimization.

Revisiting LLM Reasoning via Information Bottleneck One-shot Entropy Minimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.371047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.371047Z digest=sha256:c1b004470dfbe0ce7fe78b61da51a673c7edeb4978ff05162150e2fdfabdfa26

Observation b606fae2-3e58-4654-a66b-1e6442e05a64 · outbound

This paper cites Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models.

Revisiting LLM Reasoning via Information Bottleneck Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.381567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.381567Z digest=sha256:838d6c2d44be6c7647105cd8a09ce7f7e78d9c3cd94c440c81a3a04e5d217537

Observation ef29466a-18b2-4ffd-8c6b-54af4b9f63a0 · outbound

This paper cites The information bottleneck method.

Revisiting LLM Reasoning via Information Bottleneck The information bottleneck method

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.394685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.394685Z digest=sha256:c55fcbaae4a93d8dfae3bb5e4c1a19bcdc0791c460cebd05bfef33b2d65abad9

Observation 52ec288a-67c7-434f-91a2-2adcb23feb4e · outbound

This paper cites Deep learning and the information bottleneck principle.

Revisiting LLM Reasoning via Information Bottleneck Deep learning and the information bottleneck principle

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.406828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.406828Z digest=sha256:df0f4fe23dccba2f68ac2767b2cf88cf98f10f600173e7a053ecc06adc9a63ff

Observation d8a6cbce-758d-40af-ac2a-5dcad0fbbe61 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Revisiting LLM Reasoning via Information Bottleneck Proximal Policy Optimization Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.417886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.417886Z digest=sha256:a090b52bc88a026f7347a3b4d5974dc82a201846d983f610790d690506cd94d7

Observation 8b276f41-8b64-4c5f-a3b0-89984e75a915 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Revisiting LLM Reasoning via Information Bottleneck DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.440400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.440400Z digest=sha256:a6fea904058c023049dd0e26d319ff974dd8d1cc4774e96bc44096b390a1b2b9

Observation 5baeb517-c260-4fb6-8bec-8c4401bb6c21 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Revisiting LLM Reasoning via Information Bottleneck DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.450942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.450942Z digest=sha256:1795f67ad22968ab23ed5385f5576b4ed26e0e9d1e1fa92ca7d2792766103a0c

Observation a45967e8-f1c7-4192-b346-7941246b4e6d · outbound

This paper cites Qwen2.5 Technical Report.

Revisiting LLM Reasoning via Information Bottleneck Qwen2.5 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.458013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.458013Z digest=sha256:ab0ab6d8b5518566c75ea5a8ff5ec3db710caa78f0f764ee5570123a28619c58

Observation 76b44746-b628-4b69-88d0-f6d7f6eeed61 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Revisiting LLM Reasoning via Information Bottleneck Chain-of-thought prompting elicits reasoning in large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.466984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.466984Z digest=sha256:93f85a936dc11fb2aad331bf4039ef859980c51e872b162c501ecb5499b14863

Observation 49e2e3fa-71f4-48ae-a96f-bb2b524efc66 · outbound

This paper cites Alemi, Ian Fischer, Joshua V.

Revisiting LLM Reasoning via Information Bottleneck Alemi, Ian Fischer, Joshua V

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:21:09.440840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:21:08.476481Z digest=sha256:f4c557e53d3fe3e95b120643f93d9ea82a482d82eded625f04e16b6961b57373

Observation 4ba06770-7dd5-46b2-8712-ba72fa0e79db · outbound

This paper cites On the information bottleneck theory of deep learning.

Revisiting LLM Reasoning via Information Bottleneck On the information bottleneck theory of deep learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:21:09.415682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:21:08.485386Z digest=sha256:15eac91dd6ade594b89923ca24cb13f72888cb309b8e5b6bd197ce6777360024

Observation 2ec6101f-588b-4bce-85e7-f5ead91e5556 · outbound

This paper cites How does information bottleneck help deep learning? In International Conference on Machine Learning, pages 16049–16096.

Revisiting LLM Reasoning via Information Bottleneck How does information bottleneck help deep learning? In International Conference on Machine Learning, pages 16049–16096

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.493845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.493845Z digest=sha256:ae15af7bf83bbb96e2035845f0ba35885d62ec624aed0c68c43624d2d8d0c07a

Observation abf63ac7-31b0-40d3-8dcb-83d2b21bed6e · outbound

This paper cites Memorization-compression cycles improve generalization.arXiv preprint arXiv:2505.08727, 2025.

Revisiting LLM Reasoning via Information Bottleneck Memorization-compression cycles improve generalization.arXiv preprint arXiv:2505.08727, 2025

Reference 20

Resolution
verified exact
raw_fallback, observed 2026-08-15T18:21:08.962435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:21:08.504449Z digest=sha256:dbce711e1373166cdc2300529e7ba9c848a93710ca1788d09dc472c6b51762bb

Observation 5a719c3a-e967-4fdf-bd8f-d2007e949da0 · outbound

This paper cites Measures of entropy from data using infinitely divisible kernels.

Revisiting LLM Reasoning via Information Bottleneck Measures of entropy from data using infinitely divisible kernels

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:21:09.373976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:21:08.518704Z digest=sha256:379c721103b302dc090aa87ba22a2f20f0741eb97eeb02c261af16c8d3d56761

Observation 3f857dcd-b5b4-4001-8e1e-5daf622450d4 · outbound

This paper cites Reinforcement learning finetunes small subnetworks in large language models.

Revisiting LLM Reasoning via Information Bottleneck Reinforcement learning finetunes small subnetworks in large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.527633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.527633Z digest=sha256:79c77d272ad61bd5bbe2a643bbd215a1f1ca44fced05e742e1ff1847241d8c40

Observation c66237ac-2488-4d6f-ab4d-b9cedacdffed · outbound

This paper cites Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them.

Revisiting LLM Reasoning via Information Bottleneck Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.532925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.532925Z digest=sha256:155c87781f1326943da3e9e2d84cc5fca59451382978ad81bd14a1338b71d5a9

Observation 6b697f1f-c099-41df-ad6c-c211b91966d0 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Revisiting LLM Reasoning via Information Bottleneck Hybridflow: A flexible and efficient rlhf framework

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.541671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.541671Z digest=sha256:ef65ea8de1b6a0478e3765ca7be0835f985c6b0493f07cb6f208939ed4a5ba58

Observation 97b2afa1-a7c5-4363-961a-7e3d293c05a7 · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

Revisiting LLM Reasoning via Information Bottleneck Skywork Open Reasoner 1 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.547652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.547652Z digest=sha256:3b4a84196e33ec76211872411e7781288cb82f71161a0546c6b527f8396b5d59

Observation fa18ef4b-7416-40ea-a7b3-0d7e23549957 · outbound

This paper cites Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training.

Revisiting LLM Reasoning via Information Bottleneck Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.555405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.555405Z digest=sha256:397d48f5e58a16ed24ffa79cd9d3cc51cb2c7b11efe02fd52b05aa043fabb635

Pith citing papers

Observation 056a1231-6383-4588-9127-0b041157bfd4 · inbound

Self-Aligned Reward: Towards Effective and Efficient Reasoners cites this paper.

Self-Aligned Reward: Towards Effective and Efficient Reasoners Revisiting LLM Reasoning via Information Bottleneck

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.613691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T18:27:23.076544Z digest=sha256:99e634f0f8e858eafcf76c397c3bdf2750de58349a355afbd032bdc792d65d54

Observation 6381c5ad-47e7-4203-855e-b7901e1a5ba8 · inbound

When Less is Enough: Efficient Inference via Collaborative Reasoning cites this paper.

When Less is Enough: Efficient Inference via Collaborative Reasoning Revisiting LLM Reasoning via Information Bottleneck

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:41:42.284856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T19:27:04.267404Z digest=sha256:d3d9ebf9bb9ea652524cc4f7141dd22d984d9e1bb860bb1e3b82c7a658b626bd

Observation 17f774f1-b33b-4207-8472-db459de00a8b · inbound

Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction cites this paper.

Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction Revisiting LLM Reasoning via Information Bottleneck

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T03:48:31.314623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T03:48:31.314623Z digest=sha256:3e022c93e576d4564220b1b2a20670b38cdfd9efc7e85a7db2d377cd9d74d37e

Observation 9874472e-b237-404e-a4d5-143919ad7257 · inbound

Information-Theoretic Limits of Reliability and Scaling in Language Models cites this paper.

Information-Theoretic Limits of Reliability and Scaling in Language Models Revisiting LLM Reasoning via Information Bottleneck

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T14:45:37.426688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:45:37.426688Z digest=sha256:c53f64da5fc0c9d20e1610da659035ec3a68ac34ed2583e876082f23b42dfe89

Observation 2d47ed9d-be70-458f-ac3e-f1b4e783acb8 · inbound

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization cites this paper.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Revisiting LLM Reasoning via Information Bottleneck

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:05.465608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:05.465608Z digest=sha256:c9d74f4c4a081b2177ca359ad67cc3cdf63e321eedefcb4495e4bd889f9476d6