Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:09:05.688277Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 10 inbound Pith citation observations for arXiv:2510.11686.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:09:05.688277Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:39:40.961086Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T12:16:56.948632Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b79db884-03c4-410c-a06e-c586ff7d0c59 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f85980d2-63eb-46cc-87e6-4a9ca443aefc · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Program Synthesis with Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b02aeed5-5e6d-4ec3-a9fd-2532a5403d52 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Online Preference Alignment for Language Models via Count-based Exploration
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 463e9e3b-2338-4e17-98df-71867a3154c2 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training InfAlign: Inference-aware language model alignment
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e65d4f4-0fa9-4cf3-969d-4b8fad1fc30a · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 926cde3f-18cb-45a0-b833-92284fb1fcdb · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training PAD: Personalized Alignment of LLMs at Decoding-Time
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fa1d1e6-d42d-4e58-8ba4-68cd89d5d387 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Enhancing diversity in large language models via determinantal point processes.arXiv preprint arXiv:2509.04784, 2025a
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 232a33d4-98a9-4513-b3f7-66e92f757705 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Uniform sampling for matrix approximation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ce1eab0-aff1-43ad-9c2c-3bb3c6e43ded · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Weight ensembling improves reasoning in language models.arXiv:2504.10478,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb676e2d-2c07-4cb0-ad95-199e248ce63b · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f146bee1-87d4-47a7-bdbf-e45a658605f7 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Navigate the unknown: Enhancing llm reasoning with intrinsic motivation guided exploration.arXiv preprint arXiv:2505.17621,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97269dc6-35f8-4512-a8c4-85965883311d · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Large-Scale Data Selection for Instruction Tuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90838d63-31a2-46c7-8a5a-8ebd41642d03 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80061e44-3d5b-4f76-975c-ae0241157149 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training ARGS: Alignment as Reward-Guided Search
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 196d27bc-b522-4d0f-92b2-79ce5f46be03 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Diverse Preference Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65341c38-834d-40f2-af35-ac56ad9d22fb · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Jointly Reinforcing Diversity and Quality in Language Model Generations
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 490f8c85-57a1-43fb-88ca-1c61f0020d16 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Let's Verify Step by Step
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8768ed37-c50e-4366-b06c-99eaf0815fb1 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 136895b3-952e-41b6-8241-e64f54cc20fb · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Decoding-time Realignment of Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeada442-4a94-4af3-ab47-4efc1031ee6a · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Proximal Policy Optimization Algorithms
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a1e0385-5950-46db-a390-5826abdc1213 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f4bd61a-8b01-435a-a2d8-3e36ab7d83bc · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d08d0480-1de6-469f-b032-079c34fa1c10 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training HybridFlow: A Flexible and Efficient RLHF Framework
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a313990a-e2cf-4fe1-8386-e1d359a077c2 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Decoding-Time Language Model Alignment with Multiple Objectives
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 504fac42-bdc4-4eb3-97e4-d44016c45cf7 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training The invisible leash: Why rlvr may not escape its origin.arXiv preprint arXiv:2507.14843,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee320082-c0ca-4daf-9bb1-c229ce6aa84f · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ec0e46e-639f-4459-a353-d9fc20ce615e · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Formalizing Learning from Language Feedback with Provable Guarantees
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb2b05b5-f625-4fd5-bee9-2c5abec4ff16 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 671555f6-4b85-40d5-8ae0-234b84091bd5 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d97fbf4-922a-42c9-976d-88763ceff4fa · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f00e0298-0149-4177-88b1-90d28a37585b · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Expo: Unlocking hard reasoning with self-explanation- guided reinforcement learning.arXiv preprint arXiv:2507.02834,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d292a63-4aa6-40c5-80df-e91d9fc64065 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Most closely related to our work, Setlur et al
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 265fcef4-fa43-420c-ad84-10b4275e9c7c · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training breadth" (batch size) and “depth
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fce4406-4920-4fe3-bdc2-5af4915478fc · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Since the dataset does not come with any train or test splits, we use the full set of questions for our experiments
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db3ece86-d22c-42a7-b5f9-ed972f630f41 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training 1− n−c k n k # = 1 |D| |D|X i=1
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d833705-21d9-4738-a7c2-5b930c08a1ac · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 739a1e3e-cd08-40e0-9283-e7ea6a0f8a89 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7810da19-7d76-4ce5-99b0-80de902e2f49 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training 5See also Arumugam and Griffiths (2025), which uses a pre-trained model to simulate posterior sampling in-context for multi-turn sequential decision making tasks
Reference 512
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c4b6958-efad-43b3-a733-470f9bcc0a3c · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
Reference 1992
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7093e53f-8790-4494-adc5-f70b9f0c8b2f · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training The Llama 3 Herd of Models
Reference 2006
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69fd3b8d-2aa5-4606-b93a-117f4365fd7e · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Training Verifiers to Solve Math Word Problems
Reference 2010
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eb5adf9-106e-4306-9550-230e97de115d · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Phi-4 Technical Report
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f0c4e45-d421-4efe-ba34-a2461a8ed359 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8162654f-9f9b-44bc-8130-bdfa01ddf0bb · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Evaluating Large Language Models Trained on Code
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8c1a04d-130e-402d-a097-0545c2cb0b30 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 715b425a-62a1-453e-ba41-0ef7d61e0ce5 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Toward efficient exploration by large language model agents.arXiv preprint arXiv:2504.20997,
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ca3af1-8764-442a-9128-d844f7ebc685 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Mistral 7B
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01876772-9818-43a0-bc5d-1d36c572f2cd · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Exploration by Random Network Distillation
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8299430-4a49-47a9-a262-4b91bc6e1875 · outbound
Representation-Based Exploration for Language Models: From Test-Time to Post-Training Anti-Concentrated Confidence Bonuses for Scalable Exploration
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 060219b5-f3e1-4f7d-bc63-23d6c1e10395 · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7165350c-6e44-4731-a170-221bb6e1aab4 · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 83ffcf47-4606-4df4-9e25-4bf00c095336 · inbound
The Role of Generator Access in Autoregressive Post-Training Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0a6e9232-a3b5-4a36-8e5d-239df60335a2 · inbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dfdd086b-4452-41ad-96de-66c7666fd2de · inbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3f1deeeb-f615-4c04-ab0f-469b491d570d · inbound
The tractability landscape of diffusion alignment: regularization, rewards, and computational primitives Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a4c882a5-141e-4c30-8a1d-7f1111eae931 · inbound
On Advantage Estimates for Max@K Policy Gradients Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c6196c4d-4b05-4577-9504-dfebd55d8b57 · inbound
OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8554cc91-1b8d-49f8-8889-5cdc59e11cf5 · inbound
Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58cd48d0-5f05-41ba-bab0-696d9cec0d80 · inbound
Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Reference 150
Source-reported events for the cited work
Unavailable: canonical work link unavailable.