Pith. sign in

Paper Citation Record · LEDGER

Representation-Based Exploration for Language Models: From Test-Time to Post-Training

As of 13 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 10 inbound Pith citation observations for arXiv:2510.11686.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.11686 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:09:05.688277Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:39:40.961086Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:16:56.948632Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b79db884-03c4-410c-a06e-c586ff7d0c59 · outbound

This paper cites an unresolved cited work.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:05.406788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:05.406788Z digest=sha256:694f1850b0733da9ba79b00e8527eaf938020950dad1d27f0cb84e7772ae97c9

Observation f85980d2-63eb-46cc-87e6-4a9ca443aefc · outbound

This paper cites Program Synthesis with Large Language Models.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Program Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.109898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.109898Z digest=sha256:7c78c1baea10d8d6ecf8bc4d5cee37b0bd47c54a3ebc13486f1377b102f150ab

Observation b02aeed5-5e6d-4ec3-a9fd-2532a5403d52 · outbound

This paper cites Online Preference Alignment for Language Models via Count-based Exploration.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Online Preference Alignment for Language Models via Count-based Exploration

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.222940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.222940Z digest=sha256:00fd0e1c604b68572bc703243df3b31686110174886d305f7d984e1442c07e47

Observation 463e9e3b-2338-4e17-98df-71867a3154c2 · outbound

This paper cites InfAlign: Inference-aware language model alignment.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training InfAlign: Inference-aware language model alignment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.325432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.325432Z digest=sha256:fbde9e6adb8ff926dc7c96da245a8266b92dab230497b39607e4dbae39587d06

Observation 4e65d4f4-0fa9-4cf3-969d-4b8fad1fc30a · outbound

This paper cites an unresolved cited work.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:05.543958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:05.543958Z digest=sha256:39f190e25c34b5698fe758ee425cd54b5aeb6637d250e9666a1fd06414914769

Observation 926cde3f-18cb-45a0-b833-92284fb1fcdb · outbound

This paper cites PAD: Personalized Alignment of LLMs at Decoding-Time.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training PAD: Personalized Alignment of LLMs at Decoding-Time

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.528178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.528178Z digest=sha256:f89e2006a035be408d35ea05b28768d77742e40c4b2042945b9caea7c652780f

Observation 4fa1d1e6-d42d-4e58-8ba4-68cd89d5d387 · outbound

This paper cites Enhancing diversity in large language models via determinantal point processes.arXiv preprint arXiv:2509.04784, 2025a.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Enhancing diversity in large language models via determinantal point processes.arXiv preprint arXiv:2509.04784, 2025a

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.586954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.586954Z digest=sha256:a4aba91fe0c46e9979285ee36f49753f5bda9cc266f3240d090b6f07da8cbc7f

Observation 232a33d4-98a9-4513-b3f7-66e92f757705 · outbound

This paper cites Uniform sampling for matrix approximation.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Uniform sampling for matrix approximation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.701918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.701918Z digest=sha256:780c7cc756c58e3797ccc5abc3fad23a9ca6f1fb63f334d61e9acebdef376fca

Observation 8ce1eab0-aff1-43ad-9c2c-3bb3c6e43ded · outbound

This paper cites Weight ensembling improves reasoning in language models.arXiv:2504.10478,.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Weight ensembling improves reasoning in language models.arXiv:2504.10478,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.753216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.753216Z digest=sha256:f70c31937f59578ec217da142519bc85e5f40212dba6927655f6fdd222f2536b

Observation cb676e2d-2c07-4cb0-ad95-199e248ce63b · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.866606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.866606Z digest=sha256:1f22b3a777cf05feefe2514ae1407d7f2f18d1cd5209bff88037dc0460c542d5

Observation f146bee1-87d4-47a7-bdbf-e45a658605f7 · outbound

This paper cites Navigate the unknown: Enhancing llm reasoning with intrinsic motivation guided exploration.arXiv preprint arXiv:2505.17621,.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Navigate the unknown: Enhancing llm reasoning with intrinsic motivation guided exploration.arXiv preprint arXiv:2505.17621,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.936913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.936913Z digest=sha256:50a4950ff5077c1183c587edbea8f68760d08b29fc6c7b946ce6f710d8410065

Observation 97269dc6-35f8-4512-a8c4-85965883311d · outbound

This paper cites Large-Scale Data Selection for Instruction Tuning.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Large-Scale Data Selection for Instruction Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:02.113206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:02.113206Z digest=sha256:dc2e4955e1f0765663aa25d30b4f62b2f69e94654ec06bf2391161ec08d81abd

Observation 90838d63-31a2-46c7-8a5a-8ebd41642d03 · outbound

This paper cites Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:02.446650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:02.446650Z digest=sha256:861a759656e766b1b2a2f6981b9e13767fd0266151dc81971a5fe2b05e0bba8a

Observation 80061e44-3d5b-4f76-975c-ae0241157149 · outbound

This paper cites ARGS: Alignment as Reward-Guided Search.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training ARGS: Alignment as Reward-Guided Search

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:02.535591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:02.535591Z digest=sha256:46cc30409799d20532acc12b718c9b1b64d2323c3811348022b6fd32ef4fff68

Observation 196d27bc-b522-4d0f-92b2-79ce5f46be03 · outbound

This paper cites Diverse Preference Optimization.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Diverse Preference Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:02.661430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:02.661430Z digest=sha256:110e2a15698023350fbc6cac71dccd3c6bfb1684f4b642fd009027f95c431424

Observation 65341c38-834d-40f2-af35-ac56ad9d22fb · outbound

This paper cites Jointly Reinforcing Diversity and Quality in Language Model Generations.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:02.784967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:02.784967Z digest=sha256:2c1eea3b518c0c43c1ead55313976c0f50c3d99e0bc81eedaf91d88aa140f397

Observation 490f8c85-57a1-43fb-88ca-1c61f0020d16 · outbound

This paper cites Let's Verify Step by Step.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Let's Verify Step by Step

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:02.858946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:02.858946Z digest=sha256:fe1de409ceb2ccb387d9824857ec9ae27c109c265c47634219267864b66ef9e7

Observation 8768ed37-c50e-4366-b06c-99eaf0815fb1 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:02.979065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:02.979065Z digest=sha256:0f66e16a453170a2766941e015d3fea15dba574fcbd257cccc3f2b449dc51cdc

Observation 136895b3-952e-41b6-8241-e64f54cc20fb · outbound

This paper cites Decoding-time Realignment of Language Models.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Decoding-time Realignment of Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:03.075477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:03.075477Z digest=sha256:1e4f735b8ff5cf1d8f9c0017740123f14e85f107edb3527eceb649e733fa05c4

Observation aeada442-4a94-4af3-ab47-4efc1031ee6a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Proximal Policy Optimization Algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:03.215735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:03.215735Z digest=sha256:12e11ce4e8279f709e00aa726b2e1b07745494cbd374564cfe6ea672c55c1dfb

Observation 6a1e0385-5950-46db-a390-5826abdc1213 · outbound

This paper cites e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:03.509893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:03.509893Z digest=sha256:089ae512556a3baf213d5fac53c14d05c3971dbb15b071d327972663c7a75981

Observation 0f4bd61a-8b01-435a-a2d8-3e36ab7d83bc · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:03.708863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:03.708863Z digest=sha256:8b19bf4abd9774c0f1629359e5476fdaaaeef749c5d9b95fc3c76ed56c34e156

Observation d08d0480-1de6-469f-b032-079c34fa1c10 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training HybridFlow: A Flexible and Efficient RLHF Framework

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:03.780031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:03.780031Z digest=sha256:f98258fef18f08a10acdf6728cdd7538e8f1774aff0310cce74be09f6b256681

Observation a313990a-e2cf-4fe1-8386-e1d359a077c2 · outbound

This paper cites Decoding-Time Language Model Alignment with Multiple Objectives.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Decoding-Time Language Model Alignment with Multiple Objectives

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:03.873616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:03.873616Z digest=sha256:45bc43901924e139fd8dc38df2a19559dcaf88d58a7870adebf3bbc2c1441151

Observation 504fac42-bdc4-4eb3-97e4-d44016c45cf7 · outbound

This paper cites The invisible leash: Why rlvr may not escape its origin.arXiv preprint arXiv:2507.14843,.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training The invisible leash: Why rlvr may not escape its origin.arXiv preprint arXiv:2507.14843,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:03.992324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:03.992324Z digest=sha256:424dace3250638ac0d94bbd50e94dbf661add6b0a8bea34aa853631427e91b70

Observation ee320082-c0ca-4daf-9bb1-c229ce6aa84f · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:04.061180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:04.061180Z digest=sha256:1c99e4c7390df566369f12e95ae73f2eed597f58ede3de2a6de5032af0343880

Observation 9ec0e46e-639f-4459-a353-d9fc20ce615e · outbound

This paper cites Formalizing Learning from Language Feedback with Provable Guarantees.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Formalizing Learning from Language Feedback with Provable Guarantees

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:04.124049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:04.124049Z digest=sha256:35abaeffac65b115d90eef26b6067a41865869dc17a9db8927a64f913b31503d

Observation bb2b05b5-f625-4fd5-bee9-2c5abec4ff16 · outbound

This paper cites Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:04.221022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:04.221022Z digest=sha256:820e379d60cecaf5d08ba28bb1a60a1d68742ea8ef4ee5089e1c4e53fa0e9e4b

Observation 671555f6-4b85-40d5-8ae0-234b84091bd5 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:04.354294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:04.354294Z digest=sha256:3ed68a42f8a213c68b2612b5bea0218e6cde70649454d2521dbd782ab0c5959a

Observation 6d97fbf4-922a-42c9-976d-88763ceff4fa · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:04.463137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:04.463137Z digest=sha256:ac57767b05fe009ea3a04dcd3781be66c3f577f6ec50b975aa271ddb8e8a724f

Observation f00e0298-0149-4177-88b1-90d28a37585b · outbound

This paper cites Expo: Unlocking hard reasoning with self-explanation- guided reinforcement learning.arXiv preprint arXiv:2507.02834,.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Expo: Unlocking hard reasoning with self-explanation- guided reinforcement learning.arXiv preprint arXiv:2507.02834,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:04.603388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:04.603388Z digest=sha256:fcee91f9446190d1470bf2f2944e56dc17f2c89a38e2c6ed3b23af3148512bd5

Observation 5d292a63-4aa6-40c5-80df-e91d9fc64065 · outbound

This paper cites Most closely related to our work, Setlur et al.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Most closely related to our work, Setlur et al

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:04.694700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:04.694700Z digest=sha256:d31dfeef8b15833458501c0375191e0ac4e472f112957eb06b3194811950ffd4

Observation 265fcef4-fa43-420c-ad84-10b4275e9c7c · outbound

This paper cites breadth" (batch size) and “depth.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training breadth" (batch size) and “depth

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:04.846966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:04.846966Z digest=sha256:1cd484a7a758be4b2a835cc034eef585d2708d5fe9045435bc80dd20b3080c94

Observation 1fce4406-4920-4fe3-bdc2-5af4915478fc · outbound

This paper cites Since the dataset does not come with any train or test splits, we use the full set of questions for our experiments.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Since the dataset does not come with any train or test splits, we use the full set of questions for our experiments

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:05.091913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:05.091913Z digest=sha256:2ef588bc59eaeb37f684a9c33633ab78f1634a468499903b7b043dc83ae96a5b

Observation db3ece86-d22c-42a7-b5f9-ed972f630f41 · outbound

This paper cites 1− n−c k n k # = 1 |D| |D|X i=1.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training 1− n−c k n k # = 1 |D| |D|X i=1

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:05.184426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:05.184426Z digest=sha256:fd9d36dc2b927684c6f1a80be153268cfa50e505be7181cbbea55f0ab2e6501f

Observation 4d833705-21d9-4738-a7c2-5b930c08a1ac · outbound

This paper cites an unresolved cited work.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:05.306218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:05.306218Z digest=sha256:2f2ab1a91fdcbb1b1e71b502ee1c4633a30d8605116dc28198c987cb1fc8abbc

Observation 739a1e3e-cd08-40e0-9283-e7ea6a0f8a89 · outbound

This paper cites an unresolved cited work.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:05.688277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:05.688277Z digest=sha256:97676b10acabc792d16d23598a5134b4688c3ffacafc4b72a97c2530e8ba89fe

Observation 7810da19-7d76-4ce5-99b0-80de902e2f49 · outbound

This paper cites 5See also Arumugam and Griffiths (2025), which uses a pre-trained model to simulate posterior sampling in-context for multi-turn sequential decision making tasks.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training 5See also Arumugam and Griffiths (2025), which uses a pre-trained model to simulate posterior sampling in-context for multi-turn sequential decision making tasks

Reference 512

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:04.980774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:04.980774Z digest=sha256:27ac04bfaa2101d6383898b385b7dadd80ea5fd6761dffc5ab4eb64ac8f1f1a8

Observation 9c4b6958-efad-43b3-a733-470f9bcc0a3c · outbound

This paper cites Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:03.938953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:03.938953Z digest=sha256:842505d2b5daca133f1ae5b7d88f0f76d19d6114b2952a75aad9f8a1a5a0c82f

Observation 7093e53f-8790-4494-adc5-f70b9f0c8b2f · outbound

This paper cites The Llama 3 Herd of Models.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training The Llama 3 Herd of Models

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.816289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.816289Z digest=sha256:b50bf0a129d2b899624df1faa07bbbe197e26abba2537314e563f6350a43e4f7

Observation 69fd3b8d-2aa5-4606-b93a-117f4365fd7e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Training Verifiers to Solve Math Word Problems

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.638862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.638862Z digest=sha256:318d30fe0107424efef3634357d283c8fad2cd8106f6425192f8640bf83f2b5b

Observation 0eb5adf9-106e-4306-9550-230e97de115d · outbound

This paper cites Phi-4 Technical Report.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Phi-4 Technical Report

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:00.796356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:00.796356Z digest=sha256:f991dcf6fb2c63bbc35131948a712d6743c3d676572c1cd1946fb9f7e24c7872

Observation 0f0c4e45-d421-4efe-ba34-a2461a8ed359 · outbound

This paper cites Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:02.013031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:02.013031Z digest=sha256:7e91c5d15ee01413e9600de9d0c957e7c6bef4fbc9cf07885db44fddd39a5ad9

Observation 8162654f-9f9b-44bc-8130-bdfa01ddf0bb · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Evaluating Large Language Models Trained on Code

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.459764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.459764Z digest=sha256:d1da6d1932e5c8358ec5861e02c8073235ce26c0f924aa38eac11c91783d57ce

Observation e8c1a04d-130e-402d-a097-0545c2cb0b30 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:03.372158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:03.372158Z digest=sha256:a6c76ffdb24d67ca80b50f4938c9847666e552e310974ee138fb80b5c2c5f46d

Observation 715b425a-62a1-453e-ba41-0ef7d61e0ce5 · outbound

This paper cites Toward efficient exploration by large language model agents.arXiv preprint arXiv:2504.20997,.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Toward efficient exploration by large language model agents.arXiv preprint arXiv:2504.20997,

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:00.918533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:00.918533Z digest=sha256:54c815a8dce713d4b8b4b5e59d437437a0976e42f623187e72aa3e8397a17164

Observation 23ca3af1-8764-442a-9128-d844f7ebc685 · outbound

This paper cites Mistral 7B.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Mistral 7B

Reference 2023

Resolution
malformed identifier
no resolver link, observed 2026-08-04T10:09:02.327154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:02.327154Z digest=sha256:5cc096e43aa456e7969ba839777d48527c3217bb8e0dfd5a7a9e65a70724c9e0

Observation 01876772-9818-43a0-bc5d-1d36c572f2cd · outbound

This paper cites Exploration by Random Network Distillation.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Exploration by Random Network Distillation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.404089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.404089Z digest=sha256:f30198f42b24e063e8db4671a77bf075d1263c22c43fa5e785ee794e70148fa8

Observation a8299430-4a49-47a9-a262-4b91bc6e1875 · outbound

This paper cites Anti-Concentrated Confidence Bonuses for Scalable Exploration.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Anti-Concentrated Confidence Bonuses for Scalable Exploration

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:01.013036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:01.013036Z digest=sha256:51adb302da131c6167fbef3056c6527737298a779964719e934f30912e7eedfc

Pith citing papers

Observation 060219b5-f3e1-4f7d-bc63-23d6c1e10395 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Representation-Based Exploration for Language Models: From Test-Time to Post-Training

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-16T02:22:28.823656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:db035d648cee106124593499de804ea1a66e5ffff8486a14c12134fb8b1ea6a9

Observation 7165350c-6e44-4731-a170-221bb6e1aab4 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Representation-Based Exploration for Language Models: From Test-Time to Post-Training

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-16T02:22:28.823656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:7c52f99ad84bdd2b452c46a06f11ac20937bc33257f24a48a61ae980551b6273

Observation 83ffcf47-4606-4df4-9e25-4bf00c095336 · inbound

The Role of Generator Access in Autoregressive Post-Training cites this paper.

The Role of Generator Access in Autoregressive Post-Training Representation-Based Exploration for Language Models: From Test-Time to Post-Training

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-16T02:22:28.823656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T20:16:46.370831Z digest=sha256:32cb8b2e37156a44ed8bea3c897bbb12a7ad38e5f1673ddba725449b2ca9b3b8

Observation 0a6e9232-a3b5-4a36-8e5d-239df60335a2 · inbound

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback cites this paper.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Representation-Based Exploration for Language Models: From Test-Time to Post-Training

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-07-16T02:22:28.823656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-08T17:01:04.571087Z digest=sha256:6fddeeb13911048ba7b0c532a0b5906e72663ab749a091cf724e571b61738fa2

Observation dfdd086b-4452-41ad-96de-66c7666fd2de · inbound

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback cites this paper.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Representation-Based Exploration for Language Models: From Test-Time to Post-Training

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-16T02:22:28.823656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:2e0bd5b230ea688db82623d83a60ea4407589541da1604e723198475fa084ede

Observation 3f1deeeb-f615-4c04-ab0f-469b491d570d · inbound

The tractability landscape of diffusion alignment: regularization, rewards, and computational primitives cites this paper.

The tractability landscape of diffusion alignment: regularization, rewards, and computational primitives Representation-Based Exploration for Language Models: From Test-Time to Post-Training

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-16T02:22:28.823656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T02:49:53.547490Z digest=sha256:fd16e793ad53020d4400d3d5c42fef841b028167e2985af06a3cf635e7af7ac8

Observation a4c882a5-141e-4c30-8a1d-7f1111eae931 · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients Representation-Based Exploration for Language Models: From Test-Time to Post-Training

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-16T02:22:28.823656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:928e989b7a66a78b542ebe7aa77a99da16a27304e7144a820c30262895cb62c5

Observation c6196c4d-4b05-4577-9504-dfebd55d8b57 · inbound

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation cites this paper.

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation Representation-Based Exploration for Language Models: From Test-Time to Post-Training

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-07-16T02:22:28.823656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T02:17:30.974692Z digest=sha256:035e2e32f5c0e9e2cdf9ad9df2276a532151c0a28941a94b9cf32b5ab6eb54b6

Observation 8554cc91-1b8d-49f8-8889-5cdc59e11cf5 · inbound

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models cites this paper.

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models Representation-Based Exploration for Language Models: From Test-Time to Post-Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T23:43:07.011210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T23:43:07.011210Z digest=sha256:552046f56854cc0ccb8762c85bc948f0f9b98e2562531ea1a158f98d20bc4fc8

Observation 58cd48d0-5f05-41ba-bab0-696d9cec0d80 · inbound

Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models cites this paper.

Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models Representation-Based Exploration for Language Models: From Test-Time to Post-Training

Reference 150

Resolution
unresolved
no resolver link, observed 2026-08-12T00:39:40.961086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:39:40.961086Z digest=sha256:30df0254c360691612218b51bc84fc1ace15d396d1eb7f7bff717f4c09edcedc