Pith. sign in

Paper Citation Record · LEDGER

Self-Improvement in Language Models: The Sharpening Mechanism

As of 22 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 18 inbound Pith citation observations for arXiv:2412.01951.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01951 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:11:14.891530Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:18:13.990974Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:08:21.732193Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1fb15654-2f89-4084-b60d-78fbafbc5140 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Self-Improvement in Language Models: The Sharpening Mechanism Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.604915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.604915Z digest=sha256:c0cf8b6f390076442647b6b64f7bee4e6c01351592e6c976f9439acb96471b27

Observation 14864a9c-3e66-4e79-9a7e-946e425a786e · outbound

This paper cites an unresolved cited work.

Self-Improvement in Language Models: The Sharpening Mechanism Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:11:15.675684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:11:14.855033Z digest=sha256:b33d76366011c5f6b8917771dc5435dc97bd1b4877f279aadc0862e4af5b2309

Observation 27f039b6-6ea1-4fda-9230-a402e54bcd7f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Self-Improvement in Language Models: The Sharpening Mechanism Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.623071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.623071Z digest=sha256:6df78d96ffb84a9d58864c26b1d60bf83b7e7652aad9b417637098ee978fa347

Observation 7f531f67-c03e-433f-bd4b-19dcb247b78a · outbound

This paper cites J.1.2 Proof of Theorem 4.2′ Proof of Theorem 4.2′.

Self-Improvement in Language Models: The Sharpening Mechanism J.1.2 Proof of Theorem 4.2′ Proof of Theorem 4.2′

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.609379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:11:14.879283Z digest=sha256:4c318652d7468122db6ba6371cfe9b7631b9a1262fcd3a51bef8ea37a4ea0678

Observation f9750400-8623-4dcb-8275-b44a276f1434 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Self-Improvement in Language Models: The Sharpening Mechanism Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.635630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.635630Z digest=sha256:49b53b30b4cd0e802a5a92aceb02a6e48b16511c4b97fcd075961a5cd9707aaf

Observation a1a19f5b-306a-4106-87a7-90cc46f728b3 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Self-Improvement in Language Models: The Sharpening Mechanism BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.652665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.652665Z digest=sha256:f43705273366c53a6163fc792228227b60f5065c8da4a282eae4ce99b8b3bcec

Observation de0cfaa2-1a94-4681-9f56-ce9a6e749911 · outbound

This paper cites Foundations of Reinforcement Learning and Interactive Decision Making.

Self-Improvement in Language Models: The Sharpening Mechanism Foundations of Reinforcement Learning and Interactive Decision Making

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.664926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.664926Z digest=sha256:9d4ae17ac592f367153c35dca8f1fdb70bbb974efb18fa8b3331fd46bef39bbe

Observation 9d41769f-0976-44c9-8b0e-c04e47963bfe · outbound

This paper cites The Statistical Complexity of Interactive Decision Making.

Self-Improvement in Language Models: The Sharpening Mechanism The Statistical Complexity of Interactive Decision Making

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.668664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.668664Z digest=sha256:7ed4745291ae173c83264ff09ab03ac9c0fede7cacb26d6bec513b79f04f864e

Observation 334e2da3-8166-41b8-9891-3350675f0b0e · outbound

This paper cites REBEL: Reinforcement Learning via Regressing Relative Rewards.

Self-Improvement in Language Models: The Sharpening Mechanism REBEL: Reinforcement Learning via Regressing Relative Rewards

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.672865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.672865Z digest=sha256:1857e513e1171eb32a8f4ea5925bb6aa5df15a2f9c3f9b81faf32e376373721f

Observation 5d599889-188e-48a1-b903-6776752e606f · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Self-Improvement in Language Models: The Sharpening Mechanism Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.686228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.686228Z digest=sha256:d0c8a557382b13f7d2411aaedcfb388fe0445812983777009432fd2217717191

Observation ecca62f2-d97e-4dc6-a1f0-813308f74f10 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Self-Improvement in Language Models: The Sharpening Mechanism Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.690922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.690922Z digest=sha256:13404d8cdad516eb859aab2832d1135d997f70faa3e62165699fdfd7982a94aa

Observation f99de5e0-db70-491d-a30c-3b37a29e4a47 · outbound

This paper cites Amortizing intractable inference in large language models.

Self-Improvement in Language Models: The Sharpening Mechanism Amortizing intractable inference in large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.703729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.703729Z digest=sha256:800d4802227903b6a07b2f5613a952e914c591d2c18231002e138cef4f8eae2b

Observation 4f835749-1798-421a-b54d-e5b277411b0a · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Self-Improvement in Language Models: The Sharpening Mechanism Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.707926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.707926Z digest=sha256:cad7a1d386b905f8894439e27ddfb506cabc44148c1ace6711f0286f22ec517e

Observation ee9f7080-a4e1-43f1-8178-68a2f29df2df · outbound

This paper cites Large Language Models Can Self-Improve.

Self-Improvement in Language Models: The Sharpening Mechanism Large Language Models Can Self-Improve

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.712014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.712014Z digest=sha256:5d3d1b73bd8599fbecef8bf998288fb7a2f7a05a39f57d273862e89c806eda99

Observation 0a225960-04a1-4a92-8e2f-a658629ba70f · outbound

This paper cites Mistral 7B.

Self-Improvement in Language Models: The Sharpening Mechanism Mistral 7B

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.715920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.715920Z digest=sha256:aca66dbc6200eda9da70d02f777a30b3c79758a253ef86042e4ae05e050a8319

Observation 4ecb5d6a-a9cb-44b6-a322-befa3abcb542 · outbound

This paper cites Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer.

Self-Improvement in Language Models: The Sharpening Mechanism Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.724462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.724462Z digest=sha256:3659dd4476223663c45c04eb1919169036a0633c7a47771fe9f20be971a5632e

Observation d7cea87d-19b0-4b6c-8247-e08c0aace764 · outbound

This paper cites Auto-Regressive Next-Token Predictors are Universal Learners.

Self-Improvement in Language Models: The Sharpening Mechanism Auto-Regressive Next-Token Predictors are Universal Learners

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.728418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.728418Z digest=sha256:536a9dc6f2fcd0e68ed4c59f5132abe48d54928317e2f3d262162fd816151066

Observation e9d616e4-1187-4a71-9e08-6a2ced45bf4b · outbound

This paper cites If beam search is the answer, what was the question?.

Self-Improvement in Language Models: The Sharpening Mechanism If beam search is the answer, what was the question?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.732058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.732058Z digest=sha256:f0d951f734b383561f2cc7f9c124fae37ba6b75025620cab686d113b19c8cd6d

Observation de860048-80a8-4c98-928b-e37c2eb7d89a · outbound

This paper cites Controlled Decoding from Language Models.

Self-Improvement in Language Models: The Sharpening Mechanism Controlled Decoding from Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.736440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.736440Z digest=sha256:00a6eee02cbfb39cd9015b82c1e0ea7ce7321943eb69c21937c39234f45e0d0e

Observation 5a8d4734-9aa9-4cac-9a4e-e10413807cee · outbound

This paper cites West-of-N: Synthetic Preferences for Self-Improving Reward Models.

Self-Improvement in Language Models: The Sharpening Mechanism West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.744562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.744562Z digest=sha256:7da9117620a4ec4c95b8056d14274c48ec41b1a663b6364b971bd37cc4300974

Observation e0a80c5d-efc9-4ee1-ac42-e7fcbf1dbff1 · outbound

This paper cites Language Model Self-improvement by Reinforcement Learning Contemplation.

Self-Improvement in Language Models: The Sharpening Mechanism Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.748447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.748447Z digest=sha256:695346fba6b76bec975ad5a2d4a0b07683da689fc6ccb081e29ace0da8ab30be

Observation 4bfc5069-9fde-48bd-a408-74f5244ffb56 · outbound

This paper cites Understanding the Gains from Repeated Self-Distillation.

Self-Improvement in Language Models: The Sharpening Mechanism Understanding the Gains from Repeated Self-Distillation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.752473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.752473Z digest=sha256:7b4c2a9b828a88d98b9cd62dfe391c03c37b46e8d2c9c8d81752f03ba0db7ca4

Observation 87f34528-7e9e-47f8-b25f-7d3625855b21 · outbound

This paper cites The Entropy Enigma: Success and Failure of Entropy Minimization.

Self-Improvement in Language Models: The Sharpening Mechanism The Entropy Enigma: Success and Failure of Entropy Minimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.756635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.756635Z digest=sha256:d637e8809347b7afc97a7fcd49385435a950c4c2f9eeb8e03141564c2167679c

Observation 9c566fc9-7d28-41df-b585-eb9d12e9ea1b · outbound

This paper cites Recursive Introspection: Teaching Language Model Agents How to Self-Improve.

Self-Improvement in Language Models: The Sharpening Mechanism Recursive Introspection: Teaching Language Model Agents How to Self-Improve

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.760513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.760513Z digest=sha256:6384ba32173e7a5a51ac316e3f9e470b8cca9105597809057fa07415a4f7a03e

Observation 5580a3f5-f4cf-46f6-872d-ae9f8d2da078 · outbound

This paper cites BOND: Aligning LLMs with Best-of-N Distillation.

Self-Improvement in Language Models: The Sharpening Mechanism BOND: Aligning LLMs with Best-of-N Distillation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.773191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.773191Z digest=sha256:5ec423cc812163117133339e69be2f91d9f5aaab602f415745cd3f1894869075

Observation b6402f5d-daae-4bb7-bf04-dab266799e55 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Self-Improvement in Language Models: The Sharpening Mechanism Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.777382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.777382Z digest=sha256:3d205ad6c908c620abdf5a71904320833a97f44976ebf8e33e9222e2da269eff

Observation d2b150d1-58cd-4487-aa4b-176f7418581f · outbound

This paper cites The Importance of Online Data: Understanding Preference Fine-tuning via Coverage.

Self-Improvement in Language Models: The Sharpening Mechanism The Importance of Online Data: Understanding Preference Fine-tuning via Coverage

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.781971Z digest=sha256:bc04869ba4012669cc72b671d1a6aa73c9925fdd88f2be77ce5f4f84648ae177

Observation 2cece19f-bf61-46fd-b1b8-ff72e0ce75c4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Self-Improvement in Language Models: The Sharpening Mechanism Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.786588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.786588Z digest=sha256:88fd5f900e0addadd9a0ef118541ed2c29eff2f59663f38fb8ceac14099a0567

Observation e213e65d-7b99-4c4f-99fd-08adad6f00cb · outbound

This paper cites Tent: Fully Test-time Adaptation by Entropy Minimization.

Self-Improvement in Language Models: The Sharpening Mechanism Tent: Fully Test-time Adaptation by Entropy Minimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.791360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.791360Z digest=sha256:4edc2e2061a2dd8c2c2e72808fb538666a1cb29e82d8ced3addbbd2eb9f37f1c

Observation 2ecd20b6-2550-40e1-b732-f15ab39e5c62 · outbound

This paper cites Self-Taught Evaluators.

Self-Improvement in Language Models: The Sharpening Mechanism Self-Taught Evaluators

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.796547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.796547Z digest=sha256:745094d89dbfdb2ce9d12163f56060c8152ed406ef53a7ac923c4d0f814afd05

Observation 23e2b478-f42c-4ddb-a92f-465bd58563a5 · outbound

This paper cites Chain-of-Thought Reasoning Without Prompting.

Self-Improvement in Language Models: The Sharpening Mechanism Chain-of-Thought Reasoning Without Prompting

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.800609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.800609Z digest=sha256:ff93208aa0015f41e9a7dfc153cb9da6d451132bd1fc5f1c66698ad564deb109

Observation 2f244995-23bd-4bb2-ab14-0cd4770b9397 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Self-Improvement in Language Models: The Sharpening Mechanism Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.805880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.805880Z digest=sha256:4e8d4146b093a7312b740cbcd29bcd4828795144fe637867a8938fdba60d59a4

Observation 33a550c8-070a-4b15-bf0b-f0ed9f0d6ce8 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Self-Improvement in Language Models: The Sharpening Mechanism Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.810167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.810167Z digest=sha256:e87aae278cd502e2ec79cd53bf53407269e865f698359e3e3c28f787f07b149c

Observation ea76dee6-6ef3-4b61-bfe5-e21682a2bc57 · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Self-Improvement in Language Models: The Sharpening Mechanism Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.814539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.814539Z digest=sha256:72f4545fabc1c233b70dbb46959d6203a7cf03d062b069e6508d3d4d5f1ed797

Observation 182c56cc-2c66-43db-b674-7b15889833d9 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Self-Improvement in Language Models: The Sharpening Mechanism Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.818560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.818560Z digest=sha256:ad1ddbf00f141cae9abe76327f50778f8e17a37c21d7b673500899158f3aac62

Observation e5212328-04eb-4a7d-ac0d-ed1f8725245c · outbound

This paper cites Asymptotics of Language Model Alignment.

Self-Improvement in Language Models: The Sharpening Mechanism Asymptotics of Language Model Alignment

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.823508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.823508Z digest=sha256:f7e6bdac07ff3d2452b21357b63a65c6731139e547eb94afb0ef2d0b117cb1ba

Observation e98b4f2e-9bb9-48cf-aa18-cc9c55bd8095 · outbound

This paper cites Online Iterative Reinforcement Learning from Human Feedback with General Preference Model.

Self-Improvement in Language Models: The Sharpening Mechanism Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.827757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.827757Z digest=sha256:e562afc36145c7903901afba57a589ccc412c47c00f9394eab697f44e26d2e67

Observation 1fe62b64-63cb-4a12-91b6-bf7765fe99d4 · outbound

This paper cites Self-Rewarding Language Models.

Self-Improvement in Language Models: The Sharpening Mechanism Self-Rewarding Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.831684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.831684Z digest=sha256:41373c31a268edef25a4a0b93868658a153c55b9544040aa1e395cfd8efbbca5

Observation 474be5c1-9b8f-4b67-8682-abe60503e6c7 · outbound

This paper cites an unresolved cited work.

Self-Improvement in Language Models: The Sharpening Mechanism Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:11:15.733206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:11:14.835570Z digest=sha256:04100b46e7c2c6931e862f55ad9b71665b1aac4e2c907d69812320d0cbefebb1

Observation b0004a9f-e054-467e-9568-9e911add078d · outbound

This paper cites LLM-as-a-Judge.

Self-Improvement in Language Models: The Sharpening Mechanism LLM-as-a-Judge

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.722498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:11:14.839322Z digest=sha256:70f04960dae5e3a719e55eb558c90d9c79352b10010267b96e69fee62c238f30

Observation bcae0b3a-9da3-4dd5-b997-43a78a11b2a8 · outbound

This paper cites an unresolved cited work.

Self-Improvement in Language Models: The Sharpening Mechanism Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:11:15.711680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:11:14.843452Z digest=sha256:a2c432cb320ffafb6340f5908d0b9e0f08f042fe651f483dee1fb02ce3416f04

Observation 485d398b-78f6-4fdd-9fab-b7a7f7f79349 · outbound

This paper cites More sophisticated inference-time search strategies such tree search and MCTS (Yao et al., 2024; Wan et al., 2024; Mudgal et al., 2023; Zhao et al.,.

Self-Improvement in Language Models: The Sharpening Mechanism More sophisticated inference-time search strategies such tree search and MCTS (Yao et al., 2024; Wan et al., 2024; Mudgal et al., 2023; Zhao et al.,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.699555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:11:14.847320Z digest=sha256:644c8c770e48204f4ba44240bd7ded5712ab1f3377b529d4ec6d389b4d7cf1b7

Observation d15a36d1-b38a-4be8-a229-e1d28059e7e3 · outbound

This paper cites Perhaps most closely related to our work is Frei et al.

Self-Improvement in Language Models: The Sharpening Mechanism Perhaps most closely related to our work is Frei et al

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.687294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:11:14.851177Z digest=sha256:cfdf1718274132d49210bcfa5a7d93d6d6a180d3a7f46c26a68212790b200dff

Observation cd1fac52-502e-41e7-9dad-ac4b99c7bc6d · outbound

This paper cites Proof of Proposition C.1.

Self-Improvement in Language Models: The Sharpening Mechanism Proof of Proposition C.1

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.663215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:11:14.858827Z digest=sha256:1ab20159d32b0e2fb4e0bc47d2705268900dee567468dfb986ef2489ed444ac8

Observation 7cdb21b2-5c02-465a-a7d3-4441a2eb07e4 · outbound

This paper cites We quantify the quality of a sharpened model as follows.

Self-Improvement in Language Models: The Sharpening Mechanism We quantify the quality of a sharpened model as follows

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.651219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:11:14.863340Z digest=sha256:08273bbc8e8a7087c5aa46ad5c6d3e3cab216ca2c56b0722c41870cc25eaec0b

Observation 43c7ee13-b33b-4892-a93f-3f6453d3e602 · outbound

This paper cites Now a⊤Γ−1 t−1a ≤ 1/λ with probability 1, where λ = λmin(Γ0).

Self-Improvement in Language Models: The Sharpening Mechanism Now a⊤Γ−1 t−1a ≤ 1/λ with probability 1, where λ = λmin(Γ0)

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.637243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:11:14.870651Z digest=sha256:50bdff2787e49740832e8216deaf94560d49a2505e72efa99ad6194dfadbee94

Observation 02465288-f85a-49e2-b606-2e43efbb3d91 · outbound

This paper cites Note that unlike the non-adaptive framework, the distribution overmi depends on the underlying instance I with which the algorithm interacts.

Self-Improvement in Language Models: The Sharpening Mechanism Note that unlike the non-adaptive framework, the distribution overmi depends on the underlying instance I with which the algorithm interacts

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.623676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:11:14.875258Z digest=sha256:3b92a2ce080f48524b56c67123299e41b10560c05225e41c48432dea6de45794

Observation 4f13ea1f-2315-4d54-bf2b-dd97e64c56da · outbound

This paper cites an unresolved cited work.

Self-Improvement in Language Models: The Sharpening Mechanism Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:11:15.593265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:11:14.883441Z digest=sha256:382bad4b93657a33376eba64ed2ff8ddf6a897f181bfb0a497e287067f0bf6da

Observation f05ee8cb-53e6-40b7-b2b1-2996a591e203 · outbound

This paper cites Initialize: π(1) ← πbase, D(0) ← ∅.

Self-Improvement in Language Models: The Sharpening Mechanism Initialize: π(1) ← πbase, D(0) ← ∅

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.580691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:11:14.887423Z digest=sha256:6d10a155bf2f40f0847605295ce414f543dfc4e987a00e097208b8a81391d710

Observation b7cd25d5-9199-46fb-b847-93b6ad41f2f4 · outbound

This paper cites Also suppose thatπ⋆ β ∈ Π where π⋆ β(y | x) ∝ π1+β−1 base (y | x).

Self-Improvement in Language Models: The Sharpening Mechanism Also suppose thatπ⋆ β ∈ Π where π⋆ β(y | x) ∝ π1+β−1 base (y | x)

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.566237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:11:14.891530Z digest=sha256:bb9234b15e013795b036a059a951b2d71c0f012f5d3bdd6e788a893adeb54832

Observation b57a37d9-567b-42a2-b260-324885c83b69 · outbound

This paper cites Chain of Thought Empowers Transformers to Solve Inherently Serial Problems.

Self-Improvement in Language Models: The Sharpening Mechanism Chain of Thought Empowers Transformers to Solve Inherently Serial Problems

Reference 1973

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.720249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.720249Z digest=sha256:22e08a95be15ad5227638e2636d77dca349615b316d14bfa9206c430b68d1130

Observation 100ba0b1-6cce-431d-ad6d-aa9ddcc9f180 · outbound

This paper cites GPT-4 Technical Report.

Self-Improvement in Language Models: The Sharpening Mechanism GPT-4 Technical Report

Reference 1983

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.740875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.740875Z digest=sha256:3edcb29c522d8392a7e30606a544c36e2bc05d861a030db838b52580782f70f6

Observation fbbf4766-f474-43ae-a5f0-3e499cdf04bf · outbound

This paper cites Towards a theory of model distillation.

Self-Improvement in Language Models: The Sharpening Mechanism Towards a theory of model distillation

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.631474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.631474Z digest=sha256:0747e65c50051d957d3b859e512a1bdb8bd1c2073808bc711bddb7105c07d730

Observation 1fa65628-c83b-4a35-a0b8-504f9e4fc9bb · outbound

This paper cites Butterfly Effects of SGD Noise: Error Amplification in Behavior Cloning and Autoregression.

Self-Improvement in Language Models: The Sharpening Mechanism Butterfly Effects of SGD Noise: Error Amplification in Behavior Cloning and Autoregression

Reference 1995

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.627241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.627241Z digest=sha256:21a6813a83655b5a0e86456a6f01a8ff8128ff28c6a0b4139c2c0e2e571a8745

Observation f23a1978-c153-4140-af44-9127a7b0e72f · outbound

This paper cites BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling.

Self-Improvement in Language Models: The Sharpening Mechanism BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.681741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.681741Z digest=sha256:c12ae91cf02acee2a9b460a4110936800249d40ddd19c16258c49ad7cc371c3f

Observation c5642192-9f09-4e58-9c3c-65b637a8bb1f · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Self-Improvement in Language Models: The Sharpening Mechanism Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.640335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.640335Z digest=sha256:59571a7753734fff012ca97a84ea1f05357b8f1c1db6d5888f4d0caa331de4ad

Observation c9870b06-ed88-4bdf-8ac6-0216350d89f4 · outbound

This paper cites In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised Learning.

Self-Improvement in Language Models: The Sharpening Mechanism In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised Learning

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.764758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.764758Z digest=sha256:4b59a8e173c831e3df068907957a9af489248e8135f630b9f6668241a40481d6

Observation 645a2964-83fb-4947-b139-d1a25cb4b978 · outbound

This paper cites PaLM 2 Technical Report.

Self-Improvement in Language Models: The Sharpening Mechanism PaLM 2 Technical Report

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.677287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.677287Z digest=sha256:71df045d50273a6da76764fa6a3aa3572453adbd4accbf04518c86bae876a0de

Observation 0308ba6c-f907-459d-a12b-5da8b4789a63 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Self-Improvement in Language Models: The Sharpening Mechanism LoRA: Low-Rank Adaptation of Large Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.699627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.699627Z digest=sha256:2ff11338a662ef8d7d55cdd2bef9a4abe7869fb63721bfa461801fad1d10553e

Observation 76e57ec9-04a5-4467-86ee-0c576f8db4f9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Self-Improvement in Language Models: The Sharpening Mechanism Proximal Policy Optimization Algorithms

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.768719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.768719Z digest=sha256:ffa7414c5b9fa869a0822fc1dcac5a87deff4b93c1ad24da7a33c540af90436a

Observation a668f288-be40-42d7-9f6d-5f541cebedfa · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Self-Improvement in Language Models: The Sharpening Mechanism Training Verifiers to Solve Math Word Problems

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.644344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.644344Z digest=sha256:fda9c45c1d61e33530c8bac0f47805a15ca19a3d9d0fffd79dcdecf25a001bfe

Observation 8f21e54c-60e5-48c2-8b3b-6e81dd482993 · outbound

This paper cites Distillation $\approx$ Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network.

Self-Improvement in Language Models: The Sharpening Mechanism Distillation $\approx$ Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.656592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.656592Z digest=sha256:f5e0acb2198951d0f285ef6065961fa32f2ac9178488e52432aa9f05845a7c7c

Observation b84f714c-230d-40f4-acba-d4382f15c153 · outbound

This paper cites The Llama 3 Herd of Models.

Self-Improvement in Language Models: The Sharpening Mechanism The Llama 3 Herd of Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.660901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.660901Z digest=sha256:b99a59bb8a391b516371e9a8bdd7e1ebfe95af68146c4077227c31464f0c82d0

Observation d7948727-c316-488a-b7bc-9d176cc9fae7 · outbound

This paper cites Variational Best-of-N Alignment.

Self-Improvement in Language Models: The Sharpening Mechanism Variational Best-of-N Alignment

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.618718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.618718Z digest=sha256:0d998a1d0a81862f316cfa5fc759102bce6f82e1876234ade28dc1efc7a0f14e

Observation 9489302e-5150-489d-8ff0-5969331fd422 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Self-Improvement in Language Models: The Sharpening Mechanism Distilling the Knowledge in a Neural Network

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.695697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.695697Z digest=sha256:5115171d0e20d6b7a164c26b4d4c62926e4de8ff294f8f008c394722c2c75391

Observation b33b96a4-3cf4-41af-84b2-d8c087597fa8 · outbound

This paper cites Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning.

Self-Improvement in Language Models: The Sharpening Mechanism Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.614276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.614276Z digest=sha256:ee91ca21e690246ab2c774773f814d3b95351409f4c7bada385b5db01e22fc5a

Observation 8bbafa47-6fa0-4958-8e87-dfad73cd2737 · outbound

This paper cites Retraining with Predicted Hard Labels Provably Increases Model Accuracy.

Self-Improvement in Language Models: The Sharpening Mechanism Retraining with Predicted Hard Labels Provably Increases Model Accuracy

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.648794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.648794Z digest=sha256:402d4aaff67ceaacc52436615a2c082b81db107bb9733d993c9d57329b5de7a7

Observation 58aa67ff-9efd-4d3b-beea-8186a5fb3bec · outbound

This paper cites Transferring Inductive Biases through Knowledge Distillation.

Self-Improvement in Language Models: The Sharpening Mechanism Transferring Inductive Biases through Knowledge Distillation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.609911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.609911Z digest=sha256:57bc0295baf5b15d3e33a8c5ac73b54addbed551b067c26577f498f79804c9f5

Pith citing papers

Observation f5053f7f-ac80-44c3-aa4a-e6ed349a16f9 · inbound

LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds cites this paper.

LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds Self-Improvement in Language Models: The Sharpening Mechanism

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:53:37.179431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:53:37.179431Z digest=sha256:a53df71b6bc0cd39a17fe3b6a779db486dc4de2bb95b6843395c42f02904ce47

Observation 715ea1a9-cf4b-474f-b290-e9e02e7b081d · inbound

Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges cites this paper.

Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges Self-Improvement in Language Models: The Sharpening Mechanism

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T14:54:29.183688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:54:29.183688Z digest=sha256:f50d9bc113808ccdda772a4076d26cbfea3b71180387918d839c02387ed04db7

Observation c97aaf18-51de-4356-adfe-fe42ae340224 · inbound

Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards cites this paper.

Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards Self-Improvement in Language Models: The Sharpening Mechanism

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:13.990974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:13.990974Z digest=sha256:130a20d859d868a897946f13bfa54f3bf14bc0b0c19fc94e45a9372184d03ce9

Observation e68c7525-f7e4-48db-906e-ced8dded2667 · inbound

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning cites this paper.

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning Self-Improvement in Language Models: The Sharpening Mechanism

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:58:33.372875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T15:58:33.219451Z digest=sha256:4cad2c88172e82169c52d10ed6b4d64cde4db463c8029528331a2b202f8dec43

Observation b479d45d-7c4b-4ae8-800e-17e7f52fb058 · inbound

Reinforcing General Reasoning without Verifiers cites this paper.

Reinforcing General Reasoning without Verifiers Self-Improvement in Language Models: The Sharpening Mechanism

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.425938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.425938Z digest=sha256:3f05b53cb680bd9a2f30e41c73d6a66e7124bfbb3d9054c3473dcc5f68af9faf

Observation 179cc3d2-eb8c-4a8f-a54d-9aa6771e01ee · inbound

Sample Complexity and Representation Ability of Test-time Scaling Paradigms cites this paper.

Sample Complexity and Representation Ability of Test-time Scaling Paradigms Self-Improvement in Language Models: The Sharpening Mechanism

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:35.274446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:35.274446Z digest=sha256:177c6f2a27468ace7097bd5ff3375b84338984df0cb5736d3972d9e4ef133b0c

Observation 2244849d-c21a-4ff2-82fb-0ec8ec7e670a · inbound

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs cites this paper.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Self-Improvement in Language Models: The Sharpening Mechanism

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.163556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.163556Z digest=sha256:80c1157b9543bae339920f07f5013c1a10fbef3238570f3f115a90a67792250e

Observation 46c8a55c-b236-4164-bb68-6b9cd86bfda9 · inbound

Post-Completion Learning for Language Models cites this paper.

Post-Completion Learning for Language Models Self-Improvement in Language Models: The Sharpening Mechanism

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:52:15.326811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:52:15.326811Z digest=sha256:5b87398c29d0a763f6517fd76dd9e0d2969c8ac3b4ce15e25d5d260d54be8e8b

Observation c99e57b6-dc00-4d09-b4a1-c0c4d0d6ae38 · inbound

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future cites this paper.

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Self-Improvement in Language Models: The Sharpening Mechanism

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:49.580054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:06:49.580054Z digest=sha256:08f6c2da8f8b47fef86beb475da9e6e5553f072f942f596f7f22180ba09bd3b0

Observation c83422e7-447b-4537-a2bf-d80056733181 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Self-Improvement in Language Models: The Sharpening Mechanism

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.702241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.702241Z digest=sha256:05686fed385513e73aaf5b8684f77986070b69c2f1e4b0a507625b5ec35e7d2d

Observation d9e3b4c8-2b19-4b7e-8c88-eaf2aad9f39a · inbound

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula cites this paper.

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula Self-Improvement in Language Models: The Sharpening Mechanism

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T02:43:35.074111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:43:35.074111Z digest=sha256:499f1125fc0e976bd8b9018e6e17c626a96727666b728ace2d90a1d3775b170a

Observation b8022d91-06f9-4e48-80f0-26072928b072 · inbound

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning cites this paper.

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning Self-Improvement in Language Models: The Sharpening Mechanism

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T23:46:32.301737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:46:32.301737Z digest=sha256:1211f37d6d2eeae3c2bab51cdd22cb5713f79031bac86218b8b0069af04f940a

Observation 4d7ed3f2-720a-485e-9dfe-8f0a24d7d7ac · inbound

The Role of Generator Access in Autoregressive Post-Training cites this paper.

The Role of Generator Access in Autoregressive Post-Training Self-Improvement in Language Models: The Sharpening Mechanism

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:05:48.110480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T20:16:46.370831Z digest=sha256:f517e6a1168167792f53ff664721408bda64815860a6550c7051d83efff53235

Observation 0e98ced7-f560-41da-bdbd-eb0e9d7fb3c8 · inbound

Beyond Distribution Sharpening: The Importance of Task Rewards cites this paper.

Beyond Distribution Sharpening: The Importance of Task Rewards Self-Improvement in Language Models: The Sharpening Mechanism

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:08:27.037680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T08:07:14.691463Z digest=sha256:76c11a205eccd8f0d4a8d5c28a3e70cd2c4809187bb3d8009d279e1c58e16976

Observation 73f72c26-913e-4e00-9097-f07e002fca28 · inbound

On the Generalization Gap in Self-Evolving Language Model Reasoning cites this paper.

On the Generalization Gap in Self-Evolving Language Model Reasoning Self-Improvement in Language Models: The Sharpening Mechanism

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.862765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:e2e470627ae3c521c43c5d1088f36c171789f335f1669532ce236a06e97e9e6f

Observation 51af3dc1-1f41-4b67-8962-6fdb0545fc90 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Self-Improvement in Language Models: The Sharpening Mechanism

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.665777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:520c4ca8b32316c6d9648f84739dd938a76459b9354b3935bf48f7caeb38a642

Observation 70c40a63-4681-4129-b016-7a5f5524e943 · inbound

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning cites this paper.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Self-Improvement in Language Models: The Sharpening Mechanism

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:26.274337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:b8e4d35159e4f5e699bea84efee2a755d8f5d5b1dd51c092722af62dee0b86af

Observation fc408979-3023-4995-b367-ff8699d39f6a · inbound

Select and Improve: Understanding the Mechanics of Post-Training for Reasoning cites this paper.

Select and Improve: Understanding the Mechanics of Post-Training for Reasoning Self-Improvement in Language Models: The Sharpening Mechanism

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:08:21.734036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T07:14:47.302277Z digest=sha256:ff6bc71c7793d96689a2a4881ad7fb89a02e1c06921f2312f32df8c1c57fb347