Pith. sign in

Paper Citation Record · LEDGER

Self-Improvement in Language Models: The Sharpening Mechanism

As of 15 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 16 inbound Pith citation observations for arXiv:2412.01951.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01951 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:11:14.891530Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:53:37.179431Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:08:21.732193Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1fb15654-2f89-4084-b60d-78fbafbc5140 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Self-Improvement in Language Models: The Sharpening Mechanism Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.604915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.604915Z digest=sha256:1f26900124ceec411ac751fde3494fa716487942a07469fa008a3cc1a35d23b2

Observation 14864a9c-3e66-4e79-9a7e-946e425a786e · outbound

This paper cites an unresolved cited work.

Self-Improvement in Language Models: The Sharpening Mechanism Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:11:15.675684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:11:14.855033Z digest=sha256:2bfb29dcfd2ec38735883ffc2bb41edcd1b45eb7ee79d52ac5130bbd6c3a2a79

Observation 27f039b6-6ea1-4fda-9230-a402e54bcd7f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Self-Improvement in Language Models: The Sharpening Mechanism Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.623071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.623071Z digest=sha256:de40ce85e9718a81a4ce6aec7db9b0d32c419f267603d2b7c9298f8312b05307

Observation 7f531f67-c03e-433f-bd4b-19dcb247b78a · outbound

This paper cites J.1.2 Proof of Theorem 4.2′ Proof of Theorem 4.2′.

Self-Improvement in Language Models: The Sharpening Mechanism J.1.2 Proof of Theorem 4.2′ Proof of Theorem 4.2′

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.609379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:11:14.879283Z digest=sha256:c60037d34f31ac83010e05e1679055e6a8cf1d44cfc13bca98ffcbea2fd4485d

Observation f9750400-8623-4dcb-8275-b44a276f1434 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Self-Improvement in Language Models: The Sharpening Mechanism Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.635630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.635630Z digest=sha256:a48a600b33f1c1b6d4ab7351813cb8873e15c42b55b4c45d25fb90abd2753763

Observation a1a19f5b-306a-4106-87a7-90cc46f728b3 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Self-Improvement in Language Models: The Sharpening Mechanism BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.652665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.652665Z digest=sha256:bbb8ad1d48d9da585890e0b47303820919f56212133ddd8c65972cea668d00d7

Observation de0cfaa2-1a94-4681-9f56-ce9a6e749911 · outbound

This paper cites Foundations of Reinforcement Learning and Interactive Decision Making.

Self-Improvement in Language Models: The Sharpening Mechanism Foundations of Reinforcement Learning and Interactive Decision Making

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.664926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.664926Z digest=sha256:ce0b19b46daf863cd97bc349c0950e4faa474c4ea41a1b33fcc29c07f9f099e1

Observation 9d41769f-0976-44c9-8b0e-c04e47963bfe · outbound

This paper cites The Statistical Complexity of Interactive Decision Making.

Self-Improvement in Language Models: The Sharpening Mechanism The Statistical Complexity of Interactive Decision Making

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.668664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.668664Z digest=sha256:973e34fc32fdac5e6f4bb03d3033729aa3b6aa0181d34e229c6626e227f497d1

Observation 334e2da3-8166-41b8-9891-3350675f0b0e · outbound

This paper cites REBEL: Reinforcement Learning via Regressing Relative Rewards.

Self-Improvement in Language Models: The Sharpening Mechanism REBEL: Reinforcement Learning via Regressing Relative Rewards

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.672865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.672865Z digest=sha256:303a43b969426f937b005e1dbe74e781d726f4b98d097f69668bdb0f3b718e8f

Observation 5d599889-188e-48a1-b903-6776752e606f · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Self-Improvement in Language Models: The Sharpening Mechanism Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.686228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.686228Z digest=sha256:f29f40fe0d9a3bd13bd0af9db3c9eadebcc0235002e523869e5829bf89084b24

Observation ecca62f2-d97e-4dc6-a1f0-813308f74f10 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Self-Improvement in Language Models: The Sharpening Mechanism Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.690922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.690922Z digest=sha256:d86495cd852649f013eea4728c99701a9e371a967f74a7d66119df6a1c8c7d47

Observation f99de5e0-db70-491d-a30c-3b37a29e4a47 · outbound

This paper cites Amortizing intractable inference in large language models.

Self-Improvement in Language Models: The Sharpening Mechanism Amortizing intractable inference in large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.703729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.703729Z digest=sha256:63146cb0d259aca6ac52afd8e0e2b8dee83a743ca6fffaaaadaf7ef1b07dd535

Observation 4f835749-1798-421a-b54d-e5b277411b0a · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Self-Improvement in Language Models: The Sharpening Mechanism Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.707926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.707926Z digest=sha256:83425fb9614719a80064df8dda7e4459de8dd06e0a0ba7c53ea8424cea164ac5

Observation ee9f7080-a4e1-43f1-8178-68a2f29df2df · outbound

This paper cites Large Language Models Can Self-Improve.

Self-Improvement in Language Models: The Sharpening Mechanism Large Language Models Can Self-Improve

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.712014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.712014Z digest=sha256:7c83c675c6ba304b75772fdaf4e74d3402d2a69e26b1cd8e2fa2a3d22f1029bc

Observation 0a225960-04a1-4a92-8e2f-a658629ba70f · outbound

This paper cites Mistral 7B.

Self-Improvement in Language Models: The Sharpening Mechanism Mistral 7B

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.715920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.715920Z digest=sha256:0ed3b8b035b25edf417962e29d61815afa5954525ceff1ee616cf493bc0259ca

Observation 4ecb5d6a-a9cb-44b6-a322-befa3abcb542 · outbound

This paper cites Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer.

Self-Improvement in Language Models: The Sharpening Mechanism Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.724462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.724462Z digest=sha256:7444ae7b1eff070959d152b93fed4e021a5cd7d3104ea6d5df59b31bc1a7dc36

Observation d7cea87d-19b0-4b6c-8247-e08c0aace764 · outbound

This paper cites Auto-Regressive Next-Token Predictors are Universal Learners.

Self-Improvement in Language Models: The Sharpening Mechanism Auto-Regressive Next-Token Predictors are Universal Learners

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.728418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.728418Z digest=sha256:a60da75176bc6afb4f6e537726a430e4e620ebd42ed52c9875c4e7b00ef83337

Observation e9d616e4-1187-4a71-9e08-6a2ced45bf4b · outbound

This paper cites If beam search is the answer, what was the question?.

Self-Improvement in Language Models: The Sharpening Mechanism If beam search is the answer, what was the question?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.732058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.732058Z digest=sha256:3c8e9b668c465c99e0b07000e9a5b535d9dcd017083e1f3e61a630fa8f698525

Observation de860048-80a8-4c98-928b-e37c2eb7d89a · outbound

This paper cites Controlled Decoding from Language Models.

Self-Improvement in Language Models: The Sharpening Mechanism Controlled Decoding from Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.736440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.736440Z digest=sha256:02eb7e9a24342d8e6005f9df8905fd009c3ffba118b860191f641b6724bfe049

Observation 5a8d4734-9aa9-4cac-9a4e-e10413807cee · outbound

This paper cites West-of-N: Synthetic Preferences for Self-Improving Reward Models.

Self-Improvement in Language Models: The Sharpening Mechanism West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.744562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.744562Z digest=sha256:30f84a2e140b030b9b3ff9b255cb8517a6316e66e8dd59c10b64249b585a5b8a

Observation e0a80c5d-efc9-4ee1-ac42-e7fcbf1dbff1 · outbound

This paper cites Language Model Self-improvement by Reinforcement Learning Contemplation.

Self-Improvement in Language Models: The Sharpening Mechanism Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.748447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.748447Z digest=sha256:e5f0172c3a54739c48d5b741c102007b901c6ab1d2c2aa9497bfc470d3fdcd66

Observation 4bfc5069-9fde-48bd-a408-74f5244ffb56 · outbound

This paper cites Understanding the Gains from Repeated Self-Distillation.

Self-Improvement in Language Models: The Sharpening Mechanism Understanding the Gains from Repeated Self-Distillation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.752473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.752473Z digest=sha256:b232e0d8dc50b57b14f11e9c985f0cef9d3da5f3887650ec532d6f72a41d1905

Observation 87f34528-7e9e-47f8-b25f-7d3625855b21 · outbound

This paper cites The Entropy Enigma: Success and Failure of Entropy Minimization.

Self-Improvement in Language Models: The Sharpening Mechanism The Entropy Enigma: Success and Failure of Entropy Minimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.756635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.756635Z digest=sha256:8dfecedf1cd41726da325e2fb4eafa91ca451eebea762edc796dca7ae0e41d4b

Observation 9c566fc9-7d28-41df-b585-eb9d12e9ea1b · outbound

This paper cites Recursive Introspection: Teaching Language Model Agents How to Self-Improve.

Self-Improvement in Language Models: The Sharpening Mechanism Recursive Introspection: Teaching Language Model Agents How to Self-Improve

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.760513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.760513Z digest=sha256:f354e145d7c4ad4007e8e96e64cbcca409e6edfc55bf752503ce1267b5e861eb

Observation 5580a3f5-f4cf-46f6-872d-ae9f8d2da078 · outbound

This paper cites BOND: Aligning LLMs with Best-of-N Distillation.

Self-Improvement in Language Models: The Sharpening Mechanism BOND: Aligning LLMs with Best-of-N Distillation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.773191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.773191Z digest=sha256:d63c920909a808b24ca6a91f04440e965797d7e89ddb3a5beeb5e5c7767d56bb

Observation b6402f5d-daae-4bb7-bf04-dab266799e55 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Self-Improvement in Language Models: The Sharpening Mechanism Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.777382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.777382Z digest=sha256:0def3cdded1e77b9de9b769c364d9025cae589564dda8f596cfa90d0eda16286

Observation d2b150d1-58cd-4487-aa4b-176f7418581f · outbound

This paper cites The Importance of Online Data: Understanding Preference Fine-tuning via Coverage.

Self-Improvement in Language Models: The Sharpening Mechanism The Importance of Online Data: Understanding Preference Fine-tuning via Coverage

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.781971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.781971Z digest=sha256:0bb3211f58d6ce605d4e1f87313292078ffdeda6e04c8058b1e5006b76f9b280

Observation 2cece19f-bf61-46fd-b1b8-ff72e0ce75c4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Self-Improvement in Language Models: The Sharpening Mechanism Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.786588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.786588Z digest=sha256:305c64b0bd97f816dbe8a8a3485b99bd845ff40afbc94b590339c638761bcce2

Observation e213e65d-7b99-4c4f-99fd-08adad6f00cb · outbound

This paper cites Tent: Fully Test-time Adaptation by Entropy Minimization.

Self-Improvement in Language Models: The Sharpening Mechanism Tent: Fully Test-time Adaptation by Entropy Minimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.791360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.791360Z digest=sha256:9246d4a5dab6efabb71b3de09b06d9453f909fb4f242db0cc2804d2aa9f89b8f

Observation 2ecd20b6-2550-40e1-b732-f15ab39e5c62 · outbound

This paper cites Self-Taught Evaluators.

Self-Improvement in Language Models: The Sharpening Mechanism Self-Taught Evaluators

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.796547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.796547Z digest=sha256:a765d46e93647dc14efeb495e543c7c8d48f5090e5da265ed5a7937445fe729c

Observation 23e2b478-f42c-4ddb-a92f-465bd58563a5 · outbound

This paper cites Chain-of-Thought Reasoning Without Prompting.

Self-Improvement in Language Models: The Sharpening Mechanism Chain-of-Thought Reasoning Without Prompting

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.800609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.800609Z digest=sha256:6e7b20306cdeab71cb0ed1f887c64fa08000a7b5ad0c031c3805cf256d8c5418

Observation 2f244995-23bd-4bb2-ab14-0cd4770b9397 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Self-Improvement in Language Models: The Sharpening Mechanism Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.805880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.805880Z digest=sha256:0e408115bc6287881391c08877feb38e5da9c5f3ec5e8a5d89c95ee21dab4210

Observation 33a550c8-070a-4b15-bf0b-f0ed9f0d6ce8 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Self-Improvement in Language Models: The Sharpening Mechanism Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.810167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.810167Z digest=sha256:cf2e017ed1ebf54d1bde5fec19bdcdd9f8304f011b34d000f865ae6fb0cc5ebe

Observation ea76dee6-6ef3-4b61-bfe5-e21682a2bc57 · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Self-Improvement in Language Models: The Sharpening Mechanism Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.814539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.814539Z digest=sha256:53bc59505a6884b42cb7db9f1d58c0a8af184be297f29e338053d5ec753eaf9e

Observation 182c56cc-2c66-43db-b674-7b15889833d9 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Self-Improvement in Language Models: The Sharpening Mechanism Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.818560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.818560Z digest=sha256:f741e57b044bfd760b0780c9a6cfccd5497669d75dca2ff4cc3b39e4307b685c

Observation e5212328-04eb-4a7d-ac0d-ed1f8725245c · outbound

This paper cites Asymptotics of Language Model Alignment.

Self-Improvement in Language Models: The Sharpening Mechanism Asymptotics of Language Model Alignment

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.823508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.823508Z digest=sha256:325d491af5d250abb8586d5daceeac44182cd6d8ebab5702de3fc9d5a8a6b557

Observation e98b4f2e-9bb9-48cf-aa18-cc9c55bd8095 · outbound

This paper cites Online Iterative Reinforcement Learning from Human Feedback with General Preference Model.

Self-Improvement in Language Models: The Sharpening Mechanism Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.827757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.827757Z digest=sha256:2793f024f200f146ab6616510b62daee2cb214079802bd5cad898adee75984fe

Observation 1fe62b64-63cb-4a12-91b6-bf7765fe99d4 · outbound

This paper cites Self-Rewarding Language Models.

Self-Improvement in Language Models: The Sharpening Mechanism Self-Rewarding Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.831684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.831684Z digest=sha256:33f0553603a5b4b4d28bf7d3296e0f4694cddac89b02c16c651377a64a8a66d5

Observation 474be5c1-9b8f-4b67-8682-abe60503e6c7 · outbound

This paper cites an unresolved cited work.

Self-Improvement in Language Models: The Sharpening Mechanism Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:11:15.733206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:11:14.835570Z digest=sha256:c86b616f2f36d2c83d13575533dcd1c2a79fa14cd11d2454a882316eefff99c8

Observation b0004a9f-e054-467e-9568-9e911add078d · outbound

This paper cites LLM-as-a-Judge.

Self-Improvement in Language Models: The Sharpening Mechanism LLM-as-a-Judge

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.722498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:11:14.839322Z digest=sha256:9204b7c46b6cefb9275255d8b13d79384db6ca6bb90c6c1917f4bcac926ef928

Observation bcae0b3a-9da3-4dd5-b997-43a78a11b2a8 · outbound

This paper cites an unresolved cited work.

Self-Improvement in Language Models: The Sharpening Mechanism Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:11:15.711680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:11:14.843452Z digest=sha256:e9d110a0c3d97b5342601ca8894f1c5a5aa61e8fdd193c1fa09c00c3e7558c91

Observation 485d398b-78f6-4fdd-9fab-b7a7f7f79349 · outbound

This paper cites More sophisticated inference-time search strategies such tree search and MCTS (Yao et al., 2024; Wan et al., 2024; Mudgal et al., 2023; Zhao et al.,.

Self-Improvement in Language Models: The Sharpening Mechanism More sophisticated inference-time search strategies such tree search and MCTS (Yao et al., 2024; Wan et al., 2024; Mudgal et al., 2023; Zhao et al.,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.699555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:11:14.847320Z digest=sha256:941a97dde8510ea2e91a0d3222c7277c0d8694650df17c8ac038d355154b0195

Observation d15a36d1-b38a-4be8-a229-e1d28059e7e3 · outbound

This paper cites Perhaps most closely related to our work is Frei et al.

Self-Improvement in Language Models: The Sharpening Mechanism Perhaps most closely related to our work is Frei et al

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.687294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:11:14.851177Z digest=sha256:283ef627fc5db4c9705ed4a7508fa9988a51e435a1f876e37bdb3d1ac511e52b

Observation cd1fac52-502e-41e7-9dad-ac4b99c7bc6d · outbound

This paper cites Proof of Proposition C.1.

Self-Improvement in Language Models: The Sharpening Mechanism Proof of Proposition C.1

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.663215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:11:14.858827Z digest=sha256:2f1d5474b2f10e83adf2f1878699a7cff454532b902f46859b59c33e07c7642a

Observation 7cdb21b2-5c02-465a-a7d3-4441a2eb07e4 · outbound

This paper cites We quantify the quality of a sharpened model as follows.

Self-Improvement in Language Models: The Sharpening Mechanism We quantify the quality of a sharpened model as follows

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.651219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:11:14.863340Z digest=sha256:12e54bac8dd4ca3a5ee0b6c6a659d742c5425480336601f9eb95978cde659eed

Observation 43c7ee13-b33b-4892-a93f-3f6453d3e602 · outbound

This paper cites Now a⊤Γ−1 t−1a ≤ 1/λ with probability 1, where λ = λmin(Γ0).

Self-Improvement in Language Models: The Sharpening Mechanism Now a⊤Γ−1 t−1a ≤ 1/λ with probability 1, where λ = λmin(Γ0)

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.637243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:11:14.870651Z digest=sha256:0bd96a659166436041290e6fb25fe1743ed7b24f818c2ee1d157f1b3f5698042

Observation 02465288-f85a-49e2-b606-2e43efbb3d91 · outbound

This paper cites Note that unlike the non-adaptive framework, the distribution overmi depends on the underlying instance I with which the algorithm interacts.

Self-Improvement in Language Models: The Sharpening Mechanism Note that unlike the non-adaptive framework, the distribution overmi depends on the underlying instance I with which the algorithm interacts

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.623676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:11:14.875258Z digest=sha256:8e937769dfba4d0a4ad9d5a4df8d41a55030bdc0d93d74fd074ec6dfac207902

Observation 4f13ea1f-2315-4d54-bf2b-dd97e64c56da · outbound

This paper cites an unresolved cited work.

Self-Improvement in Language Models: The Sharpening Mechanism Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:11:15.593265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:11:14.883441Z digest=sha256:a684ba2ecd553c46e4c20684271cd94aead94ec4051b867121688cedaac743fa

Observation f05ee8cb-53e6-40b7-b2b1-2996a591e203 · outbound

This paper cites Initialize: π(1) ← πbase, D(0) ← ∅.

Self-Improvement in Language Models: The Sharpening Mechanism Initialize: π(1) ← πbase, D(0) ← ∅

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.580691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:11:14.887423Z digest=sha256:2e23fb9d4d851acab71c2a6c3165317f34c11b47e139905a48340f17d9e3b2f9

Observation b7cd25d5-9199-46fb-b847-93b6ad41f2f4 · outbound

This paper cites Also suppose thatπ⋆ β ∈ Π where π⋆ β(y | x) ∝ π1+β−1 base (y | x).

Self-Improvement in Language Models: The Sharpening Mechanism Also suppose thatπ⋆ β ∈ Π where π⋆ β(y | x) ∝ π1+β−1 base (y | x)

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:11:15.566237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T00:11:14.891530Z digest=sha256:1875bfd3d8488df9f8b2bb84dc03f53e4eee6e6d1f66255e06c32bae5eaba94d

Observation b57a37d9-567b-42a2-b260-324885c83b69 · outbound

This paper cites Chain of Thought Empowers Transformers to Solve Inherently Serial Problems.

Self-Improvement in Language Models: The Sharpening Mechanism Chain of Thought Empowers Transformers to Solve Inherently Serial Problems

Reference 1973

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.720249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.720249Z digest=sha256:7b6b2c128faeabbe048e93ddba671a2cb3e289f963519a8444953976b12febd3

Observation 100ba0b1-6cce-431d-ad6d-aa9ddcc9f180 · outbound

This paper cites GPT-4 Technical Report.

Self-Improvement in Language Models: The Sharpening Mechanism GPT-4 Technical Report

Reference 1983

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.740875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.740875Z digest=sha256:023db945130afc2d14ef3dcd0d9ecaa1601581800790bfe2915998cf799a5eed

Observation fbbf4766-f474-43ae-a5f0-3e499cdf04bf · outbound

This paper cites Towards a theory of model distillation.

Self-Improvement in Language Models: The Sharpening Mechanism Towards a theory of model distillation

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.631474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.631474Z digest=sha256:8aa5755a54536d4a2e5a91a33b90714ddd957b87a2477abf09276e0ecc0db0ad

Observation 1fa65628-c83b-4a35-a0b8-504f9e4fc9bb · outbound

This paper cites Butterfly Effects of SGD Noise: Error Amplification in Behavior Cloning and Autoregression.

Self-Improvement in Language Models: The Sharpening Mechanism Butterfly Effects of SGD Noise: Error Amplification in Behavior Cloning and Autoregression

Reference 1995

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.627241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.627241Z digest=sha256:7965332f38292bd39499b6ab4519ce7d702893645e0e675b82239f09327e0fc1

Observation f23a1978-c153-4140-af44-9127a7b0e72f · outbound

This paper cites BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling.

Self-Improvement in Language Models: The Sharpening Mechanism BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.681741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.681741Z digest=sha256:cb1229d53c89ded7a258e0c39f77ac94255faff0853e5d110a23ec43496aa690

Observation c5642192-9f09-4e58-9c3c-65b637a8bb1f · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Self-Improvement in Language Models: The Sharpening Mechanism Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.640335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.640335Z digest=sha256:ad45184583eb3f2932d354013758cb5e604ec2451fd737956539ca6d499a63b1

Observation c9870b06-ed88-4bdf-8ac6-0216350d89f4 · outbound

This paper cites In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised Learning.

Self-Improvement in Language Models: The Sharpening Mechanism In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised Learning

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.764758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.764758Z digest=sha256:9f77072ce7ea498855309545eeb0e60fe2c0ba5db2441ac75620baacb50fcde4

Observation 645a2964-83fb-4947-b139-d1a25cb4b978 · outbound

This paper cites PaLM 2 Technical Report.

Self-Improvement in Language Models: The Sharpening Mechanism PaLM 2 Technical Report

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.677287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.677287Z digest=sha256:7e1237444f06054a570059d9531a93261423399d1820eb7265808b08a85cc5e3

Observation 0308ba6c-f907-459d-a12b-5da8b4789a63 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Self-Improvement in Language Models: The Sharpening Mechanism LoRA: Low-Rank Adaptation of Large Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.699627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.699627Z digest=sha256:810f4598befc6f4e58449c00d8655bd11d46f9b4a6652492f9240c8688c6c51c

Observation 76e57ec9-04a5-4467-86ee-0c576f8db4f9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Self-Improvement in Language Models: The Sharpening Mechanism Proximal Policy Optimization Algorithms

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.768719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.768719Z digest=sha256:4a6dd5f83416485e6c92245ad3bcf7bccac7da3e21d51e8a9734be04136525d4

Observation a668f288-be40-42d7-9f6d-5f541cebedfa · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Self-Improvement in Language Models: The Sharpening Mechanism Training Verifiers to Solve Math Word Problems

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.644344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.644344Z digest=sha256:3006dfbe3e0c4319c0b33deb1c186ad5850c3383641633f99e62f176dc500192

Observation 8f21e54c-60e5-48c2-8b3b-6e81dd482993 · outbound

This paper cites Distillation $\approx$ Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network.

Self-Improvement in Language Models: The Sharpening Mechanism Distillation $\approx$ Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.656592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.656592Z digest=sha256:657eaf90e58945790e7b73c444da553846912924a110121ddd0481c6b03aff1c

Observation b84f714c-230d-40f4-acba-d4382f15c153 · outbound

This paper cites The Llama 3 Herd of Models.

Self-Improvement in Language Models: The Sharpening Mechanism The Llama 3 Herd of Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.660901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.660901Z digest=sha256:e49854d4a4befa0246df7402e88182b68897ea8eca96f5f397c6f24aa6e1ca2f

Observation d7948727-c316-488a-b7bc-9d176cc9fae7 · outbound

This paper cites Variational Best-of-N Alignment.

Self-Improvement in Language Models: The Sharpening Mechanism Variational Best-of-N Alignment

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.618718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.618718Z digest=sha256:a3d7a51fb22c50d33c53bbbc427abc15c416ddae7f2dcd2dbab0f5db519b4d8d

Observation 9489302e-5150-489d-8ff0-5969331fd422 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Self-Improvement in Language Models: The Sharpening Mechanism Distilling the Knowledge in a Neural Network

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.695697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.695697Z digest=sha256:47c6c98d613862ebc6a116880c5454f14cca15ccda3449b5f691726db7a1c6b0

Observation b33b96a4-3cf4-41af-84b2-d8c087597fa8 · outbound

This paper cites Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning.

Self-Improvement in Language Models: The Sharpening Mechanism Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.614276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.614276Z digest=sha256:77ad3a9ba0f90e08ae745faae814a69c46ba945f848c93150384a62e04adf04d

Observation 8bbafa47-6fa0-4958-8e87-dfad73cd2737 · outbound

This paper cites Retraining with Predicted Hard Labels Provably Increases Model Accuracy.

Self-Improvement in Language Models: The Sharpening Mechanism Retraining with Predicted Hard Labels Provably Increases Model Accuracy

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.648794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.648794Z digest=sha256:5c5a4604d7b01ab6a2dbe5d7a1ad10b8ea6a1622d326732f80c13a9d58f63e72

Observation 58aa67ff-9efd-4d3b-beea-8186a5fb3bec · outbound

This paper cites Transferring Inductive Biases through Knowledge Distillation.

Self-Improvement in Language Models: The Sharpening Mechanism Transferring Inductive Biases through Knowledge Distillation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.609911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.609911Z digest=sha256:b080722882459aa410ba8dfd38af19fa4ea8a9d58b428bb8367d62a911397c60

Pith citing papers

Observation f5053f7f-ac80-44c3-aa4a-e6ed349a16f9 · inbound

LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds cites this paper.

LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds Self-Improvement in Language Models: The Sharpening Mechanism

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:53:37.179431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:53:37.179431Z digest=sha256:32f2b8239b2b99b03e92c68a0bb8983945f684d2a8d56ed04432428209f19cf7

Observation 715ea1a9-cf4b-474f-b290-e9e02e7b081d · inbound

Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges cites this paper.

Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges Self-Improvement in Language Models: The Sharpening Mechanism

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T14:54:29.183688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:54:29.183688Z digest=sha256:06775778501eb0e8c07f1f86d4172830885b8717057dc003891c95c35c389594

Observation e68c7525-f7e4-48db-906e-ced8dded2667 · inbound

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning cites this paper.

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning Self-Improvement in Language Models: The Sharpening Mechanism

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:58:33.372875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T15:58:33.219451Z digest=sha256:270a2bac824aa585ad53e155eacba6bb01b31399ce814d04e49f43108daca84e

Observation b479d45d-7c4b-4ae8-800e-17e7f52fb058 · inbound

Reinforcing General Reasoning without Verifiers cites this paper.

Reinforcing General Reasoning without Verifiers Self-Improvement in Language Models: The Sharpening Mechanism

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:51.425938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:51.425938Z digest=sha256:d405d24d151290b9c2e4d996d476caf87ccfb069708dae4576edf08af55d632e

Observation 179cc3d2-eb8c-4a8f-a54d-9aa6771e01ee · inbound

Sample Complexity and Representation Ability of Test-time Scaling Paradigms cites this paper.

Sample Complexity and Representation Ability of Test-time Scaling Paradigms Self-Improvement in Language Models: The Sharpening Mechanism

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:35.274446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:35.274446Z digest=sha256:9511515fa9982d41a52149571ad1018d59dd7d826edb12c0eca7c4ee7bf1eea5

Observation 2244849d-c21a-4ff2-82fb-0ec8ec7e670a · inbound

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs cites this paper.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Self-Improvement in Language Models: The Sharpening Mechanism

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.163556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.163556Z digest=sha256:6561efcaa036de99b5a07f7626c9ee173ea0ee21105748ba043a8b69eb9151f1

Observation c99e57b6-dc00-4d09-b4a1-c0c4d0d6ae38 · inbound

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future cites this paper.

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Self-Improvement in Language Models: The Sharpening Mechanism

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:49.580054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:06:49.580054Z digest=sha256:e2acbd732c458fc5bc956c901e8efe973f1879d03535247c795f07a64001e73d

Observation c83422e7-447b-4537-a2bf-d80056733181 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Self-Improvement in Language Models: The Sharpening Mechanism

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.702241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.702241Z digest=sha256:0f896d5149171a2ea54a1031d612c96d8a1b2a43457b8816b81f1fc5f02158c0

Observation d9e3b4c8-2b19-4b7e-8c88-eaf2aad9f39a · inbound

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula cites this paper.

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula Self-Improvement in Language Models: The Sharpening Mechanism

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T02:43:35.074111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:43:35.074111Z digest=sha256:2b91d0dea13c322009612c8c598d82f669d6b0083d311f592ec3162a2fcd2c58

Observation b8022d91-06f9-4e48-80f0-26072928b072 · inbound

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning cites this paper.

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning Self-Improvement in Language Models: The Sharpening Mechanism

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T23:46:32.301737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:46:32.301737Z digest=sha256:01665f411334bc26b8771aaa23f61e63b418aa17fa3cee1d76ee7310e3152c05

Observation 4d7ed3f2-720a-485e-9dfe-8f0a24d7d7ac · inbound

The Role of Generator Access in Autoregressive Post-Training cites this paper.

The Role of Generator Access in Autoregressive Post-Training Self-Improvement in Language Models: The Sharpening Mechanism

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:05:48.110480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T20:16:46.370831Z digest=sha256:8ae489c10c264b4245227d58ca4321ac8b0583907a45e749d53d7094a449f6d0

Observation 0e98ced7-f560-41da-bdbd-eb0e9d7fb3c8 · inbound

Beyond Distribution Sharpening: The Importance of Task Rewards cites this paper.

Beyond Distribution Sharpening: The Importance of Task Rewards Self-Improvement in Language Models: The Sharpening Mechanism

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:08:27.037680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T08:07:14.691463Z digest=sha256:14f9ce8b9457b7c4ed2fd5d8d1c27127d4613a540fdd71cf831e5aaff9f94b74

Observation 73f72c26-913e-4e00-9097-f07e002fca28 · inbound

On the Generalization Gap in Self-Evolving Language Model Reasoning cites this paper.

On the Generalization Gap in Self-Evolving Language Model Reasoning Self-Improvement in Language Models: The Sharpening Mechanism

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.862765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:611efcb68bf871c10e9929474d155c185c376b8873d0334a0363265f85c1f7cb

Observation 51af3dc1-1f41-4b67-8962-6fdb0545fc90 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Self-Improvement in Language Models: The Sharpening Mechanism

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.665777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:dbfc0e89fd079fc26c5134c23a95194fc2c8c1abce65e824e1a03ede92a8fbc5

Observation 70c40a63-4681-4129-b016-7a5f5524e943 · inbound

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning cites this paper.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Self-Improvement in Language Models: The Sharpening Mechanism

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:26.274337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:a3de33e0545a166494f83e8b1c156d1e7bd99477b39ab5cc8508313f6f364075

Observation fc408979-3023-4995-b367-ff8699d39f6a · inbound

Select and Improve: Understanding the Mechanics of Post-Training for Reasoning cites this paper.

Select and Improve: Understanding the Mechanics of Post-Training for Reasoning Self-Improvement in Language Models: The Sharpening Mechanism

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:08:21.734036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T07:14:47.302277Z digest=sha256:f276743580ba9989458b33d08015bf4987c408d64c1a804905540c3791cc8c04