Pith. sign in

Paper Citation Record · LEDGER

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

As of 13 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 18 inbound Pith citation observations for arXiv:2412.02674.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02674 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:15:28.793471Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:54:29.270448Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.476979Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a582408f-2d25-4c20-9cc5-d3096012ef66 · outbound

This paper cites GPT-4 Technical Report.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.483490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.483490Z digest=sha256:6121af27a228fbba93175a371efd3a14e740e05f33053d022a9a092a7570e53e

Observation 1006b770-934f-41e5-9ed3-7cad476dce90 · outbound

This paper cites Critique-out-Loud Reward Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Critique-out-Loud Reward Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.498365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.498365Z digest=sha256:998989565ed68f5382551425b600dd598da8c8c1c23b4c9df9cca790034a2c1d

Observation 3e86cb86-5877-4f2d-a9bc-0c27bb0838af · outbound

This paper cites Qwen Technical Report.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.503131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.503131Z digest=sha256:9f0611e5ff29c31ef8421b8338a6010579164d6d872e1d389a9d19501109434e

Observation 4e21ec20-274a-43bf-9623-e75a19af96ee · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Constitutional AI: Harmlessness from AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.508973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.508973Z digest=sha256:44ce1382e71e7d5f73572f8350bc26190f0c5e15a212ba282dfb7c4fcc9e1d23

Observation 8cecfc31-dcfe-4552-9b88-b57bf309ce46 · outbound

This paper cites On the Stability of Iterative Retraining of Generative Models on their own Data.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models On the Stability of Iterative Retraining of Generative Models on their own Data

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.525342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.525342Z digest=sha256:cbf6d23d34f50b303b9f811f08ba467a29cc4adfb64f888f81abccdb3c6cee3a

Observation d2c8edaf-45f1-49d0-9a5a-abaa4a59eb3e · outbound

This paper cites Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training Loop.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training Loop

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.530737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.530737Z digest=sha256:742b664e1add86caed9d90933d5ad6b174679afe65ed8e535d91a67db1584e76

Observation 7a626e7b-1941-48ca-aa31-42cb9b82f236 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.535758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.535758Z digest=sha256:a4854a713ff0cd4650fce5738559afb651d152ee5c8bb2d412b175760c609a4c

Observation 2eabb5c7-befa-4308-9e34-e7a3f7b09e84 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.539819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.539819Z digest=sha256:677386e6956586d660958c9629a9cab50ef4f724f25cfa31793081929149b963

Observation 4b86df85-5d4a-4114-bf8b-29e86fa1edcd · outbound

This paper cites Learning to Generate Better Than Your LLM.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Learning to Generate Better Than Your LLM

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.551006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.551006Z digest=sha256:5a40c5c4ea9220e1e89d9d227484d0fdd4e91bc53c17c244a57cd85f26a881e1

Observation d31c9d90-01d5-4efa-a92b-d52e383a60f7 · outbound

This paper cites Teaching Large Language Models to Self-Debug.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Teaching Large Language Models to Self-Debug

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.554478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.554478Z digest=sha256:8ed46e686422dec322ebae8edeaf4183344208cddd9d5780ab357d48b6009b55

Observation 49dffea3-7aed-42a6-aa7f-9ed251e93a59 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.559268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.559268Z digest=sha256:18f551d5d73c6a99f6fbd051e8daccdaef94420eb20c48e8bfc8a9e955dab35c

Observation e7f030f7-35dc-4570-97ee-63008b05c85a · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Can Large Language Models Be an Alternative to Human Evaluations?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.563407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.563407Z digest=sha256:200b5c6fb9b95589cd7963b31ee3f2d9e113602021d31cd10e36d67909e6bd7c

Observation 97831683-ac83-4a52-96f8-61f4818844d7 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.See https://vicuna.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.See https://vicuna

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.567793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.567793Z digest=sha256:a12aece25156627246b4a03cc692b137ce732c1cf91c8b89643b1ad315307d5d

Observation 9d8a2194-e9ad-46f9-9fa4-ba7564fe7149 · outbound

This paper cites Model Collapse Demystified: The Case of Regression.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Model Collapse Demystified: The Case of Regression

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.579987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.579987Z digest=sha256:9c5d589a134d81600458335ed05e4886aea2f3a1edebefb9e45add37871732b2

Observation bb192b01-f8f8-47b4-831f-9bd212ebdeaf · outbound

This paper cites Training on the Test Task Confounds Evaluation and Emergence.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Training on the Test Task Confounds Evaluation and Emergence

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.585346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.585346Z digest=sha256:a7c6754bd873c46d877efdfae9afd97a1741cc087d5bae66623b445b183dcf36

Observation 28e71f36-6cbb-4fcd-998d-7ad8038de913 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models RLHF Workflow: From Reward Modeling to Online RLHF

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.589443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.589443Z digest=sha256:b0f47df3abd04435185b98948c5d9aba775ec073c4ee2643359a1aaf02f352fb

Observation 20974d57-4de3-4937-bda5-070ac236be73 · outbound

This paper cites The Llama 3 Herd of Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.593584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.593584Z digest=sha256:5f7365e3b3fb2a39c4a7f7364f6bbb459d285f9fad9b2af95d0db5a14c51497b

Observation b14dd521-d354-43a2-8a13-bd9516519e42 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.598804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.598804Z digest=sha256:42ba638a1a51bef202f4171af18e5f1ae9e6bdd7ca9c70f3ce21156a66d2ac83

Observation ed77794f-ce2e-484c-8aff-76c5130c7452 · outbound

This paper cites Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Henry Sleight, John Hughes, Tomasz Korbak, Rajashree Agrawal, Dhruv Pai, Andrey Gromov, et al.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Henry Sleight, John Hughes, Tomasz Korbak, Rajashree Agrawal, Dhruv Pai, Andrey Gromov, et al

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.603090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.603090Z digest=sha256:161bf9fb05612ca640973ba68b0498070d31054aa639cccfcbcd8abc571c0fa9

Observation 95e0dace-5d23-469c-a1dd-ff4d14cc32ae · outbound

This paper cites Self-Correcting Self-Consuming Loops for Generative Model Training.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Self-Correcting Self-Consuming Loops for Generative Model Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.607337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.607337Z digest=sha256:ab54a54aaf62864fff71e54ffd7f6691d63ab13091b236f187636003ade57e6d

Observation c7419b19-0da9-4ea2-9723-92b4175c42b8 · outbound

This paper cites Textbooks Are All You Need.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Textbooks Are All You Need

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.611461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.611461Z digest=sha256:0f6d444d65332bd3bc0abaf29aa07295c68768d9fa852647411ca72ac628b187

Observation 71f6d78d-5621-4854-a1ab-cd6d43502736 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.615413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.615413Z digest=sha256:3830fee2d651bd2e88e4dbefaa5d7ee6d2e91c2a1689ea4b2297295c43836aad

Observation 99c1ad99-4373-41cc-84e8-58185689d9f1 · outbound

This paper cites Scaling laws for downstream task performance of large language models.arXiv preprint arXiv:2402.04177,.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Scaling laws for downstream task performance of large language models.arXiv preprint arXiv:2402.04177,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.637902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.637902Z digest=sha256:d1d88ad9019eeb6bcf7756933b0e458373df6542a5531837c0e41ada67d9b83b

Observation 4ba54d39-b856-4fbb-985f-6dd250ead3cd · outbound

This paper cites SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.641881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.641881Z digest=sha256:6e4d34f74fa871566dc57940138bac2bfcba3447be459661ccd82174665c017e

Observation f6b417b9-0794-4870-8504-6ea296da7493 · outbound

This paper cites Scaling Laws for Neural Language Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Scaling Laws for Neural Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.649181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.649181Z digest=sha256:dcf6d751353ae221af34dbea6e2744f71612b92306eb6ac9f5dbbf35d4c8c18b

Observation 75b51781-c29e-4c73-9188-0be6463c098a · outbound

This paper cites Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.658554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.658554Z digest=sha256:c7f919776c43f2261bba05a38bb868681ec912d3658409fabff086a54e7441cf

Observation 8b5a27ef-cb02-45dc-a395-dbd640ca7882 · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Textbooks Are All You Need II: phi-1.5 technical report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.662986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.662986Z digest=sha256:4638bc1cbd4b6bdbc7fa584e18ad831a0608bf2be9532c069de3bb4b3e8e7585

Observation 393ae529-41ec-4015-aede-d8de3469f8a7 · outbound

This paper cites I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.667466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.667466Z digest=sha256:46cd1841da702b11ab739b74047906d484bb6738e4e7f393adafee318d261152

Observation ff6abce8-9a92-4643-ae6c-6ceb759e6950 · outbound

This paper cites Let's Verify Step by Step.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Let's Verify Step by Step

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.672176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.672176Z digest=sha256:38c430e1b18912e0933bcb42f0e1c886a3c3c1bd595352c01aba92ddd88a51cd

Observation 8e461d49-9e9e-40c5-a340-1ab0792071c5 · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Statistical Rejection Sampling Improves Preference Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.677019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.677019Z digest=sha256:f77b903e29237244920c5a0d827e06a2945aa2b792fe91550855dae2a801d048

Observation 11592053-4d9d-43a4-880a-94a3c6b9652f · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.682560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.682560Z digest=sha256:173341b4ee678facec1749254d16882717f2fd08b14ca8bcb729723dada7bbf4

Observation 3f23aae0-c249-476c-bafe-64b6086acdcb · outbound

This paper cites Combining Generative Artificial Intelligence (AI) and the Internet: Heading towards Evolution or Degradation?.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Combining Generative Artificial Intelligence (AI) and the Internet: Heading towards Evolution or Degradation?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.686900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.686900Z digest=sha256:9359b4ecaf13614c914b5490fd8a281e8141e0b76319078afe71f546977ce714

Observation 4ff5c2a0-a54e-467e-a1bf-9bb59e954622 · outbound

This paper cites Iterative Reasoning Preference Optimization.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Iterative Reasoning Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.691595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.691595Z digest=sha256:43c534a07167985c095e7680e41962b95b678b548dffac96a7d06d12505a9731

Observation 2969d4bf-fb2b-480c-aab9-9f3cccf6ecd0 · outbound

This paper cites Observational Scaling Laws and the Predictability of Language Model Performance.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Observational Scaling Laws and the Predictability of Language Model Performance

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.695929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.695929Z digest=sha256:0c6b57675e989b62134b7862d1f841dcc0bd267f094d75c91943c1d5a9320ee6

Observation bd27e811-fe89-4086-94dc-46179857da5f · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.701530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.701530Z digest=sha256:9a38f4dbdc62bc2f69b1e51ba8a983ea0427c14504eec92bf09d054dd1fdda3c

Observation f2b82ff6-6639-41aa-bc6c-a542e42df54e · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.707967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.707967Z digest=sha256:883bac97eef04a7e41c953bcd02efddeeb0c3222070fe34d27e3f3e08a2b332a

Observation 0e46a5ea-b307-4b1d-bed4-778114a6a13d · outbound

This paper cites SuperCLUE: A Comprehensive Chinese Large Language Model Benchmark.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models SuperCLUE: A Comprehensive Chinese Large Language Model Benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.712803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.712803Z digest=sha256:fc754460b76d02b9875255b0862941f7579b25577c46323af8ad0ff73f19f2cd

Observation 31de1885-8f17-4558-ab61-91933eaa087e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.722455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.722455Z digest=sha256:c1b5c7505ab0ae05f26de6143e4244d159caab80cd8244d1ebaf351326f8653a

Observation 2354056d-55f4-4c98-a313-2a42c94dcf98 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.726482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.726482Z digest=sha256:71e2dc4d9f31ef488fa22a5bc64d11fc4e01a8fe922105056fcb85d67a7ab5f3

Observation b0d9e045-3538-42cf-9d5c-60c33672daa0 · outbound

This paper cites Self-Taught Evaluators.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Self-Taught Evaluators

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.731425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.731425Z digest=sha256:641fd93d0da44c2956fe850499e52fa2bb0edcd786613687568cb3e4301279f9

Observation 1fb67c28-3388-46d9-b612-cde357028e3a · outbound

This paper cites From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.742927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.742927Z digest=sha256:5044e2347c8b4e6ffe3c59323bcaa0b725abf58179e1e7b389e26cc9db2fc815

Observation 98b2924e-d1d3-456d-bba8-9439cf7f9261 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.747381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.747381Z digest=sha256:a832f9e1f7bf1a130ff7a6e6393ed77ae8bd8f4a424146616feb45806330df54

Observation 9b3c2369-6c8b-499e-866a-88410e6e5615 · outbound

This paper cites Qwen2 Technical Report.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Qwen2 Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.751851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.751851Z digest=sha256:cfad9ea98a95aef79f95683d1dd0c7d1d9a00d963e782a647b6f54a278b7c32c

Observation 4dc7458b-b4c8-4da7-b51e-d960c619ba71 · outbound

This paper cites Genie: Achieving Human Parity in Content-Grounded Datasets Generation.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Genie: Achieving Human Parity in Content-Grounded Datasets Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.755772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.755772Z digest=sha256:54dcf269af42ce239eccd9a389e70cb7e4a19d0de0b49527ee8a34b334d9ed6f

Observation 629bf445-2d55-4cc5-83a5-7cabcce69417 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Yi: Open Foundation Models by 01.AI

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.760232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.760232Z digest=sha256:4bd151bb0f201f5fa56927b746636d77fddcb987fa1b0ea6cf2dc433e8eb72d6

Observation 0489c865-ea9d-4798-8599-1d8488769214 · outbound

This paper cites Self-Rewarding Language Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Self-Rewarding Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.764847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.764847Z digest=sha256:4625eb946624e882322dc3e705a293bee45ac566d546291a6dcb0bdc6c666493

Observation d00dfe03-37cc-417d-b210-f49f2e73307d · outbound

This paper cites Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.769802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.769802Z digest=sha256:8c572ddb4a03391fd75aad9cb6e3029500bb4e3cf67e46d9a6ad8d9879e8a45e

Observation 143828bb-371f-43d0-8ba7-0e6518bb23a7 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.775413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.775413Z digest=sha256:77931d17e65c773b41b9e254e18492e78ab7128d97bb629e3e3cb88b217fee60

Observation e47622b5-ed4e-465e-a80b-cfeee8c31b99 · outbound

This paper cites an unresolved cited work.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:15:29.846814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:15:28.783877Z digest=sha256:59577d084188c806df03ef67bcab818cd256012b58be23ee04e57ff0ce0f7915

Observation 3c1fe777-9ef9-46c5-9fea-6da85f5d4fbb · outbound

This paper cites an unresolved cited work.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:15:29.834256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:15:28.788743Z digest=sha256:35ce364eb8724399c7e83e801185890f77c06eb5963ebe8942e1fb393fc1909f

Observation 9b431a19-324a-48eb-af09-6baf75bbfade · outbound

This paper cites The answer is ANSWER.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models The answer is ANSWER

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:15:29.821481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:15:28.793471Z digest=sha256:5df497d7e6fe546b5da27c8c71cf14d0d110d074ea1df366a19a0503275110d5

Observation 4bb79ffe-ae45-4439-87d8-7c3a80a411ff · outbound

This paper cites Learning How Hard to Think: Input-Adaptive Allocation of LM Computation.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Learning How Hard to Think: Input-Adaptive Allocation of LM Computation

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.575938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.575938Z digest=sha256:3ad65e2eb8f185baf7e7e2ab3ab78c08b7198a715677d4f5e2ce31c5299c6491

Observation 1faa3f38-df3a-46fd-b810-1c80e0fae732 · outbound

This paper cites From Lists to Emojis: How Format Bias Affects Model Alignment.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models From Lists to Emojis: How Format Bias Affects Model Alignment

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.779769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.779769Z digest=sha256:94da9f9bfd414b60d9e5e37c6820fd2efb3ccc0db9cf038eddbb226c6adb9afc

Observation 63d336fa-0c4c-4805-8f21-c9204165c948 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Qwen2.5-Coder Technical Report

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.633561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.633561Z digest=sha256:a688cb69ad3d5564a753e9969a1b94621391440e1918a999f890a18418fa8ded

Observation 1b028eb2-10ef-4b77-ba70-6096fc73c12b · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.619469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.619469Z digest=sha256:cbe6a6a992d873d3e0ea36dc09bebff3b464fa11fb8b4745ad6f9bb76d5ecbf8

Observation 792c3e5f-31ff-46e0-944c-8e18e8bec09d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.571867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.571867Z digest=sha256:055c83e59e419cb2426ac5bd14760ca2b7fa26b1897b4c53ed08ea609bb5fa98

Observation 1064790f-6c2e-435d-a6bf-1a55ab4286f9 · outbound

This paper cites RankGen: Improving Text Generation with Large Ranking Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models RankGen: Improving Text Generation with Large Ranking Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.654641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.654641Z digest=sha256:ee51191a262e22887738d425dc012a3997c1e139465cbc251c4b819a996b56c5

Observation 243fad25-7fbf-4e43-ad28-32c556fbd12c · outbound

This paper cites Scaling Laws for Transfer.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Scaling Laws for Transfer

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.624117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.624117Z digest=sha256:af708807c59219d4bfbb00d62fa335ddcb0e1c81e70c2dc8420171c074aeac43

Observation 96760804-8764-4cbe-b70a-66b31af334d1 · outbound

This paper cites Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.520053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.520053Z digest=sha256:273542ff15cf909712cc3073a2928dd006ff7d486c824110bb418f7f987534d0

Observation a423a56c-b440-4db2-9426-3ce9500910d7 · outbound

This paper cites Nemotron-4 340B Technical Report.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Nemotron-4 340B Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.488650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.488650Z digest=sha256:b1f5b898b18b9fcf53d2400f4306a77b5688f3bae4a50eea78662db7e2f496b5

Observation c3c60f7e-5b8b-448f-81b5-2d195bebe09b · outbound

This paper cites Self-Consuming Generative Models Go MAD.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Self-Consuming Generative Models Go MAD

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.493309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.493309Z digest=sha256:e5a86c24dab21ed4aad0e72de8e3cf28531588e515971732860c2cbfaeac12cd

Observation 1480e59d-4c59-483d-8432-dd0882773667 · outbound

This paper cites Large Language Models Can Self-Improve.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Large Language Models Can Self-Improve

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.629098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.629098Z digest=sha256:7cdde1b8bdb4a413dbaca3284ef188a7d15ab7c3602f20dff094b2f13fb8c64c

Pith citing papers

Observation 574ba37f-3705-454b-b24d-4ca3b9a0819a · inbound

Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges cites this paper.

Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T14:54:29.270448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:54:29.270448Z digest=sha256:f53a968f4ee05ef06db287deb5d0ea7a5485c6f84b3c900b56ba1df31777c322

Observation e61b5bbf-b799-4ccf-ab58-29688b857d9b · inbound

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning cites this paper.

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:22.812009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:28:22.812009Z digest=sha256:95d3f813832d0e0142b5511aef297a81650cb3a5eded797ba4c1f65eee91c92c

Observation 1c0905cb-9aa6-41d5-ac77-19d69cfcfcb6 · inbound

Sample Complexity and Representation Ability of Test-time Scaling Paradigms cites this paper.

Sample Complexity and Representation Ability of Test-time Scaling Paradigms Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:38.336495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:38.336495Z digest=sha256:0c7092ed81e2516661561534d1c22f7fddf5ec299da1858cfe5911fef3f0e581

Observation b760d7ee-a2d3-4a3e-a999-b4747dad6391 · inbound

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs cites this paper.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.400160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.400160Z digest=sha256:a80c50fc7e9bfe7b4108b97171c18369e1b16a33e619698238c21b3b26115b63

Observation ecfaabd5-011e-4fb5-8f41-d6591eece8e5 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:36.497124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:36.497124Z digest=sha256:53fd89204985775c4788eea523634407b3009e500113221214a9b9cc7fd4651a

Observation eb7f58e0-4da1-417f-8b70-2b500455401a · inbound

Because we have LLMs, we Can and Should Pursue Agentic Interpretability cites this paper.

Because we have LLMs, we Can and Should Pursue Agentic Interpretability Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T01:03:21.917517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:03:21.917517Z digest=sha256:5db337666a8bfde0a49ca629713de6f10e2929073b6cfe122c2d3531d073b6ff

Observation e2a5c216-4747-45e8-91ec-72ca80aed3cd · inbound

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability cites this paper.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.251369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.251369Z digest=sha256:d7c649b317e32709c1bf1d57bf7520c0d3401e556b80e61d78a219f32097f18f

Observation d1eb8d0e-a29f-4d6d-8e30-6db51c75e167 · inbound

Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online cites this paper.

Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:52.207779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:52:52.207779Z digest=sha256:1a8020bfeb503db385b90a1024835bb6c7b599b1c1483e28736f618fb0a225a3

Observation 73d87d62-3600-436e-8f4c-b412c9fef2d6 · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.556101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.556101Z digest=sha256:02ab89c83cb4f4ec415f1e73aa4aaf23d31096cedbd9e8a38af01a1a0a5ffae6

Observation 88fa9f2e-85a3-466f-8bbc-a2d61ed641f9 · inbound

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula cites this paper.

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T02:43:36.764821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:43:36.764821Z digest=sha256:6b80b39b8939768ac592994e95254df74eeea0019cf39a9550e19953dc07f4c0

Observation 3c17bca8-58cc-4883-b5af-fff3617f1794 · inbound

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents cites this paper.

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:25:35.849997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T12:22:13.551709Z digest=sha256:31fe2e98dd9873ed772b34bcd6e80652943760ee8b25897516add349437e3718

Observation d4970fb6-29dc-4158-9ff4-388f7f142f42 · inbound

AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems cites this paper.

AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.591039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T07:02:02.992871Z digest=sha256:55e97bdc25b6367eaf514e6aee8508b573a41efa75ff876ea6c0c566167860b4

Observation fc296c37-263e-4b37-918d-b29e828bd96c · inbound

Tree of Concepts: Interpretable Continual Learners in Non-Stationary Clinical Domains cites this paper.

Tree of Concepts: Interpretable Continual Learners in Non-Stationary Clinical Domains Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:46:37.231503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T06:44:44.353815Z digest=sha256:15528c0bdd394f4676adc9a9787d54b4704de46b58cb9f5d248c16f5fc7c51c6

Observation caec8726-3d6a-4f00-a7cb-0274bf57bdcc · inbound

Annotations Mitigate Post-Training Mode Collapse cites this paper.

Annotations Mitigate Post-Training Mode Collapse Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:46:51.283999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-12T03:58:11.179607Z digest=sha256:59bef885ba7ccb6fbcd08f9f9a057e3d70b8bc4f33dde55d8908db4f920ae8de

Observation f5591126-1c6c-4e21-b113-78511853ae51 · inbound

On the Generalization Gap in Self-Evolving Language Model Reasoning cites this paper.

On the Generalization Gap in Self-Evolving Language Model Reasoning Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.928244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:a4831483e7c8e8c183b753e2bc91e43cecfb124406b4d845ed65febe1936ec37

Observation 95b3bb05-48ac-4514-b2bf-632c27e7fbed · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.478449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:7a84a247a56340419958fec0ce894358beac3003454e89008bf5eb0bd035b689

Observation 6d689bb3-6ea7-4f52-9251-8247dbb3f52d · inbound

Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models cites this paper.

Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:05:42.232371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-07-01T04:44:56.520156Z digest=sha256:560c207d97e7d7674a2f787591adb056af73dccb7566678141e2635b69193120

Observation 4a151dce-1fbe-4ea6-af68-2b1378e91128 · inbound

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets cites this paper.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:58.750607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:58.750607Z digest=sha256:d2a8b7f02e3be2ab68f9e1886260baac57ef4f074060078817cbfe10b9193268