Pith. sign in

Paper Citation Record · LEDGER

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

As of 14 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 18 inbound Pith citation observations for arXiv:2412.02674.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02674 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:15:28.793471Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:54:29.270448Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.476979Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a582408f-2d25-4c20-9cc5-d3096012ef66 · outbound

This paper cites GPT-4 Technical Report.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.483490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.483490Z digest=sha256:4ec14ef99eb0e56e3eeedd92533be68eee19541113f5f0c5bc9f0e0d5d67de66

Observation 1006b770-934f-41e5-9ed3-7cad476dce90 · outbound

This paper cites Critique-out-Loud Reward Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Critique-out-Loud Reward Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.498365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.498365Z digest=sha256:3b35632199c89cc3e13d6e82fc0ce9f0bd358e79c7c2eb0a2400a922c9bab054

Observation 3e86cb86-5877-4f2d-a9bc-0c27bb0838af · outbound

This paper cites Qwen Technical Report.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.503131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.503131Z digest=sha256:4ac227bf5e109831c554e9c1d3e2af6440fae0809093c7900561b8930dd30ec7

Observation 4e21ec20-274a-43bf-9623-e75a19af96ee · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Constitutional AI: Harmlessness from AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.508973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.508973Z digest=sha256:49a09729ee15c2afcbd1b43554b6a55989f03dc574e34ebd47e93c3e350eaaec

Observation 8cecfc31-dcfe-4552-9b88-b57bf309ce46 · outbound

This paper cites On the Stability of Iterative Retraining of Generative Models on their own Data.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models On the Stability of Iterative Retraining of Generative Models on their own Data

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.525342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.525342Z digest=sha256:57d9dc51a7b950470f7a2843f554f8e31a436423fda83e69fe611da34ac736c9

Observation d2c8edaf-45f1-49d0-9a5a-abaa4a59eb3e · outbound

This paper cites Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training Loop.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training Loop

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.530737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.530737Z digest=sha256:64507ed809f7384a924c2f4c96d5a25f46dbcd809394a6e67670ce08e99cdcde

Observation 7a626e7b-1941-48ca-aa31-42cb9b82f236 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.535758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.535758Z digest=sha256:6dd4c5265dbd4427c6eaea659517acbbcc8d5bccfe2520e7d988dad7ad4cc2d8

Observation 2eabb5c7-befa-4308-9e34-e7a3f7b09e84 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.539819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.539819Z digest=sha256:797e8fb825750201170f08b42fc73983dfe9da39d51905a6b5685a7e13b135bd

Observation 4b86df85-5d4a-4114-bf8b-29e86fa1edcd · outbound

This paper cites Learning to Generate Better Than Your LLM.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Learning to Generate Better Than Your LLM

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.551006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.551006Z digest=sha256:0ba713ecdf9ababbc449b98734c5ef8a4ca18de2b6a79e837d06b9799f9cfe72

Observation d31c9d90-01d5-4efa-a92b-d52e383a60f7 · outbound

This paper cites Teaching Large Language Models to Self-Debug.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Teaching Large Language Models to Self-Debug

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.554478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.554478Z digest=sha256:c7ff6cf680d7dd83a0faf023e50f3b44a44259346e69a73955e5a11c4ea26805

Observation 49dffea3-7aed-42a6-aa7f-9ed251e93a59 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.559268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.559268Z digest=sha256:76e6234ca729613f07016cbe8ba943d81a7af2f2885550f3567374b99ed83e4a

Observation e7f030f7-35dc-4570-97ee-63008b05c85a · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Can Large Language Models Be an Alternative to Human Evaluations?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.563407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.563407Z digest=sha256:0ac81a233bea6fc36013f24eb0e1184fd603309e8880426ee99de1166b461c6e

Observation 97831683-ac83-4a52-96f8-61f4818844d7 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.See https://vicuna.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.See https://vicuna

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.567793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.567793Z digest=sha256:b101050b9f9a9910b13e640f4b46dba170d09a18801b8371e16b8bba8812c37b

Observation 9d8a2194-e9ad-46f9-9fa4-ba7564fe7149 · outbound

This paper cites Model Collapse Demystified: The Case of Regression.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Model Collapse Demystified: The Case of Regression

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.579987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.579987Z digest=sha256:0267548f485cc8bd332e413ae0fc5ee15ea999f00cbef108720a452ecb90aa5e

Observation bb192b01-f8f8-47b4-831f-9bd212ebdeaf · outbound

This paper cites Training on the Test Task Confounds Evaluation and Emergence.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Training on the Test Task Confounds Evaluation and Emergence

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.585346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.585346Z digest=sha256:268a805cf702cc05fb3ea149f3946c7ae456a32ac25d8bec233a648cc1b3eaa2

Observation 28e71f36-6cbb-4fcd-998d-7ad8038de913 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models RLHF Workflow: From Reward Modeling to Online RLHF

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.589443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.589443Z digest=sha256:810d73278c45c396bdd40e072540af65a10d90b634aa882af6b10a33cc2da217

Observation 20974d57-4de3-4937-bda5-070ac236be73 · outbound

This paper cites The Llama 3 Herd of Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.593584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.593584Z digest=sha256:c607d09248ed0c83c503db41d1c9fafab881f4a22d2fa249d59bf200b171ec0f

Observation b14dd521-d354-43a2-8a13-bd9516519e42 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.598804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.598804Z digest=sha256:6fbc40d6d58abbca728c0c67c64038e184edced5ea39ec9563a8f47c782b31a9

Observation ed77794f-ce2e-484c-8aff-76c5130c7452 · outbound

This paper cites Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Henry Sleight, John Hughes, Tomasz Korbak, Rajashree Agrawal, Dhruv Pai, Andrey Gromov, et al.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Henry Sleight, John Hughes, Tomasz Korbak, Rajashree Agrawal, Dhruv Pai, Andrey Gromov, et al

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.603090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.603090Z digest=sha256:010336d7fa0a6c1c1aa4a8630e552c4a115c0878fa74b6c92b3ec587e9598990

Observation 95e0dace-5d23-469c-a1dd-ff4d14cc32ae · outbound

This paper cites Self-Correcting Self-Consuming Loops for Generative Model Training.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Self-Correcting Self-Consuming Loops for Generative Model Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.607337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.607337Z digest=sha256:d4528240af1cc874c53a5cfa452f7437e00bfb07e83b24fec077cd42993d31e2

Observation c7419b19-0da9-4ea2-9723-92b4175c42b8 · outbound

This paper cites Textbooks Are All You Need.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Textbooks Are All You Need

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.611461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.611461Z digest=sha256:b1281f6e21240fe652b34bdd6b85aa5d4007fbeb1a4a215a2ba955773ed706f2

Observation 71f6d78d-5621-4854-a1ab-cd6d43502736 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.615413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.615413Z digest=sha256:c3fd885705ef0a8ddd710597dd32d499b00d3f7c83cacdef513aa021184a90df

Observation 99c1ad99-4373-41cc-84e8-58185689d9f1 · outbound

This paper cites Scaling laws for downstream task performance of large language models.arXiv preprint arXiv:2402.04177,.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Scaling laws for downstream task performance of large language models.arXiv preprint arXiv:2402.04177,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.637902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.637902Z digest=sha256:5f7bb045f7a56c74da073518546b75f6cfa517f16b7c296ad4211ffabf6b68ad

Observation 4ba54d39-b856-4fbb-985f-6dd250ead3cd · outbound

This paper cites SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.641881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.641881Z digest=sha256:c0dcb4eb92bcdd98126d510ca0862e90205f174c64e552279059f1f16948c390

Observation f6b417b9-0794-4870-8504-6ea296da7493 · outbound

This paper cites Scaling Laws for Neural Language Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Scaling Laws for Neural Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.649181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.649181Z digest=sha256:7c4ae334672f3813ba1bb7745f26bccbd24a5fd6c2759cb7b1e71feb235ce1cc

Observation 75b51781-c29e-4c73-9188-0be6463c098a · outbound

This paper cites Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.658554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.658554Z digest=sha256:b0fde600e97dc291642a92538234edbbf044f7124536d1eb941f7b4aebc1c33c

Observation 8b5a27ef-cb02-45dc-a395-dbd640ca7882 · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Textbooks Are All You Need II: phi-1.5 technical report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.662986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.662986Z digest=sha256:98fbb5849fd90fdecc10ac66fd9bb11eb1b99fd71402c45713d89d03dc6c3b81

Observation 393ae529-41ec-4015-aede-d8de3469f8a7 · outbound

This paper cites I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.667466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.667466Z digest=sha256:e26644af3d2e363e0a75ec7cc21851e8500ec81b7797ffeceffcfd0b3ee74dff

Observation ff6abce8-9a92-4643-ae6c-6ceb759e6950 · outbound

This paper cites Let's Verify Step by Step.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Let's Verify Step by Step

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.672176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.672176Z digest=sha256:3eb2c87ce0bfafe82932ca998ffa02ee2c8726d2e6791371f1402c728a6782c5

Observation 8e461d49-9e9e-40c5-a340-1ab0792071c5 · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Statistical Rejection Sampling Improves Preference Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.677019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.677019Z digest=sha256:b2b3833712f66b0a6b3387c3f79dd215791760c5f1c20a84203797b65bd00daf

Observation 11592053-4d9d-43a4-880a-94a3c6b9652f · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.682560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.682560Z digest=sha256:678053ca9188943a85b8c33bc01c9a9ed78b8b1e7aa992b7b57371e830f4da21

Observation 3f23aae0-c249-476c-bafe-64b6086acdcb · outbound

This paper cites Combining Generative Artificial Intelligence (AI) and the Internet: Heading towards Evolution or Degradation?.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Combining Generative Artificial Intelligence (AI) and the Internet: Heading towards Evolution or Degradation?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.686900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.686900Z digest=sha256:0500ff9bc801412739542aa69f3b3333b7ac517dafd8c99f401416a76448eb4d

Observation 4ff5c2a0-a54e-467e-a1bf-9bb59e954622 · outbound

This paper cites Iterative Reasoning Preference Optimization.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Iterative Reasoning Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.691595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.691595Z digest=sha256:5bd0e4d04ad0e279c1c1e4b2c6079fe24a70cd9f46e8e19fad31fe1e476a61b1

Observation 2969d4bf-fb2b-480c-aab9-9f3cccf6ecd0 · outbound

This paper cites Observational Scaling Laws and the Predictability of Language Model Performance.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Observational Scaling Laws and the Predictability of Language Model Performance

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.695929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.695929Z digest=sha256:2bf10a5477d84e9bb10a9307b22a62b9002dad7768d41cf1411e27f277d817ed

Observation bd27e811-fe89-4086-94dc-46179857da5f · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.701530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.701530Z digest=sha256:1cb16db664dea5c844a736bda729ef5ef72c2e6f4e5e9d68a9a4f70a8e1652ce

Observation f2b82ff6-6639-41aa-bc6c-a542e42df54e · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.707967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.707967Z digest=sha256:d8a1973d18fd9ac7342091c41b9adef6043c8e24976bf9b75e4e2b538d9e3f15

Observation 0e46a5ea-b307-4b1d-bed4-778114a6a13d · outbound

This paper cites SuperCLUE: A Comprehensive Chinese Large Language Model Benchmark.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models SuperCLUE: A Comprehensive Chinese Large Language Model Benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.712803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.712803Z digest=sha256:6bbe5ad168a45c5e7a164c2240f3b98089700fc09e8d0ce2a113af7097e8c149

Observation 31de1885-8f17-4558-ab61-91933eaa087e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.722455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.722455Z digest=sha256:2c9be25cbddea707b63583c87c9c23cf2345be0fb9dabefef13ebf4b48861e06

Observation 2354056d-55f4-4c98-a313-2a42c94dcf98 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.726482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.726482Z digest=sha256:89f2f550d40a7a9f5eae02a66d6f4ba07b9f66f617a27b163f700ad7adf1448d

Observation b0d9e045-3538-42cf-9d5c-60c33672daa0 · outbound

This paper cites Self-Taught Evaluators.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Self-Taught Evaluators

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.731425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.731425Z digest=sha256:71f6f3300dfab75ef42be64b4f53033db8f932ab55de15c89716da506ad93a63

Observation 1fb67c28-3388-46d9-b612-cde357028e3a · outbound

This paper cites From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.742927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.742927Z digest=sha256:34893ce05e8f695a534d8a3ff5e5fd81b3040d586c4523ea03e9d3991093bb22

Observation 98b2924e-d1d3-456d-bba8-9439cf7f9261 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.747381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.747381Z digest=sha256:12de0b6729c3c3c42d4f80037b005721c0277e5522eeff73edf89759670afff6

Observation 9b3c2369-6c8b-499e-866a-88410e6e5615 · outbound

This paper cites Qwen2 Technical Report.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Qwen2 Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.751851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.751851Z digest=sha256:85e021e0e9f845b494add5063ce4f2f5dea63840ec2f0454ba254d7e46262496

Observation 4dc7458b-b4c8-4da7-b51e-d960c619ba71 · outbound

This paper cites Genie: Achieving Human Parity in Content-Grounded Datasets Generation.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Genie: Achieving Human Parity in Content-Grounded Datasets Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.755772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.755772Z digest=sha256:ada740e5d44dc3bea68be66df46932ebc711bb39d569994e4e070ee9dda927aa

Observation 629bf445-2d55-4cc5-83a5-7cabcce69417 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Yi: Open Foundation Models by 01.AI

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.760232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.760232Z digest=sha256:372b56604e8f6a528bcaafdebd30a8137b038845c3ecad96a8b0d76c7060c820

Observation 0489c865-ea9d-4798-8599-1d8488769214 · outbound

This paper cites Self-Rewarding Language Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Self-Rewarding Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.764847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.764847Z digest=sha256:5fc5f43acac1639e0388d81f7dde632472b49bf29f9b173a6f01ed100aeea2b9

Observation d00dfe03-37cc-417d-b210-f49f2e73307d · outbound

This paper cites Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.769802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.769802Z digest=sha256:17157a028de74b9a66f4be79c08ef83ef2de429589baeae8314c5a8a1af843ad

Observation 143828bb-371f-43d0-8ba7-0e6518bb23a7 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.775413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.775413Z digest=sha256:46e3e58f57868d1fcf214eed9f5de4232d42c5a4e89df68f85bd6281d27dd2fc

Observation e47622b5-ed4e-465e-a80b-cfeee8c31b99 · outbound

This paper cites an unresolved cited work.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:15:29.846814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T23:15:28.783877Z digest=sha256:316de78bedd9fd6ae090ac1b627e6b9b585d72eb87a4b14e96cd83b33021d768

Observation 3c1fe777-9ef9-46c5-9fea-6da85f5d4fbb · outbound

This paper cites an unresolved cited work.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:15:29.834256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T23:15:28.788743Z digest=sha256:9fc189ee6acb34e37e0a68117e4d93f3c495e1260e1c840b31e55f7ebb55a1a4

Observation 9b431a19-324a-48eb-af09-6baf75bbfade · outbound

This paper cites The answer is ANSWER.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models The answer is ANSWER

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:15:29.821481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T23:15:28.793471Z digest=sha256:78c85d36e6a838e3bc5e7ed0707a08a104801f6f675280da5abed31b9b2ce4a0

Observation 4bb79ffe-ae45-4439-87d8-7c3a80a411ff · outbound

This paper cites Learning How Hard to Think: Input-Adaptive Allocation of LM Computation.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Learning How Hard to Think: Input-Adaptive Allocation of LM Computation

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.575938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.575938Z digest=sha256:185e12453ddf6f2a40483eeb03578425313eb8e2ed37355d61be4bdde3effbed

Observation 1faa3f38-df3a-46fd-b810-1c80e0fae732 · outbound

This paper cites From Lists to Emojis: How Format Bias Affects Model Alignment.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models From Lists to Emojis: How Format Bias Affects Model Alignment

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.779769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.779769Z digest=sha256:95e69f07ec8fb1f6fe88f6abcc0232ca4da3a929b2ee1864c44ba8a8e4920700

Observation 63d336fa-0c4c-4805-8f21-c9204165c948 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Qwen2.5-Coder Technical Report

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.633561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.633561Z digest=sha256:cdfa17f0452f8c9213cf1a694b627be4c5472afa37106a728fe79f83d384b416

Observation 1b028eb2-10ef-4b77-ba70-6096fc73c12b · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.619469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.619469Z digest=sha256:459b0b509d4cf35ab94530a7988f83905d431a2bc8f222a082ca29b82933e2e0

Observation 792c3e5f-31ff-46e0-944c-8e18e8bec09d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.571867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.571867Z digest=sha256:2621993123192f182309cd399b22a2e309488b6fdb8802a7284a9c22da4ac723

Observation 1064790f-6c2e-435d-a6bf-1a55ab4286f9 · outbound

This paper cites RankGen: Improving Text Generation with Large Ranking Models.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models RankGen: Improving Text Generation with Large Ranking Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.654641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.654641Z digest=sha256:6ec820165ea384ca8be66fa1094940c58cc4a1c2235e741fd6dcb1ca17fe74a8

Observation 243fad25-7fbf-4e43-ad28-32c556fbd12c · outbound

This paper cites Scaling Laws for Transfer.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Scaling Laws for Transfer

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.624117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.624117Z digest=sha256:b7be6c57fda2ad33ecf75dbf41413d8680632cbae9b8e08cc1a8dcf3305a9b43

Observation 96760804-8764-4cbe-b70a-66b31af334d1 · outbound

This paper cites Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.520053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.520053Z digest=sha256:03cea4237818ace66726294ecc1b1e7b1189ffb65e7cdc6c51e6ede8b6d36762

Observation a423a56c-b440-4db2-9426-3ce9500910d7 · outbound

This paper cites Nemotron-4 340B Technical Report.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Nemotron-4 340B Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.488650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.488650Z digest=sha256:7e211d2c823a1068fef37e56240a1d8fe7ba8af8bbeda1734751fe5b7f2a085f

Observation c3c60f7e-5b8b-448f-81b5-2d195bebe09b · outbound

This paper cites Self-Consuming Generative Models Go MAD.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Self-Consuming Generative Models Go MAD

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.493309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.493309Z digest=sha256:c2159458757ae235b4a225357fec47b02be80392d00af566062a96eb13a95c2d

Observation 1480e59d-4c59-483d-8432-dd0882773667 · outbound

This paper cites Large Language Models Can Self-Improve.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Large Language Models Can Self-Improve

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.629098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.629098Z digest=sha256:d966191a9bbb2d5a6d5ee6c299da2feea492767059c4155a316a9458df9c5777

Pith citing papers

Observation 574ba37f-3705-454b-b24d-4ca3b9a0819a · inbound

Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges cites this paper.

Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T14:54:29.270448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:54:29.270448Z digest=sha256:7ed1fed16072b289590f81005355521bc4f84150850f694417bb5010c8c698bf

Observation e61b5bbf-b799-4ccf-ab58-29688b857d9b · inbound

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning cites this paper.

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:22.812009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:28:22.812009Z digest=sha256:df33bbe33edf2656b420f4307fe49c82bc32621c43d662eaff907b0151982339

Observation 1c0905cb-9aa6-41d5-ac77-19d69cfcfcb6 · inbound

Sample Complexity and Representation Ability of Test-time Scaling Paradigms cites this paper.

Sample Complexity and Representation Ability of Test-time Scaling Paradigms Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:38.336495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:38.336495Z digest=sha256:be7ceda3deec08125d574e3de86a41dc1de130484f1754d1766ccbbcc25b0cf0

Observation b760d7ee-a2d3-4a3e-a999-b4747dad6391 · inbound

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs cites this paper.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.400160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.400160Z digest=sha256:a25c8c19e13dc0bd15e9e7e9b889b5b5cb30aecc96090a99d1581f174596e1c2

Observation ecfaabd5-011e-4fb5-8f41-d6591eece8e5 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:36.497124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:36.497124Z digest=sha256:da39686bd16defe9bdbf407ed924aa8918f649bfff170f42efd7ccf5e32ba398

Observation eb7f58e0-4da1-417f-8b70-2b500455401a · inbound

Because we have LLMs, we Can and Should Pursue Agentic Interpretability cites this paper.

Because we have LLMs, we Can and Should Pursue Agentic Interpretability Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T01:03:21.917517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:03:21.917517Z digest=sha256:8c2f012f254198026c0c97a256aa0fb0aab688cd03161c2a0fc0b1006ab1e8be

Observation e2a5c216-4747-45e8-91ec-72ca80aed3cd · inbound

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability cites this paper.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.251369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.251369Z digest=sha256:e91f28b0539955a25ebd3106cdadc55ebf0d346004c8147a75f9e14aebedd5bd

Observation d1eb8d0e-a29f-4d6d-8e30-6db51c75e167 · inbound

Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online cites this paper.

Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:52.207779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:52:52.207779Z digest=sha256:fb45b0cb8c11081d25e76cf5744878f3c8a5d527196188e99a5bdcfc2191cb3e

Observation 73d87d62-3600-436e-8f4c-b412c9fef2d6 · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.556101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.556101Z digest=sha256:d6c2f029337968ea9b722293ef3519de0b3fcdb63ef0eb94685e980c46f2422d

Observation 88fa9f2e-85a3-466f-8bbc-a2d61ed641f9 · inbound

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula cites this paper.

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T02:43:36.764821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:43:36.764821Z digest=sha256:576d811cce98d79ca6c82bc64d685e9456c82cefe67b182ce95e1dcca9075df8

Observation 3c17bca8-58cc-4883-b5af-fff3617f1794 · inbound

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents cites this paper.

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:25:35.849997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T12:22:13.551709Z digest=sha256:5f603a2295bb3155f0aaf7a78abc1e78dad6367f652c3eaa3fc55147763c8653

Observation d4970fb6-29dc-4158-9ff4-388f7f142f42 · inbound

AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems cites this paper.

AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.591039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T07:02:02.992871Z digest=sha256:104dd7c603ca67615e9f39a325cc2293a5aa77e55cd5b2a58e5e732168f1d566

Observation fc296c37-263e-4b37-918d-b29e828bd96c · inbound

Tree of Concepts: Interpretable Continual Learners in Non-Stationary Clinical Domains cites this paper.

Tree of Concepts: Interpretable Continual Learners in Non-Stationary Clinical Domains Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:46:37.231503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T06:44:44.353815Z digest=sha256:0aa7b0ddb4fd94e09578fa4f2d368f753a0dc093867a6dc1251adf86a37d4fce

Observation caec8726-3d6a-4f00-a7cb-0274bf57bdcc · inbound

Annotations Mitigate Post-Training Mode Collapse cites this paper.

Annotations Mitigate Post-Training Mode Collapse Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:46:51.283999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-12T03:58:11.179607Z digest=sha256:25599cff7035ff449512cea76acf7354110d01d2664145b303a39d51918f42d7

Observation f5591126-1c6c-4e21-b113-78511853ae51 · inbound

On the Generalization Gap in Self-Evolving Language Model Reasoning cites this paper.

On the Generalization Gap in Self-Evolving Language Model Reasoning Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.928244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:a36cc0c20042896680f02ebed484b431c2fb0c4b0e256e00403ab80175c64a1d

Observation 95b3bb05-48ac-4514-b2bf-632c27e7fbed · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.478449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:bb28e649aad350b5ee87e6fa691e99608f742b9d4d6499cd6d6e79f91da398a8

Observation 6d689bb3-6ea7-4f52-9251-8247dbb3f52d · inbound

Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models cites this paper.

Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:05:42.232371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-01T04:44:56.520156Z digest=sha256:64945b999c999fdee1ce75cbdf9ae579165ae4860853cff279dfcae6a1411333

Observation 4a151dce-1fbe-4ea6-af68-2b1378e91128 · inbound

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets cites this paper.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:58.750607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:58.750607Z digest=sha256:65e56f72da42f76fabf905dc0b95d11533a16c1b24b012ef7e564b55c1634a07