Pith. sign in

Paper Citation Record · LEDGER

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

As of 14 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 4 inbound Pith citation observations for arXiv:2507.02778.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02778 v3

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:29:03.805548Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T21:20:00.041277Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:28:58.386085Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19a388be-1abf-4a33-ba82-7b35e40fd82f · outbound

This paper cites GPT-4 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:55.823384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:55.823384Z digest=sha256:1e38e9b51428580926faa8ab59b0fb7e2fdc92349cea111f6084c7cfd94c45f6

Observation 73f0cd9a-0b1b-408e-8072-9d94b04998e2 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku, Mar 2024.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The claude 3 model family: Opus, sonnet, haiku, Mar 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:08.781161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:28:55.944606Z digest=sha256:867f1059401f48096436a181db70c401363476558c3558b60395669b497ea3e3

Observation 50bbfc04-ce8c-4b4b-b5b0-3fff3e8b3f5a · outbound

This paper cites Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities., June 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities., June 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:08.531479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:28:56.097504Z digest=sha256:0a69d30aa540af7876cd1b7159c6bd4a8e4b586b6cd4ea8abc0d113b02522da9

Observation ae1989c8-73a6-49a0-8fc4-2d16699ff84d · outbound

This paper cites Qwen3 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Qwen3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.227672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.227672Z digest=sha256:55790898f08c6801a698014f6fa5797e702ee7b215c6a9d6debda0951b340838

Observation 586457e5-ab3b-4b4a-9464-8e20506298a9 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, Apr 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, Apr 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:08.248924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:28:56.333821Z digest=sha256:476f4994b9708de8a1a406b32016ad99a6213d4c02e57f436c9c8447c2d7ac66

Observation 91ab01f4-557c-433b-b672-94dbef010bc0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.531856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.531856Z digest=sha256:964617213cb3004ff8841bcbddec55841301f4fed0ee3f99c3ed61f7565f858e

Observation f38e80ac-bd0f-42de-b022-b8b7c6500e47 · outbound

This paper cites On faithfulness and factuality in abstractive summarization.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models On faithfulness and factuality in abstractive summarization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.646944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.646944Z digest=sha256:1892214022659174e7bf3ec2dd8ddb8a54a3a4ebc69cda442a1b79fc727d46ca

Observation 77338fbe-e90f-4bc5-bba8-76b029fc8ea4 · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.783467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.783467Z digest=sha256:944e09de86c5f5d14e8f2e0a3633dca2c154af1141666ff3901a1f7fcc0cb693

Observation 098af2f9-edf9-49c1-a735-ebf863240b3d · outbound

This paper cites Do, Yan Xu, and Pascale Fung.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Do, Yan Xu, and Pascale Fung

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:56.946027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:56.946027Z digest=sha256:2b25f40f4c41930c516157d5dbfcede9d45fe27ff239987f4154ee7a1ad92676

Observation caa6fbd6-6cbc-4f13-b88a-1bcb810818e9 · outbound

This paper cites Large language models can be easily distracted by irrelevant context.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Large language models can be easily distracted by irrelevant context

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.071428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.071428Z digest=sha256:d5b416c79e6bbaa9dbb9c7bfdae1c830a02df119ffdc66f7c054e240fcdcf0ac

Observation d87bcd25-bf69-48cd-b7fa-c14ce7548102 · outbound

This paper cites Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.217863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.217863Z digest=sha256:024128d878c5ee8b5e0ae63da1dc187d4db2bd2ff2661a38cc533cabe309dd71

Observation fae702c4-d4da-4075-8c5e-a0238f612dfe · outbound

This paper cites Reflexion: language agents with verbal reinforcement learning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Reflexion: language agents with verbal reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:08.005937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:28:57.332058Z digest=sha256:cc5b2242033044754dc61a4efdbefc2c9eb80451940230085eef9a046d0d7fc8

Observation 1cecc2d2-7db8-48a5-94cf-cff9ceed325b · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Self-refine: Iterative refinement with self-feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.429687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.429687Z digest=sha256:3b8cebfc33b7f2b930b0ade237febe1316addaea35e91b29577d73037bdf5471

Observation 0952229d-8e2e-4192-9c73-59e1a4b6dcff · outbound

This paper cites Language models can solve computer tasks.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Language models can solve computer tasks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:07.724534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:28:57.560434Z digest=sha256:3034fbe216a82c19ac76f9f0e62be625ddb9ef7ef6faf9697df42be857b414cb

Observation ee98ddf2-8acf-4313-9d25-4c6955fe54ee · outbound

This paper cites When can LLM s actually correct their own mistakes? a critical survey of self-correction of LLM s.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models When can LLM s actually correct their own mistakes? a critical survey of self-correction of LLM s

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.717223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.717223Z digest=sha256:c52ca747a347d4c72c8fc5e1e82c3895eb642038cd1371d23a629df59c6142ec

Observation 7cb640df-e27c-4882-80c9-919eeba976dc · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Large Language Models Cannot Self-Correct Reasoning Yet

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.838761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.838761Z digest=sha256:4b65fb404b9d2575a4b89c276362acb6e74c1037a6664398e1f279411c0c24a8

Observation 8923266a-b790-4707-a354-bce566c6368d · outbound

This paper cites LLM s cannot find reasoning errors, but can correct them given the error location.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models LLM s cannot find reasoning errors, but can correct them given the error location

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:57.998432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:57.998432Z digest=sha256:21f90ee0745a73fc3251af000c9e9ed93f24e7bd3d48c25e8900224e92a0e5bb

Observation ed5bf506-a601-41b5-a800-544053498431 · outbound

This paper cites Evaluating LLM s at detecting errors in LLM responses.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Evaluating LLM s at detecting errors in LLM responses

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:07.454677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.142786Z digest=sha256:1576e075d58aa5d02511dec692baaff2e1997dfbf00a8d9b9596caee22fc2672

Observation 5f1f55b3-3870-468a-9f67-1af9ce35a562 · outbound

This paper cites Training language models to self-correct via reinforcement learning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Training language models to self-correct via reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:07.214068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.292638Z digest=sha256:cba37fe9f15aaeb6a5383b6593d55a6909335547a4b3063964e98ea818931375

Observation efddf2d7-1f69-4927-bb0a-f4f3896e609c · outbound

This paper cites Jailbroken: how does llm safety training fail? In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS '23, Red Hook, NY, USA, 2023.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Jailbroken: how does llm safety training fail? In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS '23, Red Hook, NY, USA, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:06.943034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.454749Z digest=sha256:28516ecb085d895a71a7a08cedd42abbd9d6165b6489480926ddc5c57801723e

Observation 553309ad-0b0f-4417-b8b1-f864b3625a67 · outbound

This paper cites Formalizing and benchmarking prompt injection attacks and defenses.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Formalizing and benchmarking prompt injection attacks and defenses

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:06.592654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.558097Z digest=sha256:04df52b2c2040cc35e56cd65b53964795c31cb31d71900ed376c70ffe247f10b

Observation 7caeb070-8c3f-4b21-a32e-ab3575a5ea7d · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:58.704896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:58.704896Z digest=sha256:2bce10f411d103af9924b2a8c0cdbfd83de10354c485dbaee3576d4bb2e07550

Observation 9aced9d1-1d08-48f8-87f4-1e5b8e6a1a5b · outbound

This paper cites an unresolved cited work.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:29:06.302848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:28:58.869328Z digest=sha256:ed2636ccf6718e146ebceeaf2f17e9bbe3a0e21846efa2162ff15d12fdb21b5d

Observation 6e96adf5-08a3-472a-8d93-9315cfd851d5 · outbound

This paper cites Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.074923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.074923Z digest=sha256:963826393b0b9d2f8bffb287843714631ff0803b4cbdc31a4e4d8bba33800747

Observation a61249e0-c3ef-455b-b871-6c8f7c79c70c · outbound

This paper cites s1: Simple test-time scaling.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models s1: Simple test-time scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.270560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.270560Z digest=sha256:db14728f4485a85f2eb1824e07f3aab3eebc81ba15f1b09215bd464589ce904f

Observation f0954cbd-2c14-4b6b-8801-899f2a583ed1 · outbound

This paper cites Benchmarking cognitive biases in large language models as evaluators.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Benchmarking cognitive biases in large language models as evaluators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.400211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.400211Z digest=sha256:4541ce8627d70e1a65bee91f8c8876c1d825c3e31b315caa7568cbfcd4ec9e46

Observation 78c11968-cc45-47a7-b086-2d93729aa853 · outbound

This paper cites Cognitive bias in decision-making with LLM s.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Cognitive bias in decision-making with LLM s

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.570749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.570749Z digest=sha256:f655fd421577869e4e242f53a018a8f228b9050f624d4276d7b8ca7b28e2ad6b

Observation bc521204-141d-47ff-be90-5f34ec347f6e · outbound

This paper cites Capturing failures of large language models via human cognitive biases.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Capturing failures of large language models via human cognitive biases

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:06.008527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:28:59.745163Z digest=sha256:9cf97b56cec850d894ff08b2c7c88265ce193d56771a214f669b9d84ece68c98

Observation 42e7bac0-e04d-4130-8fc7-375e883a226c · outbound

This paper cites Lin, and Lee Ross.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Lin, and Lee Ross

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:28:59.906202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:28:59.906202Z digest=sha256:7118ae5eae7e4d04fcec7b48e8c567be457c7c79ac042c45b4753c8fda67bd2c

Observation 2c893750-911c-4989-a3ab-7264afd41c4f · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.072960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.072960Z digest=sha256:62a0e10aebb91a07421bce7beda2fd76641cc3676af3b849a749e720aa6aac2f

Observation af15ffc1-4b55-4d7b-893f-55462028cbe3 · outbound

This paper cites PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.154021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.154021Z digest=sha256:bfcb3619cf841799e031732a96c3a61ee0681a5ea3faa2a34d967d636eb03496

Observation 242c2d42-a346-4983-a5f9-e9fb052d06a4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Training Verifiers to Solve Math Word Problems

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.206673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.206673Z digest=sha256:2ebfd5a8cdfcca1c0a50bdfb519f52438cf7839af5bbdd3836986204d4c1a14c

Observation 0a682beb-9d8b-4ffc-8185-1c377d0db050 · outbound

This paper cites Introducing gpt-4.1 in the api, Apr 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Introducing gpt-4.1 in the api, Apr 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:05.741704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:29:00.270483Z digest=sha256:ffb4bae6f06e2c865d384aa4a825349083876b6eba538eeafaf2e9a65d59991a

Observation 547769ab-98dd-42f4-8e5d-f44fecc985c4 · outbound

This paper cites Let's verify step by step.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Let's verify step by step

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.355018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.355018Z digest=sha256:ba2b920ac7ccfca8e58029f3126d7760e135d42653595c6780f583d1058e95f2

Observation cdac7ea2-3ff3-405a-9392-f1aa43da387b · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Measuring mathematical problem solving with the MATH dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.438344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.438344Z digest=sha256:8335618ce10a57c8aed644c208fcd84adb433d6d91fc74e8557477493beb8a22

Observation d1a30e71-b262-458f-8efa-c396237e35b6 · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Transformers: State-of-the-art natural language processing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.521903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.521903Z digest=sha256:8eb7498e00fb756dc86d44982264b9ac1c2775b28d0d0b309f515b73c116ea61

Observation 98372d2d-19f9-45e5-8238-be050bbd90ed · outbound

This paper cites DeepSeek-V3 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models DeepSeek-V3 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.633907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.633907Z digest=sha256:67be1dfd78811892b2f827aa57429d9a8c2bf50f429084a01fd70f976e8e40a2

Observation 5e52510a-4855-4949-9e1d-a50fd817f0f4 · outbound

This paper cites Qwen2.5 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.711530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.711530Z digest=sha256:c779f74e935526abb0745d5382c5f375e2f89fa560f2e9e81d855f4790d42aad

Observation a5b12f90-c1f5-4320-988b-69bde06197c9 · outbound

This paper cites Llama 3.3, Dec 2024.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Llama 3.3, Dec 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:05.454473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:29:00.790776Z digest=sha256:6d1f80221dc4156c5a90d66617ec82e8de4ac8349bd64cccab224583257b8f21

Observation 070d8f16-3cc9-4caa-ba86-afcde95ccfc4 · outbound

This paper cites Phi-4 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Phi-4 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:00.903257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:00.903257Z digest=sha256:943112b722248edeefe80d56bf203a62c63dabb6b954fa050fefd43ccfafed17

Observation d8898eac-0793-4381-a9b0-2df54c6e3a79 · outbound

This paper cites Qwen2 Technical Report.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Qwen2 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.073515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.073515Z digest=sha256:a43f6f0a9f6bb471820f6908a849aa08619120422a6e41e9dba65e2c3f77ec4b

Observation 598889cd-4a56-43db-8cd4-7e4e3b4c55bc · outbound

This paper cites The Llama 3 Herd of Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The Llama 3 Herd of Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.235737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.235737Z digest=sha256:6fda4892bd8de9ba247e8af4d82356364e5d942c19310ae5b5853a4e91ac9fd5

Observation 9c5dc13c-19b8-4c72-8e21-1ba13f577e88 · outbound

This paper cites Mistral small 3, Jan 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Mistral small 3, Jan 2025

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:05.122636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:29:01.386989Z digest=sha256:d4b8cc5500eb87c62f19be13de58cf2b39d26762b718428c6bad73dd6013d8da

Observation 91a1453a-6c9b-4155-aebf-55fa2131cd9f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.501958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.501958Z digest=sha256:1c7c7a9b721c5ad7bc55f9afb9326e6ca4e12d7dc2802327f9866cc10d5992f1

Observation 8a78e84b-9451-4307-9a7a-d78784925c50 · outbound

This paper cites Impact of pretraining term frequencies on few-shot numerical reasoning.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Impact of pretraining term frequencies on few-shot numerical reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.660462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.660462Z digest=sha256:0c19f5759284514523f7dfae2ad91e2acf7b170819bdf0f2aceecd3c1a150620

Observation a8943ef9-1315-4cd0-8398-f439535fbb8e · outbound

This paper cites Smith, Sarah Wiegreffe, and Yanai Elazar.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Smith, Sarah Wiegreffe, and Yanai Elazar

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:04.782194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:29:01.845245Z digest=sha256:8f06fef0636bed7321a4906d095b4119617fd71cf20214f129f51475c1010dc6

Observation e5ca97be-2dab-4115-a7b7-171ee449086e · outbound

This paper cites o pf, Yannic Kilcher, Dimitri von R \.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models o pf, Yannic Kilcher, Dimitri von R \

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:01.991836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:01.991836Z digest=sha256:b7cfb8927b1be299b4514a4ad3715be31ae92bc01515e3e5b1b7d77a417d6928

Observation 9472acb6-eb08-412e-a62b-0d21af73a4f9 · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.140460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.140460Z digest=sha256:7a42abfea76d8dff4eb4ff3fb9dd70050c02d9c212e770b404c996a9d0889a04

Observation c317eade-a7e4-48ad-8654-5687149260d8 · outbound

This paper cites Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.244807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.244807Z digest=sha256:8b43330a42774f8fada0c4d0e48e28d1126c4b0b1c18fe4a16b6e5a9f7113dc3

Observation 61ed6acb-e56a-49e0-8e0a-49c8c23c40b5 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2024.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Ultrafeedback: Boosting language models with high-quality feedback, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:29:04.501080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:29:02.356141Z digest=sha256:10819b1e64fe2fbe39b0452699902d261fc3dd40b11b2452807e96e641eb1981

Observation d42f5cd2-00f6-4eeb-98c0-831a16acdeb7 · outbound

This paper cites Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.487929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.487929Z digest=sha256:a14a41cafad77bb1f4bc084cb6c62fd61b040cb9b85189108394309df3b4036c

Observation fa7a74bf-6fcc-4198-83eb-b128ade148e9 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.599732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.599732Z digest=sha256:9620cdd9a90b08b69cb4a31aca26c00575150e3503678ea6df51a03c9bd3df23

Observation fab9b386-0834-49e7-9db5-dc727f9131ea · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models OpenThoughts: Data Recipes for Reasoning Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.783626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.783626Z digest=sha256:95f6b1d09d2cf7316d26ac1efdc7445b2dda6fed31c962620ced1561644afdcb

Observation 432e34ca-6d12-4311-9df2-c0977733d550 · outbound

This paper cites Training language models to follow instructions with human feedback.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Training language models to follow instructions with human feedback

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:02.938680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:02.938680Z digest=sha256:e86f7707028fc272e3d460eae1b58485bf07ec3a3f9c6e49a19d887069d31a91

Observation 5a580e71-de61-4a51-819e-81f6b955611e · outbound

This paper cites Learning From Mistakes Makes LLM Better Reasoner.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Learning From Mistakes Makes LLM Better Reasoner

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.095630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.095630Z digest=sha256:7bd75a7ee44005aecb2530fa4fd0a456769424be16a784c30595128d695df8b4

Observation f256587f-fd18-4918-be5b-1e797c5c666e · outbound

This paper cites Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.216236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.216236Z digest=sha256:d5a414ec72a95c046c5a8eb4433931ed86dcc9d64bcad5657db05bdc0e018d12

Observation 4eb03d92-b6a5-4658-90b2-9dd56e085279 · outbound

This paper cites The effect of sampling temperature on problem solving in large language models.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models The effect of sampling temperature on problem solving in large language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.409054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.409054Z digest=sha256:38164ab32f6958466407a258ad1c4572c0dcd6b8b4c3abd61a7c73dce41a6258

Observation d70d4393-6219-408a-8892-8d96da99a79c · outbound

This paper cites @esa (Ref.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models @esa (Ref

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.530733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.530733Z digest=sha256:f091af353bdcbd7a9e7599c1e72010aa02748bdd03193ad43ce0957962165026

Observation a4614c9e-01f7-4332-adb6-eec84c16fca5 · outbound

This paper cites an unresolved cited work.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.675873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.675873Z digest=sha256:76defd574751a224541f956fa6f301b38841f2e3eb4a9df69f8ce9ea95a871aa

Observation e8dda0b1-32fe-40b6-bebf-104c5fd66015 · outbound

This paper cites after incorrect reasoning or answer to prompt LLMs to self-correct, without finetuning. We observe significant reductions in the blind spot after appending ``Wait.

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models after incorrect reasoning or answer to prompt LLMs to self-correct, without finetuning. We observe significant reductions in the blind spot after appending ``Wait

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:03.805548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:03.805548Z digest=sha256:aa0374c2fe16a386313696907900c2ae6521ea6d46c77f8ab7f9e84e0ec03426

Pith citing papers

Observation 3547e909-d01a-42bc-8fbe-10036c058fef · inbound

ReFlect: An Effective Harness System for Complex Long-Horizon LLM Reasoning cites this paper.

ReFlect: An Effective Harness System for Complex Long-Horizon LLM Reasoning Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-08-04T02:29:36.375484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-08T11:35:25.204350Z digest=sha256:7582a45accb6d5f57e16f212cc91e1a77da73a54f5be929942c82fdba853fe38

Observation 96d7ffc6-6dae-45a3-989d-37e0f37d0836 · inbound

ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection cites this paper.

ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-08-04T02:29:36.375484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T02:35:37.724212Z digest=sha256:864bbfbfdf71ef0a453e6192c37742f7c4a7119c1cac4ac31277d2e937a6862c

Observation 3a91d3a0-6f38-4726-acc8-06b923d8d159 · inbound

Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning cites this paper.

Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-08-04T02:29:36.375484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T07:11:02.464556Z digest=sha256:0dccf392d241d2109efe097a3e79e5971e62b8e2a8e88853c39049056432dd65

Observation eaa9921f-da88-4f3c-9369-e2f5a7f177a6 · inbound

ESC: Emotional Self-Correction for Reliable Vision-Language Models cites this paper.

ESC: Emotional Self-Correction for Reliable Vision-Language Models Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-08-04T02:29:36.375484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-03T21:20:00.041277Z digest=sha256:cffc4d235a8c903d2074689cb25ffac0f72b1402d33b32475b1c0215deb501d4