Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T23:34:10.768149Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2607.15388.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T23:34:10.768149Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 946d907e-8464-449c-b318-dbdc1bfefe59 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41556648-636e-46e7-9950-5324951d6638 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Benchmarks saturate when the model gets smarter than the judge.arXiv preprint arXiv:2601.19532, 2026
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a403b750-4dea-4d47-873b-ac7062439b01 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning xverify: Efficient answer verifier for reasoning model evaluations.arXiv preprint arXiv:2504.10481, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 846e75f2-a91c-47a2-a240-9bde89501526 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning A coefficient of agreement for nominal scales.Educational and Psychological Measurement, 20(1):37–46, 1960
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4893f67-e0f0-4aa1-8f5e-89e3ef673b17 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Improving Factuality and Reasoning in Language Models through Multiagent Debate
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f3ef200-f229-4067-a7f3-498e6f885606 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21176412-77ba-4a5c-a2ab-0f877c91a028 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Large Language Models Cannot Self-Correct Reasoning Yet
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30eceb0e-c75e-4f61-917c-707457da1c27 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning AI Agents That Matter
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aabdee1f-1d04-43fe-b11a-61957d61f5fd · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Towards a Science of Scaling Agent Systems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42ff4daa-3c5a-46ba-9b67-c49dd8300c2e · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Richard Landis and Gary G
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a7c4be0-789c-49f5-8b13-e97999bf956f · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 341b1928-7631-4c62-bdb3-adb4cdaf3eca · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Decomposing LLM self-correction: The accuracy-correction paradox and error depth hypothesis.arXiv preprint arXiv:2601.00828, 2026
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70139a04-5810-41de-a13e-81fe8a44fcc1 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning AgentBench: Evaluating LLMs as Agents
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9beaba57-822b-4468-95b9-3594164b9aa8 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Self-Refine: Iterative Refinement with Self-Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bd452a2-4c47-4f12-b6e6-65ba3ebb23a6 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86a8efee-ccfb-4b45-9811-c868114bd03e · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 067ab163-1a96-4334-9b77-ddb1c0f14e19 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning On Evaluating the Integration of Reasoning and Action in LLM Agents with Database Question Answering
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3529d686-f8c3-4ff9-b29c-b384223e5562 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning gpt-oss-120b model, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84e8b31b-0595-4817-a0af-95f4cc49235e · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Introducing gpt-oss, 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3640d7a-eadc-4e2f-a5c7-afa68282f0d8 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Towards scientific intelligence: A survey of llm-based scientific agents.arXiv preprint arXiv:2503.24047, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb9daf5-7438-4253-9abd-ac506e51c305 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Reflexion: Language Agents with Verbal Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c60c6a2-bd0e-49b2-98f6-9a1564af8107 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Can llms correct them- selves? a benchmark of self-correction in llms.arXiv preprint arXiv:2510.16062, 2025
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8195945-d22a-4014-9600-36b09c49f097 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Multi-Agent Collaboration Mechanisms: A Survey of LLMs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03aecc9f-9a42-46e1-b96c-81e31afb5bb8 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Self-correction bench: Uncovering and addressing the self-correction blind spot in large language models.OpenReview, 2025
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5fd9750-3627-44b5-8ad7-cbd3003bf917 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc3468bd-8320-4030-b3f9-f885ad0a43be · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 920b6124-5843-427c-84b4-11e20deecfc5 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Can LLM agents really debate? a controlled study of multi-agent debate in logical reasoning.arXiv preprint arXiv:2511.07784, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee57ea8a-2eb3-48ca-91d1-7d5419098410 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0968b8c3-d473-4d4f-a221-e804f64a17f1 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10043a02-738e-4f71-8cc2-46530b4140f0 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c10dd3e5-58fc-489d-94eb-b793c9554495 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Survey on Evaluation of LLM-based Agents
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dbf8087-3485-4e26-9fd2-531ae83c4e6d · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Reinforce LLM Reasoning through Multi-Agent Reflection
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3f52520-79bc-49b9-9ab5-59ebf80e1509 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ccc4a33-b726-4b50-b95a-1191dffcb24e · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Small language models need strong verifiers to self-correct reasoning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf7f5c99-4b5e-4db2-9b3a-f92db4a927ac · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad9430a8-1636-4380-a69e-32bdca129aca · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8c4e30e-7b1d-4397-b872-0e12201c7fc5 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning useful” and “misleading
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cad18796-d777-4396-957f-386b92c24658 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4471bb6-c271-4c7c-99c8-d05ff01e1582 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dd02216-17b5-4e81-b764-9de6638cc68a · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ad57cec-5bdb-47a0-9059-68e4061bf554 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68631ab-0d3e-48e9-9b92-62d01dc6c367 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning wrong → same wrong answer,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c1ad37e-21ba-4226-bfeb-c4654a353f8a · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3cedc27-8928-457c-9ae1-94ed70a3bb02 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30cadc61-c180-4eb7-9a91-0179bf43d24b · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning On scalable oversight with weak LLMs judging strong LLMs
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e508252e-3f21-4cb6-bdd4-1aad689f6499 · outbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Why Do Multi-Agent LLM Systems Fail?
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.