Pith. sign in

Paper Citation Record · LEDGER

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning

As of 7 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2607.15388.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15388 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T23:34:10.768149Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 946d907e-8464-449c-b318-dbdc1bfefe59 · outbound

This paper cites How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:05.670060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:05.670060Z digest=sha256:b2a4fe8946a562d2ff7513dfa2a22029fca47896884cb0068cae0d05b0328595

Observation 41556648-636e-46e7-9950-5324951d6638 · outbound

This paper cites Benchmarks saturate when the model gets smarter than the judge.arXiv preprint arXiv:2601.19532, 2026.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Benchmarks saturate when the model gets smarter than the judge.arXiv preprint arXiv:2601.19532, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:05.720372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:05.720372Z digest=sha256:24ee6a0dfb216b45cc36f6db53a5392d937f1783a94602d720dc1327e97fc81a

Observation a403b750-4dea-4d47-873b-ac7062439b01 · outbound

This paper cites xverify: Efficient answer verifier for reasoning model evaluations.arXiv preprint arXiv:2504.10481, 2025.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning xverify: Efficient answer verifier for reasoning model evaluations.arXiv preprint arXiv:2504.10481, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.070973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.070973Z digest=sha256:b5422eb8493087bd9355a016a39828465d70c0245abeb337a0b7642f226a32b4

Observation 846e75f2-a91c-47a2-a240-9bde89501526 · outbound

This paper cites A coefficient of agreement for nominal scales.Educational and Psychological Measurement, 20(1):37–46, 1960.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning A coefficient of agreement for nominal scales.Educational and Psychological Measurement, 20(1):37–46, 1960

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.218181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.218181Z digest=sha256:bd656b59678bb3511c58929cc75bd8de8548d8e615d1ac84b9a14d31d5823d95

Observation f4893f67-e0f0-4aa1-8f5e-89e3ef673b17 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.309736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.309736Z digest=sha256:1578bdf03b9b749da1c3b4c33b00f41b850af273df045eab66b45737ff32a978

Observation 5f3ef200-f229-4067-a7f3-498e6f885606 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.382219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.382219Z digest=sha256:213313c224b2da03918551a0247e4f915fd8fb60543114708a8d463905cc3187

Observation 21176412-77ba-4a5c-a2ab-0f877c91a028 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Large Language Models Cannot Self-Correct Reasoning Yet

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.435035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.435035Z digest=sha256:31ee8e27b8680a65babdefc7a57f8e8a3678a41e20a8bbe046649c92185383c7

Observation 30eceb0e-c75e-4f61-917c-707457da1c27 · outbound

This paper cites AI Agents That Matter.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning AI Agents That Matter

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.567758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.567758Z digest=sha256:1048cd99796ff43c50b5130269d2adf571c5955734c95af69452068aeb4673d9

Observation aabdee1f-1d04-43fe-b11a-61957d61f5fd · outbound

This paper cites Towards a Science of Scaling Agent Systems.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Towards a Science of Scaling Agent Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.985088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.985088Z digest=sha256:f9179018ee91b79bb9b8c9aafd8e1ada44846de196f92f8a14186168d123b476

Observation 42ff4daa-3c5a-46ba-9b67-c49dd8300c2e · outbound

This paper cites Richard Landis and Gary G.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Richard Landis and Gary G

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.101341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.101341Z digest=sha256:d12331727bf9240f50df5f812aa072094bdd38f9c886aa656d5405b94cee148d

Observation 1a7c4be0-789c-49f5-8b13-e97999bf956f · outbound

This paper cites CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.165520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.165520Z digest=sha256:b0456cdca8dccc6f2e5486b61bcd86b216e6d3325d90a119411ec16ea3d0ddde

Observation 341b1928-7631-4c62-bdb3-adb4cdaf3eca · outbound

This paper cites Decomposing LLM self-correction: The accuracy-correction paradox and error depth hypothesis.arXiv preprint arXiv:2601.00828, 2026.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Decomposing LLM self-correction: The accuracy-correction paradox and error depth hypothesis.arXiv preprint arXiv:2601.00828, 2026

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.241798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.241798Z digest=sha256:97dfd19668270c1ec276c65913fef5e25dc6c96fc856de7c0470929b3691e92f

Observation 70139a04-5810-41de-a13e-81fe8a44fcc1 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning AgentBench: Evaluating LLMs as Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.318575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.318575Z digest=sha256:0a2021680ccf8e4d20bf19731363333e487d3d5d03feb3542143292594cd695c

Observation 9beaba57-822b-4468-95b9-3594164b9aa8 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Self-Refine: Iterative Refinement with Self-Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.388941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.388941Z digest=sha256:2458f947d6ebb6a0064a6e8bc30820db52ee67ad9c8e26dc9744624097fef66f

Observation 6bd452a2-4c47-4f12-b6e6-65ba3ebb23a6 · outbound

This paper cites an unresolved cited work.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.475071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.475071Z digest=sha256:072a2f7127632f1335395557ae7170293f1bcd3013285cf68a274bf363657ada

Observation 86a8efee-ccfb-4b45-9811-c868114bd03e · outbound

This paper cites an unresolved cited work.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.799710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.799710Z digest=sha256:7fb1dacda4fd12e0951fe62fe2d9c3aa7c3e8b2f11d77d3a0f9216428b14559d

Observation 067ab163-1a96-4334-9b77-ddb1c0f14e19 · outbound

This paper cites On Evaluating the Integration of Reasoning and Action in LLM Agents with Database Question Answering.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning On Evaluating the Integration of Reasoning and Action in LLM Agents with Database Question Answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.962606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.962606Z digest=sha256:f38155820b7369965a17afadbd7a7d578c19069a2bbf8486ca7ff251f95f233c

Observation 3529d686-f8c3-4ff9-b29c-b384223e5562 · outbound

This paper cites gpt-oss-120b model, 2025.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning gpt-oss-120b model, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:08.166111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:08.166111Z digest=sha256:1af1bb17ab40a63ee5123b55514a41f760241126b94977c95bff9121af201588

Observation 84e8b31b-0595-4817-a0af-95f4cc49235e · outbound

This paper cites Introducing gpt-oss, 2025.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Introducing gpt-oss, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:08.305502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:08.305502Z digest=sha256:8ea7b671cff8bea8395111bb3a109da6b896b60e6620441114f34a07dae474ca

Observation a3640d7a-eadc-4e2f-a5c7-afa68282f0d8 · outbound

This paper cites Towards scientific intelligence: A survey of llm-based scientific agents.arXiv preprint arXiv:2503.24047, 2025.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Towards scientific intelligence: A survey of llm-based scientific agents.arXiv preprint arXiv:2503.24047, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:08.446253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:08.446253Z digest=sha256:bfa33a6976faa6a149d0d254ed9970ef6b8bd3d1e66aae7c6ac92c2c8900a5ba

Observation 1cb9daf5-7438-4253-9abd-ac506e51c305 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:08.642754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:08.642754Z digest=sha256:20054e4fadeed2c608b2ddb91fccc965dfa5c186045680eb8f5409a92391e43d

Observation 0c60c6a2-bd0e-49b2-98f6-9a1564af8107 · outbound

This paper cites Can llms correct them- selves? a benchmark of self-correction in llms.arXiv preprint arXiv:2510.16062, 2025.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Can llms correct them- selves? a benchmark of self-correction in llms.arXiv preprint arXiv:2510.16062, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:08.734542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:08.734542Z digest=sha256:15bdbeb73b855cf9a76f634288b1c42057e3c3f4342a59e38138e90ab8063191

Observation d8195945-d22a-4014-9600-36b09c49f097 · outbound

This paper cites Multi-Agent Collaboration Mechanisms: A Survey of LLMs.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Multi-Agent Collaboration Mechanisms: A Survey of LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:08.868842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:08.868842Z digest=sha256:06a936cc6ebdac66b3d7bb7eb16f8cfc2d441628c3cab80cc7dea443426e69da

Observation 03aecc9f-9a42-46e1-b96c-81e31afb5bb8 · outbound

This paper cites Self-correction bench: Uncovering and addressing the self-correction blind spot in large language models.OpenReview, 2025.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Self-correction bench: Uncovering and addressing the self-correction blind spot in large language models.OpenReview, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:08.999850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:08.999850Z digest=sha256:7e864f4e1e622912b93f9bf66e3c019e67521689056f4b2f827ed5011333386f

Observation c5fd9750-3627-44b5-8ad7-cbd3003bf917 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.101806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.101806Z digest=sha256:4a838317e2de1f47a0894f9e3ec14290cfb8916fcc12495c1540268d1e10e89e

Observation bc3468bd-8320-4030-b3f9-f885ad0a43be · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.206984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.206984Z digest=sha256:e87e9a91fc547fc42a9e7a4eaf391de9b303ba4d69e3955ee8c4a089e500db59

Observation 920b6124-5843-427c-84b4-11e20deecfc5 · outbound

This paper cites Can LLM agents really debate? a controlled study of multi-agent debate in logical reasoning.arXiv preprint arXiv:2511.07784, 2025.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Can LLM agents really debate? a controlled study of multi-agent debate in logical reasoning.arXiv preprint arXiv:2511.07784, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.317994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.317994Z digest=sha256:f66ea5fa96235e745f987cf532bb96647ab8abca9a92253f9e0faa9685df0a91

Observation ee57ea8a-2eb3-48ca-91d1-7d5419098410 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.454099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.454099Z digest=sha256:606a8ae34451481c04f8353047febcd29ffe9bb585fbe6df93de3c55f0c4f304

Observation 0968b8c3-d473-4d4f-a221-e804f64a17f1 · outbound

This paper cites Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.539719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.539719Z digest=sha256:119d5a757b9aa137ddcc0855c44eb11ca5491df6914387ea24be9a13b7ebd1d9

Observation 10043a02-738e-4f71-8cc2-46530b4140f0 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.629909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.629909Z digest=sha256:ea23cb5c5d01b986d38504e3779269b5cc0393e366d85ce7bcee31f570a80593

Observation c10dd3e5-58fc-489d-94eb-b793c9554495 · outbound

This paper cites Survey on Evaluation of LLM-based Agents.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Survey on Evaluation of LLM-based Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.710058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.710058Z digest=sha256:f0a5c6d8407af535f308e254f12de50a647651eaa714cea1e2323e1c715f4a62

Observation 7dbf8087-3485-4e26-9fd2-531ae83c4e6d · outbound

This paper cites Reinforce LLM Reasoning through Multi-Agent Reflection.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Reinforce LLM Reasoning through Multi-Agent Reflection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.797820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.797820Z digest=sha256:3f8b273345a1dc70b748658971d6b9ad29904587130639c8a7be033a671f039f

Observation c3f52520-79bc-49b9-9ab5-59ebf80e1509 · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.881076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.881076Z digest=sha256:b025dfb7f2b842c532bdaee088d9e7be80b678fb1cebfd02cfbf19af54fcc94f

Observation 7ccc4a33-b726-4b50-b95a-1191dffcb24e · outbound

This paper cites Small language models need strong verifiers to self-correct reasoning.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Small language models need strong verifiers to self-correct reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:09.964586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:09.964586Z digest=sha256:3662198f57004cc3cf113bede4cbcd3292b09c5e325cf02a7feb7b42a0e6a683

Observation cf7f5c99-4b5e-4db2-9b3a-f92db4a927ac · outbound

This paper cites Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.038013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.038013Z digest=sha256:ba5c138d0c49696c87f99a1b1179c3411ef97828eb1a805a355beac82914027a

Observation ad9430a8-1636-4380-a69e-32bdca129aca · outbound

This paper cites Establishing Best Practices for Building Rigorous Agentic Benchmarks.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-01T23:34:10.122008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.122008Z digest=sha256:0c333a7547de632430b68ddf665e70ddc2251d5a2d206ff29695be46437869cb

Observation d8c4e30e-7b1d-4397-b872-0e12201c7fc5 · outbound

This paper cites useful” and “misleading.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning useful” and “misleading

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.218624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.218624Z digest=sha256:fe7bbd2d76d985e34d58cf8b9fe7c04b0593e7e8141f7cc83d369ff974b76750

Observation cad18796-d777-4396-957f-386b92c24658 · outbound

This paper cites an unresolved cited work.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.309632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.309632Z digest=sha256:3203b469c2766b26864cf9d13808271eba72f81a62c722381eedfb616030ab00

Observation e4471bb6-c271-4c7c-99c8-d05ff01e1582 · outbound

This paper cites an unresolved cited work.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.388391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.388391Z digest=sha256:4f1facae4dfc64086d0d3473a0756f56b089aadf54221c81ea2f0a570d439e2f

Observation 9dd02216-17b5-4e81-b764-9de6638cc68a · outbound

This paper cites an unresolved cited work.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.478432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.478432Z digest=sha256:2df3f710d4f9dce2319ab1ae42dac12528ed79e98e9c72985b85c3b82d64a5ee

Observation 9ad57cec-5bdb-47a0-9059-68e4061bf554 · outbound

This paper cites an unresolved cited work.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.578951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.578951Z digest=sha256:a7a42596017d4e49941dd79b7a54daaaa7e917e1f2643c5b116b17dbb8da9f4c

Observation d68631ab-0d3e-48e9-9b92-62d01dc6c367 · outbound

This paper cites wrong → same wrong answer,.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning wrong → same wrong answer,

Reference 47

Resolution
malformed identifier
no resolver link, observed 2026-08-01T23:34:10.677408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.677408Z digest=sha256:71edd8d8bb08c927ff11e7f5be60fd097cb83cd2894d01786af9546ce42a0248

Observation 8c1ad37e-21ba-4226-bfeb-c4654a353f8a · outbound

This paper cites Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:10.768149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.768149Z digest=sha256:c472e815dbb1fccf6254bfb28d3860b943583a1d01b271c19269847c372cbaac

Observation d3cedc27-8928-457c-9ae1-94ed70a3bb02 · outbound

This paper cites an unresolved cited work.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Unresolved cited work

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:07.685579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:07.685579Z digest=sha256:20a7c86be48e3c28f223a5c2f4f2be94c8cdad008efb71d0d3ecf9820761c830

Observation 30cadc61-c180-4eb7-9a91-0179bf43d24b · outbound

This paper cites On scalable oversight with weak LLMs judging strong LLMs.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning On scalable oversight with weak LLMs judging strong LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:06.878529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:06.878529Z digest=sha256:7ce98a1c1478fe6ef1836379871597566dbe1ba05f4ab88684695e7c11377d84

Observation e508252e-3f21-4cb6-bdd4-1aad689f6499 · outbound

This paper cites Why Do Multi-Agent LLM Systems Fail?.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Why Do Multi-Agent LLM Systems Fail?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T23:34:05.935527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:05.935527Z digest=sha256:7ebb14c1093462880b67b8620b891e1aad9ec558d3cbb5c1396ce3bb627b37b5

Pith citing papers

No inbound Pith citation observations are available.