Pith. sign in

Paper Citation Record · LEDGER

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling

As of 6 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2605.01566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.01566 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T14:07:43.507761Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T03:36:57.168246Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T03:45:55.534976Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact5
  • verified fuzzy3
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch9

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b1d3d7c6-86b6-46d1-9156-6c8fed48d02f · outbound

This paper cites Stay Focused: Problem Drift in Multi-Agent Debate.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Stay Focused: Problem Drift in Multi-Agent Debate

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:06:03.954920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:5405f41430ff2500557df34662e9d186dc24ed741aba2d6febaf94034b6f39cc

Observation 327d15b7-e6a5-469b-b23c-5fb5bac26362 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T17:06:03.958989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:a1744d324322a8e9316f9947b3ffb620ce122ac13282b29fad9553b4f7795d9a

Observation 94f9400c-5db8-4f85-842a-4162020c0d57 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:45.301651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:bbc8ed01d9cfcd145cfbd74573a783a458d9d759273f22de653ceaacb651b910

Observation 3d6190e3-2adc-4199-8762-fdb03b7bf5ef · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Large Language Models Cannot Self-Correct Reasoning Yet

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:48:28.335998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:941176ca86e9544558e19ec4c9ecdc46ac7182dae02c6428a0975cbed14f65fd

Observation 11b868db-ae8d-4e5e-a1c4-d667044d409c · outbound

This paper cites InProceedings of the Fourth Workshop on Scholarly Document Processing (SDP 2024), pages 105–119, Bangkok, Thailand.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling InProceedings of the Fourth Workshop on Scholarly Document Processing (SDP 2024), pages 105–119, Bangkok, Thailand

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T02:26:30.158161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:151e551136fb54aae17067410d00380916c35384907539e8f977c51b00735b08

Observation cd757e76-605d-481c-a94f-b81911b71bce · outbound

This paper cites Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T17:06:03.926913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:2d51e998c8dd4b89990690ee715d282dbe34fa26708b31a778429572771bd387

Observation 325bfc76-8678-4d41-946d-5c2757f585bf · outbound

This paper cites Scaling Laws for Neural Language Models.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Scaling Laws for Neural Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:06:03.962458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:630d112ced0174956ff36edb8569239370976b8878121267f7ad91675d10bf33

Observation 39c0b790-40d8-401f-9589-4db5ccafcdf4 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Self-Refine: Iterative Refinement with Self-Feedback

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:06:03.965474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:33e229ff4ad9609acdf5304cb71fbf5e065815360e71106fe8e6246f31ac4c48

Observation eb943f7a-b03a-46bd-9a42-0283e59703eb · outbound

This paper cites The Llama 3 Herd of Models.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling The Llama 3 Herd of Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:06:03.932948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:28c67a527f2c8b8735e52700267b9a222a622109b3410fafcca1022a7e0a2abd

Observation 5039904a-1f1f-4379-809d-79398ce67b24 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T17:06:03.950400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:0a656b67c97606cb46fbac8d4e0e60cd1b96ee68b30f3748417c2129decc5949

Observation 7552b230-5f20-47e3-8b64-b7e88137f3b4 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:01:10.451911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:67584605edeb1ed01f1078f4baa78f1363b50967159727e687631b74b751c24d

Observation b205710e-600d-4d3e-b26c-2923fc9e67b3 · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T19:29:34.516338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:1a485e4442bfcca8629faff96795dc0b7040b2ba2b01efbc7a461b2141991249

Observation 42617628-578f-44d7-9742-9f36916819c3 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T17:01:10.436307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:061870e04e079d7f20661975122c7d66b46b6e9491ff7bcc172d9da3e36a9366

Observation d8f3b0b5-94f1-4391-9da1-2a8cdff683a3 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T17:01:10.429053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:8fca5d42825e06ed3b3fbbf3aee83437495e66f2480e7637b369a28104af4f17

Observation 334b9a48-f8cd-4efd-8a9f-efb92cf0e7ad · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:38:37.536547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:fb4d630f88ddf120cd8d697f035c20215b83b2b89d8471bd4208a25d367e1b4d

Observation 69bdb243-3903-40f2-b014-874be9a2d89a · outbound

This paper cites FLOPs are calculated with the formulas from Table 4 of Appendix B.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling FLOPs are calculated with the formulas from Table 4 of Appendix B

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T02:26:30.149968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:33b300e5174461f3e26808c1029adbbf428ca9716426502475e95154b1a563e3

Observation 37cc218c-9373-4643-a7c3-e62d04852c65 · outbound

This paper cites FLOPs are estimated by assuming that each parameter is used for one multiplication and one addition per token.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling FLOPs are estimated by assuming that each parameter is used for one multiplication and one addition per token

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T02:26:30.146256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:3221708f416e963a0d4bb5499c932f517bceb8d2183e256060da44c9aa5af891

Observation 8ca8adbe-9cf7-43b5-8ded-9f3a74a0f64d · outbound

This paper cites an unresolved cited work.

Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-26T02:26:30.154013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:07:43.507761Z digest=sha256:49fe28e6560c3d831c80cc45ac004d67b6b8a345adf5fe29c75520e4e726c86b

Pith citing papers

Observation 532cbc44-7530-4413-97c4-723818559cab · inbound

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops cites this paper.

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:45:55.536724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T03:36:57.168246Z digest=sha256:0c0f07485b79b39cc1046429681bd0e7bdbb7583585cfe3fc3b48d428bb80bb3