Pith. sign in

Paper Citation Record · LEDGER

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

As of 15 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 7 inbound Pith citation observations for arXiv:2505.22960.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22960 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:02:45.317745Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:17:50.828950Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:09:33.482951Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 50165dfc-02ec-42c5-b4e2-4589e29b9d64 · outbound

This paper cites Critique-out-Loud Reward Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Critique-out-Loud Reward Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:41.917221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:41.917221Z digest=sha256:18a57318f86fb98dbcee6c7243c85d785054ebb00500ee945e01ab45b009d9f8

Observation f7bcbcfb-5cb3-40a2-a7fc-a9c9d564f61c · outbound

This paper cites AIME Problems and Solutions, 2025.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness AIME Problems and Solutions, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.974827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:02:41.963916Z digest=sha256:9e5df258383bb73e49497dae8c36a3858d0f515226468e12f8a2989dded0f4d9

Observation a3736bdf-ec6e-4833-ac82-cf84d1106980 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.027532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.027532Z digest=sha256:241801bcb3afaab94cf45806c959aeca871d95a4c643bc4a9e7171f1c744b756

Observation eb61c74b-3ac3-42f5-a418-7f9e56dd0bfa · outbound

This paper cites Why Do Multi-Agent LLM Systems Fail?.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Why Do Multi-Agent LLM Systems Fail?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.095621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.095621Z digest=sha256:d8ff5ba20b52febc70264db20c093cf915417b4aeeba33d2745a3741fa2696f1

Observation 7b8ba3e2-d821-4e8a-8e6a-24cf24afd817 · outbound

This paper cites Reconcile: Round-table conference improves reasoning via consensus among diverse llms.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Reconcile: Round-table conference improves reasoning via consensus among diverse llms

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.822189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:02:42.153098Z digest=sha256:a5dfb4859ddfdf28d159128b19ccc3b0dc948d34a8bc4e0eb9deb058148c7e45

Observation 24d4a73d-fccf-4b25-9b77-7150adbbbef6 · outbound

This paper cites Combating Adversarial Attacks with Multi-Agent Debate.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Combating Adversarial Attacks with Multi-Agent Debate

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.213315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.213315Z digest=sha256:b1adffef33a8ec808dfcb45f0077d0a427116186711f38d59f7e0a46e0290bf7

Observation ef59e14e-97b5-44fc-83c7-1e88397b29d5 · outbound

This paper cites Enhancing LLM Performance Through Debate: An Empirical Study on Multi-Agent Debate for Coding Tasks.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Enhancing LLM Performance Through Debate: An Empirical Study on Multi-Agent Debate for Coding Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.317472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.317472Z digest=sha256:fe142fdc067c8074d497ea3b9286ab22fdc93087d61b6c2c1c253d058438dbeb

Observation 0da4f06b-ae12-41e0-b5d5-8edcd3c37a3e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.391929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.391929Z digest=sha256:f1749c99407a4c15793436f599d27fb40ebb5be85df6630bec7e7c3cad3be04c

Observation 51a40f0e-8052-4add-906c-6f89b7618adc · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Multilingual Jailbreak Challenges in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.474919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.474919Z digest=sha256:1456788549067aab4b5a04eb6c3c6735391121aeb03604b29fbdddd941d4631b

Observation 157f11ae-88c1-4fce-959d-54bc007b4d75 · outbound

This paper cites Improving factuality and reasoning in language models through multiagent debate.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Improving factuality and reasoning in language models through multiagent debate

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.617588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:02:42.545982Z digest=sha256:36efbd8d26f3d6c481e9154824a38c2a412d8ac47d47f6ad4c8a7266484b6b8b

Observation 71e6a7c1-d19e-4968-9609-84f644783179 · outbound

This paper cites Multi-LLM debate: Framework, principals, and interventions.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Multi-LLM debate: Framework, principals, and interventions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.448146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:02:42.619790Z digest=sha256:3e30f70c10ff834d09edcf4323c1264ed42676b044e5c3de20ba5aedc2631b30

Observation 9fa65451-ec23-49c2-b288-95a87761aeac · outbound

This paper cites The Llama 3 Herd of Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.686078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.686078Z digest=sha256:8ff6ba6d74e2fb2d623eb3b4448fbf063857b557fd48934fbc1ece8002532375

Observation 28101e25-432d-44fc-bff6-3da30de146ba · outbound

This paper cites An empirical analysis of compute-optimal large language model training.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness An empirical analysis of compute-optimal large language model training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.781403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.781403Z digest=sha256:60ed26f3a060321a5bcc2375e09fd4ae8fddea642a83dfd865ea1ec6441d8259

Observation 7cc32020-f4cf-4c49-9778-ccfe62600de2 · outbound

This paper cites The curious case of neural text degeneration.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness The curious case of neural text degeneration

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:48.205048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:02:42.875033Z digest=sha256:c368997c1f3750fb7b305189359ef84316ab2de67fcaa5cc2d8823f535762ab6

Observation 91b8e223-e1c5-4e6c-bfcf-6fe735e1549b · outbound

This paper cites T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:42.962354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:42.962354Z digest=sha256:167c7d88b406e8126cba8e2225a48ed48784c7e8af16e448fd5e2e08ec5ce830

Observation 11a8cff6-5d82-4963-a225-0b949976c3b5 · outbound

This paper cites GPT-4o System Card.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.051383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.051383Z digest=sha256:4965c97917f957cb37c8c05f2fa13c88bd8508d73aee49541e798e8f41b1ece3

Observation cb22c870-4a60-4254-a16f-2110294f554c · outbound

This paper cites Scaling Laws for Neural Language Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Scaling Laws for Neural Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.128633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.128633Z digest=sha256:cb6f0aef1d59075ecbac99d84a86a9995ce619a4114eb5e596ca83da241d7b66

Observation b8eb9045-ce00-42b4-8d31-5e0cb97648ca · outbound

This paper cites Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.211841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.211841Z digest=sha256:bcc82cc1d5333918c197409312ba73a812d9bed08891465bd99c107bb1d277f4

Observation f4ad240d-ab0d-47f3-89c7-cd5e72f2eb31 · outbound

This paper cites A Simple Model of Inference Scaling Laws.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness A Simple Model of Inference Scaling Laws

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.310559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.310559Z digest=sha256:4946f2416061941bd4954d0152691f2dabb781fbe2900878551ee47d6a6bf821

Observation caef8f91-cf52-4165-ae02-bf8114cb98fe · outbound

This paper cites Encouraging divergent thinking in large language models through multi-agent debate.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Encouraging divergent thinking in large language models through multi-agent debate

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.390182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.390182Z digest=sha256:94b5f9fb1af47171072bec7510316ae44ace017b61e2546d9cbf4f7a8dcee4a7

Observation 9576c0c3-7c93-4a84-b6aa-2f631dc9a626 · outbound

This paper cites Let’s verify step by step.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Let’s verify step by step

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.995037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:02:43.461797Z digest=sha256:7509331aa9ebd25cd6c464589d9c032cb897cd52275d87ee86ab4674c361bf39

Observation 991a5605-1452-4318-974a-5ebd74924ed7 · outbound

This paper cites Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.514780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.514780Z digest=sha256:a66ae8d36a4d91177cbb20bf228e6556a362591ff8516c744a30101810c3d17e

Observation 6892ea3c-2473-4033-95cb-f5b7f5aa43d2 · outbound

This paper cites Breaking mental set to improve reasoning through diverse multi-agent debate.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Breaking mental set to improve reasoning through diverse multi-agent debate

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.806699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:02:43.622627Z digest=sha256:13811b3e6951e25b8ac184bb80a9284160ec7747a710e588c24c182b1b66d2dc

Observation 3f9e51d1-fd16-4550-bf19-7f0c113d4437 · outbound

This paper cites Large Language Model Guided Tree-of-Thought.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Large Language Model Guided Tree-of-Thought

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.689015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.689015Z digest=sha256:06fb747d1257cace47cc225385e75c594926dabb5ca7bf6efe6da0022284d147

Observation eed94867-4db9-458d-afd1-9f9147eb6b66 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Self-refine: Iterative refinement with self-feedback

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.629374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:02:43.809755Z digest=sha256:a1d9346c88fe9827a7679f0e2241e27577cc38055b08819582b8839a1bb3ad6b

Observation b05f67d7-eaa3-46a8-beea-016adb6c588f · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:43.876587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:43.876587Z digest=sha256:24ad346cdaea47b7b248af357e5fac64790dff5e8e453ed6639b366d71c4039e

Observation b63599e5-2e99-4f04-8f44-090da8d8c1a9 · outbound

This paper cites Should we be going mad? a look at multi-agent debate strategies for llms.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Should we be going mad? a look at multi-agent debate strategies for llms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.361841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:02:43.943342Z digest=sha256:9935b9165c587ff41b7cb7a920491df9adb13786fe30828819810e6f88873f63

Observation 0447339c-e3d6-426b-8e53-d9e5eb0858da · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.017112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.017112Z digest=sha256:9a94259ee36f99e7b77890ce0e5cc765a1f03106550c522506dde7e7e2ed544a

Observation 3649c79e-4a22-416f-8074-3946eb383c4b · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Gemma 2: Improving Open Language Models at a Practical Size

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.090457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.090457Z digest=sha256:b210a887f709187d444e0c3bc43b4c0c6078848b843f4e8cec0c79ba5347650c

Observation 4e7a4394-8d00-4085-a812-8ebbb20ad4f0 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.201364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.201364Z digest=sha256:b0af4d161eb73308dd7a041834704e3094b600fe95a3bba345797c815e0da546

Observation 43b7cf01-5e1c-41af-a77e-d9579df47a4a · outbound

This paper cites an unresolved cited work.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.323471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.323471Z digest=sha256:d594911b4a883e00c574a67096b1c69464ea14bd4e9a7a4a4613b8edac15a873

Observation fe252aa3-127c-4995-a1a4-238077d50c11 · outbound

This paper cites Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.400972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.400972Z digest=sha256:aeab6ce0e8877d3f77c272b55097ff1f21cd9a3bc75c6fd6cc56fc1e8c4eaa09

Observation 771059ea-f15e-46f0-9a36-fc16aa64452d · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Chain-of-thought prompting elicits reasoning in large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.474136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.474136Z digest=sha256:f45a22fdec8f53f038a9f2a76bc9f00355b98763a0055a4d6ef31c6f3f8eef16

Observation aec96bbc-035d-436b-b99f-9d49c1763475 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.548349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.548349Z digest=sha256:1191b99c0b832206397cff6096a88c21c35fe8bb693acba2731290db4b073402

Observation 712a0928-6fb9-4e55-9655-c1e19a20b67b · outbound

This paper cites Self-evaluation guided beam search for reasoning.Advances in Neural Information Processing Systems, 36:41618–41650, 2023.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Self-evaluation guided beam search for reasoning.Advances in Neural Information Processing Systems, 36:41618–41650, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.645060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.645060Z digest=sha256:efe203c5fe04cd22a1cbcdcd1179f84fc1aa619153293beec5a73c93469fc653

Observation 645603a2-e112-4f1b-98b8-34c3f1155150 · outbound

This paper cites DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.713305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.713305Z digest=sha256:d0514e220b224d6deb11f860e374ec6126de21612debf0d4f5d0ec3f740ebb87

Observation 86185fd4-e7d6-4378-96c3-c525adff82e7 · outbound

This paper cites Qwen2.5 Technical Report.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Qwen2.5 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:44.890454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:44.890454Z digest=sha256:80b75a69e1cd137e1bc21448d57cb6429219821e35ea715b960d7bb1c8435411

Observation 3f4324b8-7572-49a2-bbf9-870a6f3f9d87 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Tree of thoughts: Deliberate problem solving with large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:47.091050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:02:45.024427Z digest=sha256:e4d58883d6006b6efe6538cc945508a40fde395d36bf00a99a52f595512d3fdb

Observation 0fc19b89-482a-495c-89d1-d4e62f7c113d · outbound

This paper cites Csrt: Evaluation and analysis of llms using code-switching red-teaming dataset.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Csrt: Evaluation and analysis of llms using code-switching red-teaming dataset

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:46.896637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:02:45.101351Z digest=sha256:74a158f0d50e2694a617a7a351d8aa391e4c6919346fc7a5e6557d4f330a38c0

Observation 35c4c82a-a13b-483a-8835-afb609c51de4 · outbound

This paper cites AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:45.177061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:45.177061Z digest=sha256:40e4375a1191131fdec02103919c9d8c73872c07a81841aee39ffa94b1b1459a

Observation 5210efc6-e165-4d9a-829c-9dfb60662f0b · outbound

This paper cites Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:45.265378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:45.265378Z digest=sha256:b0e2fc7c3ab05311492abbd5067bb3073fa1baa12689631978e1f932137e81c1

Observation 3c178dd5-8622-43ff-a3da-77473e2e7b62 · outbound

This paper cites I’m sorry, but I can’t assist with creating content that promotes hate, racism, or any form of discrimination.

Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness I’m sorry, but I can’t assist with creating content that promotes hate, racism, or any form of discrimination

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:46.675825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:02:45.317745Z digest=sha256:9f2951c0cd14ee8ade067c62013f22b6f38e3f4af3f75a1012aa4df5bcde2c1e

Pith citing papers

Observation 2a2a4fa0-013f-4913-9e00-fa38553442cf · inbound

Free-MAD: Consensus-Free Multi-Agent Debate cites this paper.

Free-MAD: Consensus-Free Multi-Agent Debate Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T17:15:01.617748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:15:01.617748Z digest=sha256:41c8e2907969c3fd39d8df3cd3a49ffb7a8676f9c881e85ed141d8702472dbe8

Observation eb9a11b9-1f9e-449a-a1d0-539a9aea92c3 · inbound

The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning cites this paper.

The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:03.999666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T15:52:43.274993Z digest=sha256:4f1289ddf93c7b2ec166fb46060c5ad41e531d56f91651ad611660d9272f0cbd

Observation a4491ec6-89d0-45db-8255-668369926b0a · inbound

Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate cites this paper.

Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:52:35.927613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T18:46:50.409655Z digest=sha256:a87942c9b41a45be90ca3d3c1c2abc4ca1ac2d741d77af4914e733800f2d8700

Observation ad444918-9aa2-49bb-8725-23f1dae9f482 · inbound

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size cites this paper.

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:56:15.586924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T16:04:37.244398Z digest=sha256:a5825cc2b99d6d0a641565fc8fed077b17bc81edec83ad33373451ab5f276c57

Observation fb4142b4-7843-4053-a9fa-48531c202511 · inbound

When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning cites this paper.

When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:36:23.855401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T14:05:49.696342Z digest=sha256:51bf79a05ab1dbd320623f95d48558477364a8a5f03c04839a5687512f032c13

Observation b2ac482a-8485-4275-a117-6a62dac863b2 · inbound

Heterogeneous LLM Debate Under Adversarial Peers: Honest Gains, Replacement Costs, and Resilience cites this paper.

Heterogeneous LLM Debate Under Adversarial Peers: Honest Gains, Replacement Costs, and Resilience Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:09:33.490797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T17:17:05.107942Z digest=sha256:283c2a5647118bb00f74449a0af911d5817244badcdb3e23c8c618055782f69b

Observation 852dd9ca-6b3f-4588-8ce2-253d9e9171e9 · inbound

Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology cites this paper.

Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:17:50.828950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:17:50.828950Z digest=sha256:d1b8bdd1bd8dbf6614202a2e40d2eabb98c69eebe2e9a11af1e747244ee8ba8a