Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:02:45.317745Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 6 inbound Pith citation observations for arXiv:2505.22960.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:02:45.317745Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T17:15:01.617748Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T04:09:33.482951Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 50165dfc-02ec-42c5-b4e2-4589e29b9d64 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Critique-out-Loud Reward Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7bcbcfb-5cb3-40a2-a7fc-a9c9d564f61c · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness AIME Problems and Solutions, 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a3736bdf-ec6e-4833-ac82-cf84d1106980 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb61c74b-3ac3-42f5-a418-7f9e56dd0bfa · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Why Do Multi-Agent LLM Systems Fail?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b8ba3e2-d821-4e8a-8e6a-24cf24afd817 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Reconcile: Round-table conference improves reasoning via consensus among diverse llms
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 24d4a73d-fccf-4b25-9b77-7150adbbbef6 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Combating Adversarial Attacks with Multi-Agent Debate
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef59e14e-97b5-44fc-83c7-1e88397b29d5 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Enhancing LLM Performance Through Debate: An Empirical Study on Multi-Agent Debate for Coding Tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0da4f06b-ae12-41e0-b5d5-8edcd3c37a3e · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51a40f0e-8052-4add-906c-6f89b7618adc · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Multilingual Jailbreak Challenges in Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 157f11ae-88c1-4fce-959d-54bc007b4d75 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Improving factuality and reasoning in language models through multiagent debate
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 71e6a7c1-d19e-4968-9609-84f644783179 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Multi-LLM debate: Framework, principals, and interventions
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9fa65451-ec23-49c2-b288-95a87761aeac · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28101e25-432d-44fc-bff6-3da30de146ba · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness An empirical analysis of compute-optimal large language model training
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cc32020-f4cf-4c49-9778-ccfe62600de2 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness The curious case of neural text degeneration
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 91b8e223-e1c5-4e6c-bfcf-6fe735e1549b · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11a8cff6-5d82-4963-a225-0b949976c3b5 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness GPT-4o System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb22c870-4a60-4254-a16f-2110294f554c · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Scaling Laws for Neural Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8eb9045-ce00-42b4-8d31-5e0cb97648ca · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4ad240d-ab0d-47f3-89c7-cd5e72f2eb31 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness A Simple Model of Inference Scaling Laws
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caef8f91-cf52-4165-ae02-bf8114cb98fe · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Encouraging divergent thinking in large language models through multi-agent debate
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9576c0c3-7c93-4a84-b6aa-2f631dc9a626 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Let’s verify step by step
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 991a5605-1452-4318-974a-5ebd74924ed7 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6892ea3c-2473-4033-95cb-f5b7f5aa43d2 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Breaking mental set to improve reasoning through diverse multi-agent debate
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3f9e51d1-fd16-4550-bf19-7f0c113d4437 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Large Language Model Guided Tree-of-Thought
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eed94867-4db9-458d-afd1-9f9147eb6b66 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Self-refine: Iterative refinement with self-feedback
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b05f67d7-eaa3-46a8-beea-016adb6c588f · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b63599e5-2e99-4f04-8f44-090da8d8c1a9 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Should we be going mad? a look at multi-agent debate strategies for llms
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0447339c-e3d6-426b-8e53-d9e5eb0858da · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3649c79e-4a22-416f-8074-3946eb383c4b · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Gemma 2: Improving Open Language Models at a Practical Size
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e7a4394-8d00-4085-a812-8ebbb20ad4f0 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43b7cf01-5e1c-41af-a77e-d9579df47a4a · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe252aa3-127c-4995-a1a4-238077d50c11 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 771059ea-f15e-46f0-9a36-fc16aa64452d · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Chain-of-thought prompting elicits reasoning in large language models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aec96bbc-035d-436b-b99f-9d49c1763475 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 712a0928-6fb9-4e55-9655-c1e19a20b67b · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Self-evaluation guided beam search for reasoning.Advances in Neural Information Processing Systems, 36:41618–41650, 2023
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 645603a2-e112-4f1b-98b8-34c3f1155150 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86185fd4-e7d6-4378-96c3-c525adff82e7 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Qwen2.5 Technical Report
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f4324b8-7572-49a2-bbf9-870a6f3f9d87 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Tree of thoughts: Deliberate problem solving with large language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0fc19b89-482a-495c-89d1-d4e62f7c113d · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Csrt: Evaluation and analysis of llms using code-switching red-teaming dataset
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 35c4c82a-a13b-483a-8835-afb609c51de4 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5210efc6-e165-4d9a-829c-9dfb60662f0b · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c178dd5-8622-43ff-a3da-77473e2e7b62 · outbound
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness I’m sorry, but I can’t assist with creating content that promotes hate, racism, or any form of discrimination
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2a2a4fa0-013f-4913-9e00-fa38553442cf · inbound
Free-MAD: Consensus-Free Multi-Agent Debate Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb9a11b9-1f9e-449a-a1d0-539a9aea92c3 · inbound
The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a4491ec6-89d0-45db-8255-668369926b0a · inbound
Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ad444918-9aa2-49bb-8725-23f1dae9f482 · inbound
The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fb4142b4-7843-4053-a9fa-48531c202511 · inbound
When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b2ac482a-8485-4275-a117-6a62dac863b2 · inbound
Heterogeneous LLM Debate Under Adversarial Peers: Honest Gains, Replacement Costs, and Resilience Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.