Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:22:07.387448Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 3 inbound Pith citation observations for arXiv:2506.08379.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:22:07.387448Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T05:05:24.654803Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-08T05:05:25.028535Z
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0efa096b-d473-4a1e-87af-bf49807045c3 · outbound
Reinforce LLM Reasoning through Multi-Agent Reflection Therefore,A bπ h(sh, ah) = eAbπ h(sh, ah)
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 96680339-25da-4c3e-8ee3-c0be6630f002 · outbound
Reinforce LLM Reasoning through Multi-Agent Reflection We define the approximation error: ∆ =E sh∼dπ⋆ h ,ah∼π⋆(·|sh)[Aˆπ h(sh, ah)− eAˆπ h(sh, ah)]
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 29f7abc9-e6cb-45a5-b0d2-b9f06bb93792 · outbound
Reinforce LLM Reasoning through Multi-Agent Reflection KTO: Model Alignment as Prospect Theoretic Optimization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8cb2d5-ef93-44b2-81b2-f1841bd337c0 · outbound
Reinforce LLM Reasoning through Multi-Agent Reflection Exploring the Intersection of Large Language Models and Agent-Based Modeling via Prompt Engineering
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 128492c0-5a34-4de2-9ffd-649a728ba6c9 · outbound
Reinforce LLM Reasoning through Multi-Agent Reflection Orca-Math: Unlocking the potential of SLMs in Grade School Math
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c17167c2-90f9-431c-b86b-656d2be58f56 · outbound
Reinforce LLM Reasoning through Multi-Agent Reflection doi: 10.18653/v1/2024.naacl-long.327
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 04895f0d-918d-435f-bbd9-98701affebb3 · outbound
Reinforce LLM Reasoning through Multi-Agent Reflection Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b29a631-cbcb-498a-b6c3-87ff39f65aaf · outbound
Reinforce LLM Reasoning through Multi-Agent Reflection Therefore, Eah∼π⋆(·|sh)[Aˆπ h(sh, ah)]≈E ah∼π⋆(·|sh)[Aπ⋆ h (sh, ah)] = 0, where the last equality follows from the definition ofA π h
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 24cc874b-469e-4ca7-a8c1-920f6ffd2753 · outbound
Reinforce LLM Reasoning through Multi-Agent Reflection We then conduct DPO training on the critic, producing a refined critic modelbπc
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 030ab3c0-e20e-45e7-acd2-6ac2a409e09c · outbound
Reinforce LLM Reasoning through Multi-Agent Reflection Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7ace3744-985c-44d5-9248-4b73b41540d4 · outbound
Reinforce LLM Reasoning through Multi-Agent Reflection Justin Chih-Yao Chen, Archiki Prasad, Swarnadeep Saha, Elias Stengel-Eskin, and Mohit Bansal
Reference 862
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38a7ec8d-1f7a-4786-babf-f21180eb9c63 · outbound
Reinforce LLM Reasoning through Multi-Agent Reflection A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19fe975d-edb7-42b7-bf65-921f342a7812 · outbound
Reinforce LLM Reasoning through Multi-Agent Reflection ORPO: Monolithic Preference Optimization without Reference Model
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac843756-05ec-483a-b8e2-b6f5c403f49c · outbound
Reinforce LLM Reasoning through Multi-Agent Reflection CodeAgent: Autonomous Communicative Agents for Code Review
Reference 3021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 955d0567-4898-445d-a39e-c52286bacebd · inbound
Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving Reinforce LLM Reasoning through Multi-Agent Reflection
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dbf8087-3485-4e26-9fd2-531ae83c4e6d · inbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Reinforce LLM Reasoning through Multi-Agent Reflection
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c346b48-8005-4431-adde-304cc1ceb99b · inbound
Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning Reinforce LLM Reasoning through Multi-Agent Reflection
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.