Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:31:40.971330Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.09324.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:31:40.971330Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f30a8904-d07d-4dbf-95f9-3d25598e0620 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Training Verifiers to Solve Math Word Problems
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2154c012-1a09-493b-9037-3a2dcf974eea · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4c865d0-7b3f-4aab-9ca8-d369bf9b13a4 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning OpenAI o1 System Card
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eb025ee-b9be-4d07-b61a-2a2976290f80 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Mistral 7B
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67e570e3-ef33-40e8-8ab1-57cd9d5de482 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Language Models (Mostly) Know What They Know
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3763d5af-c35b-40e8-8efa-1acead1f0eaa · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Let’s verify step by step
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0ca7ff60-b193-43ea-80f3-3c4a1b3e7ec9 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Teaching Models to Express Their Uncertainty in Words
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecf3ec00-7a5e-45db-8cae-bce95f165443 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9918b9b5-9c6c-477e-8735-df3179ae6635 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f32949-0c2d-416b-b2ca-d565972e4262 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 972f5e9b-f015-43b8-ad1d-a35d8b32f25f · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a0a4733-879f-44ea-9178-5e4fdd7908e8 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 106e879b-f55d-4cd7-96c1-d27ab1b02433 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed71015f-8220-4ee6-bd25-0d48f95b0b1d · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Solving math word problems with process- and outcome-based feedback
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f47bff3-84ae-434f-b9e6-1be1c7a93777 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Tent: Fully Test-time Adaptation by Entropy Minimization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da00f497-998e-4b43-ac65-8843f7d92cf8 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Self-training with Noisy Student improves ImageNet classification
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db26dd7a-caae-4063-b2eb-074cbbf4155f · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3a74c6b-e452-4117-890a-58820d90bbdf · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Qwen3 Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb06f9fc-655a-4d51-8dd5-4a8d5a17b7ac · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Self-Rewarding Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9294ac53-e399-4e84-abc6-f6ad3434aec7 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Learning to Reason without External Rewards
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ababe61-88b4-4790-87a4-14418dfe98ec · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning B Use of Large Language Models Large language models were used only to improve grammar and clarity in author-written text
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3b0560b6-dc2c-4861-89c7-b1767232d739 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning role": "user
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b1e51fd3-abc4-48c4-bc3e-34771d4744c9 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Training language models to follow instructions with human feedback
Reference 1965
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 439abf01-d528-4a3e-b250-662f26ec332d · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 1978
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b567eace-03b0-4e05-a22e-705ae91c44b9 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 1988
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5224e1c-488d-4bd6-a3a4-87d6c2f7adf2 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 1997
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8155807b-da34-401a-897e-8159a418b84f · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Maximizing Confidence Alone Improves Reasoning
Reference 2007
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b76d1b3-2956-4e92-b2d1-d42c60fc789a · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Can large reasoning models self-train?arXiv preprint arXiv:2505.21444,
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8dc9f2c-f800-4540-8084-e3f16997dd95 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceb19587-0ff7-4c21-a5df-33c004284db0 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21b7eeaa-3378-4cf0-9f30-494b318e321e · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c133695-b7c3-40d1-b1fc-7bb7dea348f8 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning The Llama 3 Herd of Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f64ae17d-0407-4c4c-917c-e2cc6e5b07d0 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Concrete Problems in AI Safety
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2169df09-6c62-4dc3-81d5-3b3baee10156 · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b466cfc6-3bbf-4847-b6d6-ea2c260e53cd · outbound
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Co-rewarding: Stable self-supervised rl for eliciting reasoning in large language models
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.