Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T16:43:58.144921Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2512.12576.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T16:43:58.144921Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3497387c-12ee-49bb-801a-3631be7a3bb7 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Process Reinforcement through Implicit Rewards
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a5dfe83-8427-40a7-b41e-8dadf5934bd3 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e1fcc6-1eaf-44c7-925f-cc05714f5343 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Auto-Encoding Variational Bayes
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c5b8c08-1601-4292-88d2-341ce52d9524 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Solving Quantitative Reasoning Problems with Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2290000d-250f-4ada-bb46-e38ec905da98 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Ling, W., Yogatama, D., Dyer, C., and Blunsom, P
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3cbd91c-f7c8-41f4-983f-73cbf847eeea · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning General-Reasoner: Advancing LLM Reasoning Across All Domains
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37f9f366-a0f3-4470-bc4d-9f6148da8e13 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Con- tains 32 math questions from May 2023 SAT
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca7b58c3-e35f-410a-ae60-dae9cfdb3602 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Training language models to follow instructions with human feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70c9bf1e-8376-4080-876d-81ff21b63c2a · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Qwen2.5 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03e68c57-8a5d-4486-ac96-719212089c52 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbc0d829-1989-4f9e-a83d-b0726aab6594 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e54c9422-7261-488f-8b32-076fb86ce597 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Proximal Policy Optimization Algorithms
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d8ee088-8d68-4994-a5ee-8f3e5d9904f4 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 571222e9-820a-477a-ac3c-11adea58878b · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Defining and Characterizing Reward Hacking
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9793575a-5edd-46e4-a648-7609fd8e5b4f · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4097a892-2d58-475f-8c93-65fb44437048 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Accessed: 2025-01-23
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bb6e91e-f99c-434f-94c0-21869bc5e718 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 905da558-301d-4278-97b4-c12f208f8ab1 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02882fce-fdb2-4124-8bec-9b46003db01d · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Qwen3 Technical Report
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b66f49f-04bf-4706-aedb-d3cab0e19629 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Self-Rewarding Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aae23b8d-5928-4e2a-9202-60cc7bc8c165 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning MAmmoTH2: Scaling Instructions from the Web
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42998865-32a2-4ce8-880b-7e501606d255 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ad521f6-daf6-494d-ba12-97f25ef4b993 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Evaluating and Improving Tool-Augmented Computation-Intensive Math Reasoning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60c549bd-ce14-43bf-b44b-f2a72e3a1bf3 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Zhou, X., Liu, Z., Sims, A., Wang, H., Pang, T., Li, C., Wang, L., Lin, M., and Du, C
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44544868-b68d-4992-bbe6-c746fd65ae14 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning TTRL: Test-Time Reinforcement Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7358163e-dae2-4d17-86cb-5f3ed4e362ae · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Importance Weighted Autoencoders
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7cb0a4a-63d4-41e5-bbd3-3ede3214bafc · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c261e508-4925-434a-9d17-c4254acb339f · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning A Stable Variational Autoencoder for Text Modelling
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ce08a28-16e3-462b-afe8-c0500fc395d6 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Denoising Diffusion Probabilistic Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb0e51f-c2a5-4a38-9c2a-988bb9120003 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Greedification Operators for Policy Optimization: Investigating Forward and Reverse KL Divergences
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0545efa-1a4e-4f6a-8a2b-a39b66788a1a · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning TheoremQA: A Theorem-driven Question Answering dataset
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e026c71f-ed98-4660-954a-bb8ca40016ef · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b33c3c0-dcd0-458f-92a1-1ca91077afc9 · outbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Bootstrapping Language Models with DPO Implicit Rewards
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.