Pith. sign in

Paper Citation Record · LEDGER

Coupled Variational Reinforcement Learning for Language Model General Reasoning

As of 15 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2512.12576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.12576 v3

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T16:43:58.144921Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3497387c-12ee-49bb-801a-3631be7a3bb7 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Process Reinforcement through Implicit Rewards

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:54.836145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:54.836145Z digest=sha256:34c4fa7d55834da080262da6bdf78171e236a0c76304f3fe0bb29ddb9cdfd60b

Observation 1a5dfe83-8427-40a7-b41e-8dadf5934bd3 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Coupled Variational Reinforcement Learning for Language Model General Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:54.952300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:54.952300Z digest=sha256:33c1101e156344a613f8e48d1c09bfdc46a0275ccbb1fb662188363dd23c908c

Observation 40e1fcc6-1eaf-44c7-925f-cc05714f5343 · outbound

This paper cites Auto-Encoding Variational Bayes.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Auto-Encoding Variational Bayes

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:55.183222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:55.183222Z digest=sha256:36d7d9725b57e23819fb21888f07b34e4a7ca5af59807f810790dd97997610a3

Observation 6c5b8c08-1601-4292-88d2-341ce52d9524 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Solving Quantitative Reasoning Problems with Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:55.364185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:55.364185Z digest=sha256:c982177f61b25ca7cd31b4011fed3dd65e1b8b8a773b35ddc9ffec3e06eddc04

Observation 2290000d-250f-4ada-bb46-e38ec905da98 · outbound

This paper cites Ling, W., Yogatama, D., Dyer, C., and Blunsom, P.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Ling, W., Yogatama, D., Dyer, C., and Blunsom, P

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:55.743062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:55.743062Z digest=sha256:1dae3b4bcf09ab1e777d174186328e189ac8077f92a47d11adf694d6b7c72149

Observation f3cbd91c-f7c8-41f4-983f-73cbf847eeea · outbound

This paper cites General-Reasoner: Advancing LLM Reasoning Across All Domains.

Coupled Variational Reinforcement Learning for Language Model General Reasoning General-Reasoner: Advancing LLM Reasoning Across All Domains

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.063243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.063243Z digest=sha256:9f427bcac808bec8759d5d7217919aa99b79451fef06dee45ddf340de460402b

Observation 37f9f366-a0f3-4470-bc4d-9f6148da8e13 · outbound

This paper cites Con- tains 32 math questions from May 2023 SAT.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Con- tains 32 math questions from May 2023 SAT

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.213416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.213416Z digest=sha256:de01655dc08088345261ba6a0328144e47eeb0241b8d6c899d84146ee06787f6

Observation ca7b58c3-e35f-410a-ae60-dae9cfdb3602 · outbound

This paper cites Training language models to follow instructions with human feedback.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Training language models to follow instructions with human feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.345427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.345427Z digest=sha256:287c575d32080b52e0451a8f68c822e62c1fe1d37673ea5e56c807fcdfb2e937

Observation 70c9bf1e-8376-4080-876d-81ff21b63c2a · outbound

This paper cites Qwen2.5 Technical Report.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Qwen2.5 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.535548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.535548Z digest=sha256:bfcd895d229dfd79e23a91fddebaa0b5f72c449a19f0352a4cf57641b0a73f8f

Observation 03e68c57-8a5d-4486-ac96-719212089c52 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.609302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.609302Z digest=sha256:529359eaa6e9959bda85dfcfb067553e7c16662bb7bbf01dcc5ca2096b53f359

Observation cbc0d829-1989-4f9e-a83d-b0726aab6594 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Coupled Variational Reinforcement Learning for Language Model General Reasoning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.708534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.708534Z digest=sha256:a4ad3d1f8fe92dfb546bb35dbf8d0d2d55af79d30635a6a9cda6ae35648fea65

Observation e54c9422-7261-488f-8b32-076fb86ce597 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.772053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.772053Z digest=sha256:ddf3078fa57d39dc4b4a8ef8a6f2888b784b8aece14b09479b3062a623967eab

Observation 9d8ee088-8d68-4994-a5ee-8f3e5d9904f4 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Coupled Variational Reinforcement Learning for Language Model General Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.842532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.842532Z digest=sha256:0f2c8aaf96ec0ed8cc4b98946b726f4b539f66134548e3b31f8fc500ebc724ab

Observation 571222e9-820a-477a-ac3c-11adea58878b · outbound

This paper cites Defining and Characterizing Reward Hacking.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Defining and Characterizing Reward Hacking

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.945344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.945344Z digest=sha256:b0e8424a4455a6a18c77da467fba4dbb89ba7d486b651325f8ba28ae708436b8

Observation 9793575a-5edd-46e4-a648-7609fd8e5b4f · outbound

This paper cites Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.074776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.074776Z digest=sha256:b7e0db5f8cee50ddf25c395839a12997b4b622655a664e071d7cffe8b206d63a

Observation 4097a892-2d58-475f-8c93-65fb44437048 · outbound

This paper cites Accessed: 2025-01-23.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Accessed: 2025-01-23

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.273034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.273034Z digest=sha256:f97b8b43ab78ec9b60cb947756f58a0f8c8aa0a899aec2046c8460ed60c30a49

Observation 1bb6e91e-f99c-434f-94c0-21869bc5e718 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Coupled Variational Reinforcement Learning for Language Model General Reasoning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.348223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.348223Z digest=sha256:89a4e68d564986c1e78c051ca0bd6233d9a9cee68f8bdebbceaf1fd4e97e7750

Observation 905da558-301d-4278-97b4-c12f208f8ab1 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.528598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.528598Z digest=sha256:97f347c1ba3fc2081875e9fd491e74fe28cd3e3d82edafcb6825c56610c8e2a5

Observation 02882fce-fdb2-4124-8bec-9b46003db01d · outbound

This paper cites Qwen3 Technical Report.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Qwen3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.687482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.687482Z digest=sha256:f1813f5643161652a9ab3492821416162cd9cb000891f0a0f55328696845a6a5

Observation 6b66f49f-04bf-4706-aedb-d3cab0e19629 · outbound

This paper cites Self-Rewarding Language Models.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Self-Rewarding Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.799392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.799392Z digest=sha256:037237871dd37b1a93f1eceb6652b9e52764499fdb20dcc19eaa7436cdaf4234

Observation aae23b8d-5928-4e2a-9202-60cc7bc8c165 · outbound

This paper cites MAmmoTH2: Scaling Instructions from the Web.

Coupled Variational Reinforcement Learning for Language Model General Reasoning MAmmoTH2: Scaling Instructions from the Web

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.892407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.892407Z digest=sha256:d4a33a5146f46e7a73a7074203fe3549f530e0c0f8f6d182b082c217a3d2686e

Observation 42998865-32a2-4ce8-880b-7e501606d255 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Coupled Variational Reinforcement Learning for Language Model General Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.954210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.954210Z digest=sha256:5169b657faf88f3be83b852c6724254a11047972670858d57f82da0bdf9eecc5

Observation 4ad521f6-daf6-494d-ba12-97f25ef4b993 · outbound

This paper cites Evaluating and Improving Tool-Augmented Computation-Intensive Math Reasoning.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Evaluating and Improving Tool-Augmented Computation-Intensive Math Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:58.027815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:58.027815Z digest=sha256:cc4420b8a495ddf01287d318cd0bf0d3e5be36ff9e9e5101fa53526809a88dc0

Observation 60c549bd-ce14-43bf-b44b-f2a72e3a1bf3 · outbound

This paper cites Zhou, X., Liu, Z., Sims, A., Wang, H., Pang, T., Li, C., Wang, L., Lin, M., and Du, C.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Zhou, X., Liu, Z., Sims, A., Wang, H., Pang, T., Li, C., Wang, L., Lin, M., and Du, C

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:58.082657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:58.082657Z digest=sha256:6382f7d1edb362568b47dc4e863f7aacbc73e4cb2d01538966e4dbdb14592212

Observation 44544868-b68d-4992-bbe6-c746fd65ae14 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Coupled Variational Reinforcement Learning for Language Model General Reasoning TTRL: Test-Time Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:58.144921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:58.144921Z digest=sha256:aa4a477cff3d4ee8b8ad643bb2f5924aec7016b1660a0b53e0ef1890fc992fd9

Observation 7358163e-dae2-4d17-86cb-5f3ed4e362ae · outbound

This paper cites Importance Weighted Autoencoders.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Importance Weighted Autoencoders

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:54.075263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:54.075263Z digest=sha256:5a9e2350a00b9d5c9cfe613ab3bd4330295d9212537493df1f70308d960c9d13

Observation f7cb0a4a-63d4-41e5-bbd3-3ede3214bafc · outbound

This paper cites Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:55.921050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:55.921050Z digest=sha256:10728b377b973a82d93d1111bbbb7b1981ed684754871543dcf67f75b16b2e29

Observation c261e508-4925-434a-9d17-c4254acb339f · outbound

This paper cites A Stable Variational Autoencoder for Text Modelling.

Coupled Variational Reinforcement Learning for Language Model General Reasoning A Stable Variational Autoencoder for Text Modelling

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:55.552728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:55.552728Z digest=sha256:46e1b861d70ffb42b7f9f78fe98ecb0a44d66ee76e02ba6a710db7f1ff0ec8c4

Observation 2ce08a28-16e3-462b-afe8-c0500fc395d6 · outbound

This paper cites Denoising Diffusion Probabilistic Models.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Denoising Diffusion Probabilistic Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:55.056382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:55.056382Z digest=sha256:49935e95b0dd1a024f21f765c0f8ef51a467847f0b146123498fffae8baced27

Observation 7cb0e51f-c2a5-4a38-9c2a-988bb9120003 · outbound

This paper cites Greedification Operators for Policy Optimization: Investigating Forward and Reverse KL Divergences.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Greedification Operators for Policy Optimization: Investigating Forward and Reverse KL Divergences

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:54.214636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:54.214636Z digest=sha256:c95f59b57fac3cf839a9ad0c8a0e2640db2d11699424ce2d97ecc1a969810b27

Observation c0545efa-1a4e-4f6a-8a2b-a39b66788a1a · outbound

This paper cites TheoremQA: A Theorem-driven Question Answering dataset.

Coupled Variational Reinforcement Learning for Language Model General Reasoning TheoremQA: A Theorem-driven Question Answering dataset

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:54.710371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:54.710371Z digest=sha256:4ab6cec7a7951b45b39287ec290f08d4beb4ecf3073a6385ff65ba64cdcbc228

Observation e026c71f-ed98-4660-954a-bb8ca40016ef · outbound

This paper cites Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:54.497117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:54.497117Z digest=sha256:057b750dbbc10dae9d894010dc2632189b8fd1e5272193a18850f23b32ced151

Observation 5b33c3c0-dcd0-458f-92a1-1ca91077afc9 · outbound

This paper cites Bootstrapping Language Models with DPO Implicit Rewards.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Bootstrapping Language Models with DPO Implicit Rewards

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:54.381657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:54.381657Z digest=sha256:85848e46839fe760383d43f81036590d19237313d607c8a468fd77fc2ff7748a

Pith citing papers

No inbound Pith citation observations are available.