Pith. sign in

Paper Citation Record · LEDGER

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning

As of 16 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.09324.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09324 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:31:40.971330Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f30a8904-d07d-4dbf-95f9-3d25598e0620 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.801100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.801100Z digest=sha256:1e66406f6e71240d4a67c6b752c3fa4b5a5e6fadd373f09ff1b8b8985e98245e

Observation 2154c012-1a09-493b-9037-3a2dcf974eea · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.819331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.819331Z digest=sha256:950bb2c116cdc89b754f98afdbc170f997c10c9465d0319ca29286bc3905e657

Observation b4c865d0-7b3f-4aab-9ca8-d369bf9b13a4 · outbound

This paper cites OpenAI o1 System Card.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning OpenAI o1 System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.828258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.828258Z digest=sha256:114713b820a74aa456e27932f8647466144fad1c85dcc154f00ac1d358590651

Observation 3eb025ee-b9be-4d07-b61a-2a2976290f80 · outbound

This paper cites Mistral 7B.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Mistral 7B

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.832494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.832494Z digest=sha256:cb0636a008fd29cc16ad0e2781005d13437a9cfb79d3328e4813f0d50d65cd91

Observation 67e570e3-ef33-40e8-8ab1-57cd9d5de482 · outbound

This paper cites Language Models (Mostly) Know What They Know.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Language Models (Mostly) Know What They Know

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.837316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.837316Z digest=sha256:3041e1d7539f1313a1a430ff9045d9d81cc49a95204303d10c4365f2f4d6e905

Observation 3763d5af-c35b-40e8-8efa-1acead1f0eaa · outbound

This paper cites Let’s verify step by step.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Let’s verify step by step

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:41.893764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:31:40.852283Z digest=sha256:dc7f583ba8a2681a7f6ac0e1ee6e4d7c8527191c4c2d0a769367867821522d52

Observation 0ca7ff60-b193-43ea-80f3-3c4a1b3e7ec9 · outbound

This paper cites Teaching Models to Express Their Uncertainty in Words.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Teaching Models to Express Their Uncertainty in Words

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.858193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.858193Z digest=sha256:5d1f62bdab6f9564d5435fe059e1ddd9309f03bd38562b74c961552395cfa1e7

Observation ecf3ec00-7a5e-45db-8cae-bce95f165443 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.864651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.864651Z digest=sha256:5b9d60ecc4717c8bb552dec432b4b41e206eca44a0bc29e99b06519b84deb8ca

Observation 9918b9b5-9c6c-477e-8735-df3179ae6635 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.875842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.875842Z digest=sha256:a46a4a1d27e569cee0f7b52078206280847413761c5d915c87a985605792f35d

Observation c8f32949-0c2d-416b-b2ca-d565972e4262 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.885338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.885338Z digest=sha256:151bdadff4fd2e15091529eef9d4a0263d9e36c77dabc7d8a6ddd149252aa12b

Observation 972f5e9b-f015-43b8-ad1d-a35d8b32f25f · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.900230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.900230Z digest=sha256:8d357fc101522933e639c6d9c7e2d83897923ddf0a262890eb937935e29597d1

Observation 4a0a4733-879f-44ea-9178-5e4fdd7908e8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.904736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.904736Z digest=sha256:909d6dd6babb3f0f88d3cf2777103858da9cd0eb0e72581d776769f02bb15efa

Observation 106e879b-f55d-4cd7-96c1-d27ab1b02433 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.909954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.909954Z digest=sha256:fd33375e5c1d49d5ab541727a4078b32363872b11df981d683365419ce57d8dd

Observation ed71015f-8220-4ee6-bd25-0d48f95b0b1d · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Solving math word problems with process- and outcome-based feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.922992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.922992Z digest=sha256:7bc38480c47b047f7d7b3ee69d1c1ba399c9bf132270225c62d032766f92e812

Observation 6f47bff3-84ae-434f-b9e6-1be1c7a93777 · outbound

This paper cites Tent: Fully Test-time Adaptation by Entropy Minimization.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Tent: Fully Test-time Adaptation by Entropy Minimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.928529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.928529Z digest=sha256:22adbb8a04c8b7bc394ea0d136f671ecc6aac4d865c62512f2c54fe1c551402e

Observation da00f497-998e-4b43-ac65-8843f7d92cf8 · outbound

This paper cites Self-training with Noisy Student improves ImageNet classification.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Self-training with Noisy Student improves ImageNet classification

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.939920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.939920Z digest=sha256:726f6c4f2ae15ed715b468bf406c8241a2bbe3a3f618c44e9f8d7e6eb1aa8281

Observation db26dd7a-caae-4063-b2eb-074cbbf4155f · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.945421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.945421Z digest=sha256:fb09e8bae8b428b676fa180091e306a81d4c6b769420a31fbf6d7766fd8d218d

Observation c3a74c6b-e452-4117-890a-58820d90bbdf · outbound

This paper cites Qwen3 Technical Report.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Qwen3 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.949492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.949492Z digest=sha256:b3bdd9f37595ecbe444495f8e8d59178a2a707303312e443311fef93792aad35

Observation cb06f9fc-655a-4d51-8dd5-4a8d5a17b7ac · outbound

This paper cites Self-Rewarding Language Models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Self-Rewarding Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.953612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.953612Z digest=sha256:c059944c1dcbf4feddce4aec1f062dc3495b62edb276d13b02dcb49fa67bb632

Observation 9294ac53-e399-4e84-abc6-f6ad3434aec7 · outbound

This paper cites Learning to Reason without External Rewards.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Learning to Reason without External Rewards

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.961404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.961404Z digest=sha256:bb1e92ec7efbd1d0ea64f326b0060865d2f4cbd279c95bc10ba11241d1b25309

Observation 5ababe61-88b4-4790-87a4-14418dfe98ec · outbound

This paper cites B Use of Large Language Models Large language models were used only to improve grammar and clarity in author-written text.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning B Use of Large Language Models Large language models were used only to improve grammar and clarity in author-written text

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:41.871053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:31:40.965726Z digest=sha256:5e2f6bc6f84306acf9374825e6f4df0fc111d595fa7d2f330ef69040279e41ff

Observation 3b0560b6-dc2c-4861-89c7-b1767232d739 · outbound

This paper cites role": "user.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning role": "user

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:31:41.848806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:31:40.971330Z digest=sha256:9cd29f655f1690d5221499333ae0c54bc7ffb61062c1da3ae821a4693fcf0f48

Observation b1e51fd3-abc4-48c4-bc3e-34771d4744c9 · outbound

This paper cites Training language models to follow instructions with human feedback.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Training language models to follow instructions with human feedback

Reference 1965

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.870949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.870949Z digest=sha256:0aae62500b59f06c4fd438cda7a3e6d544b32d6238d96f7d84eca8cfe03b9d84

Observation 439abf01-d528-4a3e-b250-662f26ec332d · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.915743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.915743Z digest=sha256:a5e273529f6cb15770455522f5d2da505f34a5b88559de148c361fbf6da697f3

Observation b567eace-03b0-4e05-a22e-705ae91c44b9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 1988

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.890115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.890115Z digest=sha256:429f84b2bd35d69a52c1ad4c408d0212f4c7faabaab6b354501bb9209976434a

Observation d5224e1c-488d-4bd6-a3a4-87d6c2f7adf2 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.794957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.794957Z digest=sha256:b857d03d38e8ce6089586068a7c66db9658fd714adcb4d8ee060b04fcc6fbf56

Observation 8155807b-da34-401a-897e-8159a418b84f · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Maximizing Confidence Alone Improves Reasoning

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.880140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.880140Z digest=sha256:911087de6a707bfda7c91d8b97bfd44655df4195553c8ce16ae67932f80ea833

Observation 2b76d1b3-2956-4e92-b2d1-d42c60fc789a · outbound

This paper cites Can large reasoning models self-train?arXiv preprint arXiv:2505.21444,.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Can large reasoning models self-train?arXiv preprint arXiv:2505.21444,

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.895438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.895438Z digest=sha256:39bea2fce7d8398aef3087d31fea49bd834101306097e81bd976122978fc6410

Observation a8dc9f2c-f800-4540-8084-e3f16997dd95 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.934963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.934963Z digest=sha256:cb7f1b2a546d184a7436caeef2d6ad53d89690dc5d82255cb6af663d6f93f9eb

Observation ceb19587-0ff7-4c21-a5df-33c004284db0 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.806378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.806378Z digest=sha256:2c62930cbe1a977c0e6ca3eae7f85c94395e9fa4d0ed370a51c9b3040bec40b1

Observation 21b7eeaa-3378-4cf0-9f30-494b318e321e · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.843022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.843022Z digest=sha256:82cd6a75c2f223d7d2d54d2ed6bb60bf0204f19c5d39db001db547df30228877

Observation 2c133695-b7c3-40d1-b1fc-7bb7dea348f8 · outbound

This paper cites The Llama 3 Herd of Models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning The Llama 3 Herd of Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.814245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.814245Z digest=sha256:e6e344183b5a5a5aaf2b013a072d294ffa3a577a84447e561374277ef82d6d80

Observation f64ae17d-0407-4c4c-917c-e2cc6e5b07d0 · outbound

This paper cites Concrete Problems in AI Safety.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Concrete Problems in AI Safety

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.788731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.788731Z digest=sha256:4b86b34a92aa13973157ae65bfb2527a5568e8e2c81d72ce65dac7a525df96b0

Observation 2169df09-6c62-4dc3-81d5-3b3baee10156 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.824204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.824204Z digest=sha256:eb97ef4e14996bb0d91bee7ab02deebd650884e623f52a99cb8f073494adfc69

Observation b466cfc6-3bbf-4847-b6d6-ea2c260e53cd · outbound

This paper cites Co-rewarding: Stable self-supervised rl for eliciting reasoning in large language models.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Co-rewarding: Stable self-supervised rl for eliciting reasoning in large language models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.957387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.957387Z digest=sha256:e412402a5e5f35a4d26317252fa9a26684b42b0aefc989ffdd99b114586c01cf

Pith citing papers

No inbound Pith citation observations are available.