Pith. sign in

Paper Citation Record · LEDGER

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts

As of 20 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2510.09278.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.09278 v2

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:40:50.239122Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 361a8eb8-0174-4642-b389-c59f74f9f910 · outbound

This paper cites an unresolved cited work.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:42.323020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:42.323020Z digest=sha256:d73cf7f7c4691feb8518aa1bcc141e3f66f3bf98e26e66223ba816a63f9991c3

Observation 77f1045f-9409-4096-9730-639a51ae7717 · outbound

This paper cites Chain-of-Thought Reasoning In The Wild Is Not Always Faithful.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:42.437398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:42.437398Z digest=sha256:a3d416ec17779a5f5ed44cbb84dd34b2a35ca24c9642ba502cdd04cbb75ed2b8

Observation afba2411-c0db-4330-bd1c-911667d6a48b · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:42.596699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:42.596699Z digest=sha256:2c705fb2e2b696a25b0a37ca19a17fa8d04582dcad31a5e9174f7522fa4ecbca

Observation 82004794-885e-4aa5-a803-aaf580e011e2 · outbound

This paper cites HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:42.851428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:42.851428Z digest=sha256:6243c0400d99d7c7386cfea33aeaf2417eaaa8472cd2e0b323e61fa6ce8f0a43

Observation 399715de-4411-4b6e-a0b8-7346bfde06b6 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Reasoning Models Don't Always Say What They Think

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:42.937652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:42.937652Z digest=sha256:d12aa9a0c0df15e6472ee7d50defcb65ac51c07852756e2ce34a8ceb7cfbe1fe

Observation 2ff1e785-3ad6-44f2-a00f-3418d956fcfa · outbound

This paper cites Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:43.075961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:43.075961Z digest=sha256:4a8c3e98c4e5b214f6f193353a14bafd519cf0a7dd7ad7cddc447843cc4414f5

Observation 7bba21fd-6a2b-42b2-9553-09838920feee · outbound

This paper cites Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:43.208031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:43.208031Z digest=sha256:6a085cf9f9f4287d9ad643855ad2f441644b7bf6eb03e80202f24811a877af67

Observation 083f581b-80e7-4a24-93a0-99e2548948f7 · outbound

This paper cites DeepSeek-V3 Technical Report.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts DeepSeek-V3 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:43.382825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:43.382825Z digest=sha256:f0705a0661d0d4ec70d5990ece7f0f6721ebe6f91db6a58aa4db1e62c6865298

Observation 19582271-b0b6-431f-8b71-61e74e59e90f · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:43.561051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:43.561051Z digest=sha256:4d0678f733ddf8701b41e08e3f497565234e6adeff9dd16fcf81730fdfe9cc5f

Observation 62ef4912-82a5-48f2-89a5-b2c5682cfa19 · outbound

This paper cites ReCode: Reinforcing Code Generation with Reasoning-Process Rewards.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts ReCode: Reinforcing Code Generation with Reasoning-Process Rewards

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:43.728971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:43.728971Z digest=sha256:5ee8c256fd202913473486375fc2becc24a883826b0437cc4cd8dba42a956d04

Observation d1dc259b-d0ea-4212-8337-94bb634b95ba · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:43.915453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:43.915453Z digest=sha256:5ed66e769d0a5933322472cf9a74815ab971aad8e515175233436d0d85d27a37

Observation 16b77f73-248e-4669-95fe-35d2d5d1f121 · outbound

This paper cites Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:44.071135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:44.071135Z digest=sha256:1d1ff92a5f58956d0c1dc42975a5eca3fe78f9cf2fee7463646af40b4e109e25

Observation f66c8e89-2702-4cfc-9bc1-c620cbc15149 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:44.210807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:44.210807Z digest=sha256:f6abb0dc876909dffeba4ecd93401da7492b06add323eee1c28336bc4d246754

Observation 99c30de1-b493-4faf-b77e-26080463a2a8 · outbound

This paper cites A Survey on LLM-as-a-Judge.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts A Survey on LLM-as-a-Judge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:44.385582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:44.385582Z digest=sha256:2947d9f128c97bd58427b39d4bc0aff7067dc3cd59a379b0816455783bb16ed1

Observation 6b543e0d-9588-4705-b2ba-d72e6409fbe9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:44.588687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:44.588687Z digest=sha256:7c122ddb024745fec47cad6d9f9cf2784db9a88fd9e0b67fc88d536cc4ff8410

Observation f742fe26-bddd-4058-9644-5250871c6264 · outbound

This paper cites an unresolved cited work.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:44.717777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:44.717777Z digest=sha256:496b444133f410930422d1ffcdaed6d9bc90b1cddda7a95e25a91d837d38286a

Observation f1610e9e-0199-4a63-8962-402cc25e0f91 · outbound

This paper cites an unresolved cited work.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:44.860700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:44.860700Z digest=sha256:ccc801f5464d15e71af420e1b371df4f879d3acec390e9aedc022f63fee6c05d

Observation 2f131002-14e0-4128-b7f8-fbd9143e9613 · outbound

This paper cites RAG-RL: Advancing Retrieval-Augmented Generation via RL and Curriculum Learning.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts RAG-RL: Advancing Retrieval-Augmented Generation via RL and Curriculum Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:45.013735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:45.013735Z digest=sha256:91935ea6035aede262c18691011ed6385bd2f9baf5a8d270325108619050391f

Observation cc19292c-a333-448f-a6d9-6b92b54af040 · outbound

This paper cites an unresolved cited work.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:45.142857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:45.142857Z digest=sha256:ad004c57645b6d26d9d8072c76357bfd172e9f75925ad563da9fa9d318912628

Observation 633c5962-5bbe-4d59-af80-61a5e622d631 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:45.271923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:45.271923Z digest=sha256:e4f8e2bd736a9d848b50868b8caf18d9dbcd511516fa82e7ed7abff411400e63

Observation 65080f60-5aa7-457e-b13f-db4814363f61 · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:45.421707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:45.421707Z digest=sha256:d356251ec906e475356f899638c3fa37d907702cf3f56553efc205ee47beec7e

Observation 22b37d82-c8b9-4cda-9e10-2b91e8a3397e · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:45.563609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:45.563609Z digest=sha256:8fd4030cc7d736cc32a63ea0e94628598386e990a0e4b215e3c882bd2b7df0d1

Observation c63b186b-05fb-47b6-a7b8-365e099d1d7e · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:45.717755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:45.717755Z digest=sha256:26392435b8c920513ded2e449e2305700e920621f4998c316cb461462ec4808c

Observation dd420bf7-a3e9-4cb5-bfca-956a24b62bcf · outbound

This paper cites Prover-Verifier Games improve legibility of LLM outputs.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Prover-Verifier Games improve legibility of LLM outputs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:45.855435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:45.855435Z digest=sha256:a55e845b791f7f1ade10374908bdc971128a00581392d6cae955b7ff9f0f2514

Observation 84786868-3129-4d84-9da9-749e7413b8ec · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Gonzalez, Hao Zhang, and Ion Stoica

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:45.997496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:45.997496Z digest=sha256:2b6437c4a1946cc46411c08ea64cb0036a6efa58c5189b26abf56028c09601e2

Observation c6a6d043-c3c8-4645-8703-761607e6d954 · outbound

This paper cites Reasoning-Table: Exploring Reinforcement Learning for Table Reasoning.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Reasoning-Table: Exploring Reinforcement Learning for Table Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:46.135824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:46.135824Z digest=sha256:795c1bc70ff4899fb4dbb6b16ea179c6f074e2c9298a89d5c18b1128d14b9283

Observation 7948169f-0219-41a3-b275-fb64d3c67ad5 · outbound

This paper cites an unresolved cited work.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:46.264642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:46.264642Z digest=sha256:fc4be89c4ab6c2aa19acc3939d9a1cac40a90bd0b55cdd7c09beee4bd1c414a6

Observation 3560fc99-bac1-426f-a0dc-af3069e0e8f4 · outbound

This paper cites LIMR: Less is More for RL Scaling.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts LIMR: Less is More for RL Scaling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:46.397811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:46.397811Z digest=sha256:d6c295f9095297ede6478f5e628cc6455b5f0c1ca27af31fc59a69165541e75e

Observation 3ef278a6-c580-4aa6-8739-30e00e1c6b21 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:46.491019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:46.491019Z digest=sha256:28aa427fddf02e94905fe01e8d761da3727230004f9fb974295ea95e1f66a7a9

Observation bbffb5cc-a76d-4be1-ab53-2bab677d2646 · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Frontier Models are Capable of In-context Scheming

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:46.608654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:46.608654Z digest=sha256:29c514135d776192814b132a1c8bd24eceb4aaf5ade3bcac06ee6da0e60a460a

Observation bc1b9482-67ca-459a-a80c-163e8712ba4f · outbound

This paper cites GPT-4o System Card.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts GPT-4o System Card

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:46.751263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:46.751263Z digest=sha256:e7f0e3ea76dea2d5a9db56b548ffacef3b52bf308372f9b5938be15e0d645cdc

Observation 9b628e38-ffc3-41ed-882a-cb25f9cd7aed · outbound

This paper cites OpenAI o1 System Card.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts OpenAI o1 System Card

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:46.948269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:46.948269Z digest=sha256:36757f69a31eecf421a6baea7118566de15b161de783bb0b5fe5d12c66e5005d

Observation 26cea3b5-6e5c-4162-9666-895705db7df8 · outbound

This paper cites an unresolved cited work.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:47.107142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:47.107142Z digest=sha256:a74f34de26205752deb6545113d46333d4b2a0e14dcf9c945db301f86ab4ebfc

Observation 51da2358-127a-4fff-a8c1-6b007bace80c · outbound

This paper cites an unresolved cited work.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:47.234256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:47.234256Z digest=sha256:20f858fb44f57f815efef2782b7513df1e075b9eaa6a968e2ab95d158ccb82fc

Observation a003e3a2-a1a3-4431-917c-c05e411e3e94 · outbound

This paper cites Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:47.386849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:47.386849Z digest=sha256:f16ae2f6b5722be9e0170714a97ab9a70a4da33e0b5071d96f4c3b4f2ed64f98

Observation a26a8150-b357-4ef7-9f27-7a9b98c5a117 · outbound

This paper cites an unresolved cited work.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:47.471242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:47.471242Z digest=sha256:9436f06a81955b6b246380cce92115ce2c33b87b54ab26cdf2d51588e410a29d

Observation e32dbfc1-15da-42af-b8bf-71eda8124d6c · outbound

This paper cites Qwen2.5 Technical Report.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Qwen2.5 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:47.632256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:47.632256Z digest=sha256:4f66e2a4ecef787596fa6e78ed0cb54b9328d5cc2b654fba1758aa34ce5fc0cf

Observation 436f5eb2-4f65-4e35-ba16-5a7cf8f6c83b · outbound

This paper cites Proximal Policy Optimization Algorithms.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Proximal Policy Optimization Algorithms

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:47.812288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:47.812288Z digest=sha256:37e2fc9f64191851dbf667b6fd8d74d794db5df9caa04c164a3dda92e02b19f1

Observation 6bb43ed4-80ab-46b8-97aa-4f70436b58aa · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:47.950018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:47.950018Z digest=sha256:ca41ec0f94bda233f4bd7366479730f5021fbe2ab61b1b73341d6b3bf974b072

Observation 299ae545-5864-445e-9222-3709189eb5a1 · outbound

This paper cites an unresolved cited work.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:48.131252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:48.131252Z digest=sha256:e4263dd06ae391426a54aef9d48464c5ed6cff246c4e36607fd09989909d9109

Observation 54e137fd-72b3-4241-95c9-53f001ea1304 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:48.311730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:48.311730Z digest=sha256:533e07dcc0aa09ff560696260e98395f2c8dc390ecbddf420b556f5e7599f0d8

Observation 5678ad83-14b1-4b82-b552-9a4d16097830 · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:48.485814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:48.485814Z digest=sha256:a42980aee2cb05b4c79c88ca2582004b20ce9c0d499b577a8bd8756f15ee5d5d

Observation 580f43ad-e190-4243-a5c5-cc26060f8c41 · outbound

This paper cites CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:48.632259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:48.632259Z digest=sha256:f839cca7333d180c0ba328cdc61ec22fb475f0c65897f40141d3dc3e734a9e08

Observation 36edb7ea-7158-4493-82aa-accde33d1d79 · outbound

This paper cites Chain-of-Probe: Examining the Necessity and Accuracy of CoT Step-by-Step.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Chain-of-Probe: Examining the Necessity and Accuracy of CoT Step-by-Step

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:48.792795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:48.792795Z digest=sha256:6032acb0d46e9b7a2fe73a5b63cb026f9f26d97212d258cc2432141e71299b38

Observation 9b6e4891-99c2-4707-8f6c-8c110fcd1ba1 · outbound

This paper cites an unresolved cited work.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:48.992415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:48.992415Z digest=sha256:ff4a7a28e8ba27caa9c76933a048b7562ed818813047cca9f0963d2bc74b96d9

Observation 91b278bf-e186-4ec8-855e-1a2b944eafe5 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:49.178218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:49.178218Z digest=sha256:a4b9069333735c4e87714558ea6b7aebea55943bb5ac520d68965976369f7b09

Observation c6a232ff-50c6-4fae-bf07-a1a1895c37c8 · outbound

This paper cites ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:49.340924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:49.340924Z digest=sha256:5e1f0f267fc730e98b7dd3a60747f33b2a2cd977c0da06a03f8f22c176bfabbe

Observation 43b5f8b9-8e61-4b7c-8c6c-198a7f4189aa · outbound

This paper cites FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:49.493709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:49.493709Z digest=sha256:81a2e862d344437094f2f91572c8f92d6673ba009bb04808f8a21bc1c1ed2122

Observation a78efc93-19c0-4077-8f49-858bcd059b8e · outbound

This paper cites From Graph Retrieval to Schema Realization: Counterfactual Validation for Text-to-SPARQL over Heterogeneous Knowledge Graphs.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts From Graph Retrieval to Schema Realization: Counterfactual Validation for Text-to-SPARQL over Heterogeneous Knowledge Graphs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:49.666102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:49.666102Z digest=sha256:2a7374965f1989ac61f3b277b37412d1e5584ed495a2d07bbfec6be6aa2323ca

Observation 297cb829-c30f-433d-a7ca-7164f4b980ad · outbound

This paper cites an unresolved cited work.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:49.793271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:49.793271Z digest=sha256:8c617e5ad9f327f933f2ce4c6ee81f8c174a5d0a6ee474287d30f7fc3903a58d

Observation 7ce2f5b1-d381-4916-9404-e26dce96ef3f · outbound

This paper cites Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:49.924763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:49.924763Z digest=sha256:1792f6dda28687d1ba5e6bc6993f048ed886d9cbc8e4808d18a380d1b7e818cb

Observation 248c655b-998c-4f70-b69e-13c16eb45eb3 · outbound

This paper cites online" 'onlinestring :=.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts online" 'onlinestring :=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:50.098262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:50.098262Z digest=sha256:cc003cf486ffb9958743ce3486db0602c30afd22271160d1ff14dce1f2cddb09

Observation dd20c3ad-e6b4-43f5-acf4-8e89e68fccb7 · outbound

This paper cites write newline.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts write newline

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:50.239122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:50.239122Z digest=sha256:82ac70b5e30dd30b645d537fbc4f69170492af14516ca5d36b4ed7e810e3ac7a

Pith citing papers

No inbound Pith citation observations are available.