Pith. sign in

Paper Citation Record · LEDGER

ReFT: Reasoning with Reinforced Fine-Tuning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2401.08967.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.08967 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:25:13.011848Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2b60958c-1ddf-4199-a784-b73d877bbe4e · inbound

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs cites this paper.

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs ReFT: Reasoning with Reinforced Fine-Tuning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:36:50.152506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T12:36:50.060335Z digest=sha256:9cd81e94342a6dae922a8b3f968c66835dd5518573487cff4c0c42e6a21b1aee

Observation 1c805dae-c164-4439-9298-0f08e61d317c · inbound

Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance cites this paper.

Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance ReFT: Reasoning with Reinforced Fine-Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T10:25:13.011848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:25:13.011848Z digest=sha256:b2d323f80f23325d27630889179310618caca839943e005f0aa27139e53f7e03

Observation d3565846-792b-4bbf-8a49-6bc802995dec · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization ReFT: Reasoning with Reinforced Fine-Tuning

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T15:04:22.859163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:0616912a8086d8144b830c434ac59a2a071bfa02a0aa1a61f325253597838e1a

Observation b43b2d88-3886-4512-aada-45ba6ff87023 · inbound

A Survey of Scaling in Large Language Model Reasoning cites this paper.

A Survey of Scaling in Large Language Model Reasoning ReFT: Reasoning with Reinforced Fine-Tuning

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:22:09.255716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T21:20:07.238992Z digest=sha256:ac26fbad1a16948792136a5d31a9f334ce958b51210e53f81b6d07e46b3c89fe

Observation 8c0a4906-8f08-492a-a834-58cccecf0722 · inbound

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs cites this paper.

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs ReFT: Reasoning with Reinforced Fine-Tuning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:42:39.072862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T18:42:39.023650Z digest=sha256:0377f240941f19a2b96e5e07b58e481f13d9bab092e2eb1ce96eddf152526e09

Observation f422ba67-dcdb-42d1-abdb-77fbde5d2ca9 · inbound

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning cites this paper.

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning ReFT: Reasoning with Reinforced Fine-Tuning

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T15:58:33.423958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T15:58:33.219451Z digest=sha256:87ac29db9ffd8dd749e50b1eb14ee358395aa589e0a219abdb773ec0c793ff63

Observation b77e8eb1-cffd-4a6e-8c4c-dc90b07f2190 · inbound

Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence cites this paper.

Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence ReFT: Reasoning with Reinforced Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:07.369948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:07.369948Z digest=sha256:e53398dc461ea55330795a0ff982869ca3a452d0e4a6adab70c01c12f9481c1a

Observation d478c978-2271-4eef-acb5-47450fea440c · inbound

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making cites this paper.

ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making ReFT: Reasoning with Reinforced Fine-Tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:59.379335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:59.379335Z digest=sha256:e88243bf7d66025ccd961a7a541c8a259544777c0bc7e10f81cf551bf9230d93

Observation 77f82f1d-90bb-424d-ad69-2b7662b4c4cf · inbound

Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning cites this paper.

Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning ReFT: Reasoning with Reinforced Fine-Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:09.774303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:09.774303Z digest=sha256:b36f18143af99543c46bc3c9f920967c8948e1aebf9ba06e5395870337e61ec5

Observation 59a5836c-04db-4516-af7a-2d9bb671b982 · inbound

Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise cites this paper.

Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise ReFT: Reasoning with Reinforced Fine-Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:21.337851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:13:21.337851Z digest=sha256:dc4846d6b69834ae76df2ebf4a005f58f67cc607a2e787168caf1e4d6b723110

Observation b2cb4f65-b5b7-41c3-ab44-f1004e6de99a · inbound

Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning cites this paper.

Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning ReFT: Reasoning with Reinforced Fine-Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:07.790412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:07.790412Z digest=sha256:66cc015b09298899e3e16a79fbe6667547f31e852114378e689974d15521fa3e

Observation 3f1138a9-c49d-4495-b9d6-6c8355ea4e03 · inbound

AI Agent Behavioral Science cites this paper.

AI Agent Behavioral Science ReFT: Reasoning with Reinforced Fine-Tuning

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T11:00:53.750838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:00:53.750838Z digest=sha256:e32fee7f2bd6d9be66c8892786e6bdcc348a052b26597fa1cdca0619a402cecd

Observation bc871d76-c04e-4477-881b-39ea09382062 · inbound

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models cites this paper.

Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models ReFT: Reasoning with Reinforced Fine-Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.701431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.701431Z digest=sha256:fe73870add65f6a3847fc51dd2b6aa7bd2601923a535056d47ea6fbe43853444

Observation b1ee18de-6c4a-4f6f-8ec9-ad19fbc7fa63 · inbound

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning cites this paper.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning ReFT: Reasoning with Reinforced Fine-Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.171772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.171772Z digest=sha256:b1f7c905733b34e1299ea36c479998d0c7236432adcff5f75e62d6542f82c735

Observation 28faa975-416c-4b4a-abf9-c8e60afd8a80 · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset ReFT: Reasoning with Reinforced Fine-Tuning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.887684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.887684Z digest=sha256:22738306323cf1ce9692f725a4208224a391185b6d950f40ed939c5626a6349c

Observation 216caaee-f178-4204-b854-b162cec63e98 · inbound

Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training cites this paper.

Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training ReFT: Reasoning with Reinforced Fine-Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:31:27.373655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:31:27.373655Z digest=sha256:aa86ef3abbdf8e81a86b1f02b9a8d36b2356a55f486507ad707bc9a501ea85da

Observation f67dfa6e-fe8d-43cf-80b9-d1ddee10220c · inbound

One Token to Fool LLM-as-a-Judge cites this paper.

One Token to Fool LLM-as-a-Judge ReFT: Reasoning with Reinforced Fine-Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:37.905420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:15:37.905420Z digest=sha256:0cfd29e36667546c17f07769a9279d7b67c89fac5ce94618a39e322f70bc0d4e

Observation 60eaf0db-3db8-40e8-92d3-e5fe0f7f32dc · inbound

Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding cites this paper.

Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding ReFT: Reasoning with Reinforced Fine-Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:50:51.182259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:50:51.182259Z digest=sha256:9b3576cd95e1cac61c7a54646318beb037dea8eb317667ae7805d1803608bcb7

Observation 85fb2d14-dd27-47b3-8cea-2c384eabee1f · inbound

Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice cites this paper.

Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice ReFT: Reasoning with Reinforced Fine-Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:52:43.034343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:52:43.034343Z digest=sha256:93e4ff407ec7637a6a26e7cf6ba322f69f29e3a1367900679a3d31780b01006c

Observation 03bae2fa-c3cb-4423-8c12-78049f450a13 · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding ReFT: Reasoning with Reinforced Fine-Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.819153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.819153Z digest=sha256:77ff8329f0d9a1de5cb690b71e17df3b36e35a19ff4fd572596358d38ac06149

Observation 8318f854-3f94-46a9-8799-50700ba725a1 · inbound

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards cites this paper.

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards ReFT: Reasoning with Reinforced Fine-Tuning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:26:28.273192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T14:24:48.666197Z digest=sha256:a3f818776223fb03aae74d925f92f38e6127b8cc3f021dbf4f5b5093c1d1ba70

Observation 011d2cef-7b5a-473a-aeec-aab241eaa0a9 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning ReFT: Reasoning with Reinforced Fine-Tuning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:25:22.566198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:d3b34249ab388325aa2d93a768d32fa623ccc9431974e95386554859829a8503

Observation 393c61a2-8e87-45aa-b99a-acdb263710d5 · inbound

CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical Reasoning cites this paper.

CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical Reasoning ReFT: Reasoning with Reinforced Fine-Tuning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:20:57.551310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T13:20:49.919833Z digest=sha256:25dccc8cf35aaeee085a1a546a369292cc3741bfecb95e728257801615856407

Observation 134200e4-b179-45bb-83fb-61d63dfb2635 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning ReFT: Reasoning with Reinforced Fine-Tuning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:11.036161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T13:13:13.293921Z digest=sha256:4d8c639cb9d62dcef7e07ca7f9f9c33870d801d0c3c5001afec1189cf0d62094

Observation 60924a86-6c78-4322-9b77-3f5f1ec98ea8 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning ReFT: Reasoning with Reinforced Fine-Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:43.836467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:43.836467Z digest=sha256:185ca8745f902aef2a42e6b8bfeb495f0abaf4f40e9a7cee6eaef90828e22b0a

Observation 3931b483-e967-4b6c-b3d4-ad9276220cd8 · inbound

Towards Sparse Video Understanding and Reasoning cites this paper.

Towards Sparse Video Understanding and Reasoning ReFT: Reasoning with Reinforced Fine-Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T23:33:12.426070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:33:12.426070Z digest=sha256:fe836e92cf8a97b5699da63932740f7c316334ccdcc0dbaf2f8fc7d8b858162d

Observation b20dc4d0-0773-412b-891f-83cb66077905 · inbound

PARM: Pipeline-Adapted Reward Model cites this paper.

PARM: Pipeline-Adapted Reward Model ReFT: Reasoning with Reinforced Fine-Tuning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:28:39.504938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:15:26.015817Z digest=sha256:f2182149705f7342d72df2ea2d8fa098d2874a730904a34ee1932998cf763415

Observation e6d25f51-1e15-4dbf-b6f7-2d39a11ab1ef · inbound

Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems cites this paper.

Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems ReFT: Reasoning with Reinforced Fine-Tuning

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:06:05.471421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T23:27:54.704794Z digest=sha256:48d68d1ecb8105b9fc88e855c8a924b0d8a50413a63a97cac294883b80d0288d

Observation 604eab8c-9026-409f-8c93-f7b546795d20 · inbound

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key cites this paper.

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key ReFT: Reasoning with Reinforced Fine-Tuning

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:21:08.248562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T09:35:47.501360Z digest=sha256:cf7287a4039b600fc8b20fc130c5576ab9a74b321ef802bd39c7114b6bb37916

Observation 39b51824-bed8-48e7-b49f-e7b227ea4e8d · inbound

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key cites this paper.

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key ReFT: Reasoning with Reinforced Fine-Tuning

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:21:19.207104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T03:16:59.195706Z digest=sha256:7ca04b755cc39ac4c728c3b148589447a78174be676547c25e0cd3a49290ec67

Observation 97a3e1c7-6394-4eab-bafe-90c433301b34 · inbound

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key cites this paper.

Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key ReFT: Reasoning with Reinforced Fine-Tuning

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:39:10.435844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T22:36:13.781114Z digest=sha256:3b3ad5623d04cfa369ecc467fe4675c97bafa26cef5e58e6b09986a55b767089

Observation a5398c4e-c9e3-44eb-9592-713551c87f45 · inbound

OracleTSC: Oracle-Informed Reward Hurdle and Uncertainty Regularization for Traffic Signal Control cites this paper.

OracleTSC: Oracle-Informed Reward Hurdle and Uncertainty Regularization for Traffic Signal Control ReFT: Reasoning with Reinforced Fine-Tuning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.957307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T00:52:42.845447Z digest=sha256:cd2bca594a5baad5eddb5b4950e87b9836ba06c6f356132e0165f44ad71245d8

Observation 9b569002-79a6-411e-82d6-f71c6a131152 · inbound

Reinforcing Multimodal Reasoning Against Visual Degradation cites this paper.

Reinforcing Multimodal Reasoning Against Visual Degradation ReFT: Reasoning with Reinforced Fine-Tuning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:06:25.307181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:37:21.451146Z digest=sha256:a8c18426075142a550cc238ee5815b054449f0aac81744ace5c109af51e90691

Observation 12f10086-267a-46f8-b0c7-659981a5e0d8 · inbound

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media cites this paper.

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media ReFT: Reasoning with Reinforced Fine-Tuning

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:08:20.469682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T14:05:14.737146Z digest=sha256:469abbd7c4f1f0e42bf865d93b4909cd775511d9cfab2a17f862d8daf135aa6e

Observation 225262dd-f8f8-4b10-9332-b4e643357c8a · inbound

Hide to Guide: Learning via Semantic Masking cites this paper.

Hide to Guide: Learning via Semantic Masking ReFT: Reasoning with Reinforced Fine-Tuning

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:34:39.110322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T12:16:12.108715Z digest=sha256:28618ce2668ba19b0ce3dca442822a0165744897481a2c56471a300ad7fc13aa

Observation 52332038-93eb-4e0f-ad74-121870f0626d · inbound

Sakana Fugu Technical Report cites this paper.

Sakana Fugu Technical Report ReFT: Reasoning with Reinforced Fine-Tuning

Reference 290

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:29:38.287103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:22:37.596720Z digest=sha256:7c6a3e4a3f42bc069c3156d8d607df3f760ff261706c8d71cf7accb0e1bd8ca3

Observation fe614fb8-4ffc-4679-b762-05bf33793149 · inbound

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training cites this paper.

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training ReFT: Reasoning with Reinforced Fine-Tuning

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:09:40.781225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:14:35.109298Z digest=sha256:f8ab2e8a648528c3b39eb6cb2e24e7d781962b8e2f1acf8ed5aaf69116eef44a

Observation 195b7c74-96db-4177-97df-00a66ee514ca · inbound

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training cites this paper.

Improving Reasoning in Vision-Language Models via Perception Verified Self-Training ReFT: Reasoning with Reinforced Fine-Tuning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:35:39.633672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:29:51.635039Z digest=sha256:fbf2f7cb768a019cedf0cdf7716b9eb077b3be42c542a9f625bb9eca1131d82b

Observation a521c81a-e063-4479-9afa-ae5901e356de · inbound

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text cites this paper.

SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text ReFT: Reasoning with Reinforced Fine-Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T13:36:56.480861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:36:56.480861Z digest=sha256:a736715d26cb98ba2255d4ce9a03c988ae9ea80ab388195b66bb56effa53ed81

Observation 44030fd9-7c7c-4891-a80e-b818b9ea07ef · inbound

LeAct: Learning to Reason from Expert Actions cites this paper.

LeAct: Learning to Reason from Expert Actions ReFT: Reasoning with Reinforced Fine-Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T06:32:20.959706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:32:20.959706Z digest=sha256:1f63a1df88ac4d60e89897f29607584f406b39e6900cd720df3629d90233b332

Observation dfcfd0d3-334c-4196-9e74-1351f190d653 · inbound

MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation cites this paper.

MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation ReFT: Reasoning with Reinforced Fine-Tuning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T09:20:19.366216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:20:19.366216Z digest=sha256:21beecf80afa9936fd410a6b8257dc4081c7999149af6529aad57fdc37bceec6

Observation 9f2b05e2-0da4-4382-99b5-a5ede3c9c0dd · inbound

Self-Improving Large Language Models via Progressive Experience Evolution cites this paper.

Self-Improving Large Language Models via Progressive Experience Evolution ReFT: Reasoning with Reinforced Fine-Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:11.941692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:11.941692Z digest=sha256:25145f9d56680a9ef990dc0e98ab0a64ad4e7b9803411e027acb30fd10782361

Observation fd8b029b-c614-420d-963a-c6901beb1782 · inbound

Self-Improving Large Language Models via Progressive Experience Evolution cites this paper.

Self-Improving Large Language Models via Progressive Experience Evolution ReFT: Reasoning with Reinforced Fine-Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.493528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.493528Z digest=sha256:92d7ef9d29c3c16733bf7e0d8418eecf82f9c5d70bee5984f612d1e768a531b4