Pith. sign in

Paper Citation Record · LEDGER

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs

As of 17 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2507.02851.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02851 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:24:00.072659Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1ab79bf-bfa9-4f86-8244-f623cc80202b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.425741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.425741Z digest=sha256:88d71c83e6094089feb78e3e64b61d0c2c0343bef3b25b1b067f6b0a319024ad

Observation 4bbb6bad-8780-49ae-8cff-76d35cbd5203 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.618190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.618190Z digest=sha256:b424c37f694fb30736b4c552d7b8930909815e57a87a284744e8068861c9c3ba

Observation 10d88f10-faae-461e-92be-b5b33ae3038e · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.701438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.701438Z digest=sha256:c64d8ba0279f533982f88313e754d2e9afd6fb67682261f85cdb053a695a6f27

Observation eeac5b9b-b6d6-4df1-8850-e86650132768 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Lost in the Middle: How Language Models Use Long Contexts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.003617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.003617Z digest=sha256:ec8bc4352c29780026bd5a462f1ac4e4b1f084b883494a63bdb28436ac6f888a

Observation 25bb37db-d3ba-485c-b145-102e33531c4c · outbound

This paper cites Distributed Mixture-of-Agents for Edge Inference with Large Language Models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Distributed Mixture-of-Agents for Edge Inference with Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.097832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.097832Z digest=sha256:1e2839bfb7689d3d1f4fd360f3fbb73e3d6be3ae5908a1ae09179bdda8900ee8

Observation 13df14a6-0ba3-4bfd-9742-b2bafe2e196f · outbound

This paper cites s1: Simple test-time scaling.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs s1: Simple test-time scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.192668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.192668Z digest=sha256:6511e24dc97f2a10d6820a4764313ef4438b2feae16c3ae2f9a50a5e20052234

Observation 5e128e2a-9d15-41db-8ed4-447e90630a8e · outbound

This paper cites Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.260212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.260212Z digest=sha256:3932d240ebf677ef7a7b195d9e6244b607a094c12e7468ae0a2dccccd062297d

Observation 54dd50c3-df24-42f5-b8f8-ec4744002cf9 · outbound

This paper cites Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.391394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.391394Z digest=sha256:b5c3faf3decd82fc81bbc42fd50250a8bc875fbd66de84b9506cf42381e7a667

Observation 6f21235f-aaf2-4b8b-93d7-aa0cc0404c4c · outbound

This paper cites Iteration of Thought: Leveraging Inner Dialogue for Autonomous Large Language Model Reasoning.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Iteration of Thought: Leveraging Inner Dialogue for Autonomous Large Language Model Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.459779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.459779Z digest=sha256:2eb89c08d2559f943b38b5ba9a115ec693f6d0bbe8543ab7a9942e9af3dcc555

Observation 5624c9c5-cef8-4112-969c-fec3d4d8411b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.546311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.546311Z digest=sha256:74ee0c30475bd9e2945fb28b2c0cc6d288ae5330c016778fe0f1f4d946581989

Observation 4acbca54-faff-4b52-9264-555fd99b6fe8 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.612071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.612071Z digest=sha256:ef06c06f416ee9086d6a137864ee26e817586dc91cc76344162e295be912b844

Observation ec8f81c2-c71e-48f4-8be7-a5f134b74053 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.704363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.704363Z digest=sha256:cb0bf67e2cf136a6f3e80dc298d14a12c8546768efd423b944855cdf3de1ce32

Observation 4c107b09-0a26-4ad1-b454-e6d62fdca359 · outbound

This paper cites Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.774601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.774601Z digest=sha256:4c9f2f2b30dab763720e5ea8f0d7ce42cfadca0fc46301a3353b3588c3277156

Observation b5b6da0e-5c72-4e63-9fd5-00f508c0bc21 · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.868672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.868672Z digest=sha256:8e14ffd327168ecb029afc6e90ecaf347dcd0093afb132ff20b451ebd1cd6142

Observation 3e115bb4-420d-4c02-b88a-a10199752927 · outbound

This paper cites Inftythink: Breaking the length limits of long-context reasoning in large language models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Inftythink: Breaking the length limits of long-context reasoning in large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.964450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.964450Z digest=sha256:3c139d873724a4f5840dc912385b88a5abaa71bbaf5385e7a889a631e23f5738

Observation 19b1f7ff-1647-42a6-8aef-00e65e33485f · outbound

This paper cites Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:00.072659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:00.072659Z digest=sha256:3e8eb71f8308e97189ef7bf38f7cc507e8e08c4449285748002e2674050d0b91

Observation d697b007-2cf2-4794-be4f-2283440f5ce2 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.328473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.328473Z digest=sha256:4392718f316f7cf05a17a50bb05271785dbee279121eafbb7759cb261d56cc2b

Observation 1673fa13-dfda-421b-8440-b4887d3289ba · outbound

This paper cites OpenAI o1 System Card.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs OpenAI o1 System Card

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.821730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.821730Z digest=sha256:fc8826b457caa4932cb777682f7c4cc2a0b25ccb163063d2237fa4c416c381a8

Observation bd220da0-e9ad-43e9-b6fd-e1996fb0e4de · outbound

This paper cites Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.523176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.523176Z digest=sha256:7c4e9473a5393ea1a7e04a9a0583838f44d5d9c3aa283cf3b30f6a11021d1ce1

Observation f67615ef-5366-4947-a05b-f5a7327e10e0 · outbound

This paper cites SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.917129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.917129Z digest=sha256:e631bd1fa7e304f67792374a68bb5cb9bec578bba0803baeb518dc12217ab76d

Observation 127a6973-4793-4e5f-bb46-1190510a15f9 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Evaluating Large Language Models Trained on Code

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.197865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.197865Z digest=sha256:ae57ef9eed0fede91943161b47b3ab8a0c8426f80b7bda4cf52d8ab0bee9e492

Pith citing papers

No inbound Pith citation observations are available.