Pith. sign in

Paper Citation Record · LEDGER

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs

As of 10 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2507.02851.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02851 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:24:00.072659Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1ab79bf-bfa9-4f86-8244-f623cc80202b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.425741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.425741Z digest=sha256:ddfca1000a141169210f3877bc22b7ab6424377050bc12510e97adb4cdf8e1a1

Observation 4bbb6bad-8780-49ae-8cff-76d35cbd5203 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.618190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.618190Z digest=sha256:3f7576b2fcda7f27dfa8f6e43ca560b3d842e62de38f731ebe8a635dbdd36ef3

Observation 10d88f10-faae-461e-92be-b5b33ae3038e · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.701438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.701438Z digest=sha256:d90a127a4adbcad447337bf9cc323beb9d53ffb7c912ee6495172d117fd7da53

Observation eeac5b9b-b6d6-4df1-8850-e86650132768 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Lost in the Middle: How Language Models Use Long Contexts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.003617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.003617Z digest=sha256:fd1f66b7a983335f3e7106200876f97f4d2d6a74050fbdbc26a8a968ee4e1732

Observation 25bb37db-d3ba-485c-b145-102e33531c4c · outbound

This paper cites Distributed Mixture-of-Agents for Edge Inference with Large Language Models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Distributed Mixture-of-Agents for Edge Inference with Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.097832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.097832Z digest=sha256:189b8f5c1f9b700871bda8a114c2bdeba9cefb676e6e43487d67a00a3b126473

Observation 13df14a6-0ba3-4bfd-9742-b2bafe2e196f · outbound

This paper cites s1: Simple test-time scaling.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs s1: Simple test-time scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.192668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.192668Z digest=sha256:20054241b0e6298e70a544c7f55368003adb3a454acf9788469ed49a680d8a29

Observation 5e128e2a-9d15-41db-8ed4-447e90630a8e · outbound

This paper cites Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.260212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.260212Z digest=sha256:e2d56fd406f834152c31a3861fe20982722a8cb239555873ce4b6a7aa6956633

Observation 54dd50c3-df24-42f5-b8f8-ec4744002cf9 · outbound

This paper cites Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.391394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.391394Z digest=sha256:f2c7b1d9bce70c1d418bd315b428ea6f278c58cac2c945e907319bb072a5058d

Observation 6f21235f-aaf2-4b8b-93d7-aa0cc0404c4c · outbound

This paper cites Iteration of Thought: Leveraging Inner Dialogue for Autonomous Large Language Model Reasoning.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Iteration of Thought: Leveraging Inner Dialogue for Autonomous Large Language Model Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.459779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.459779Z digest=sha256:7c2149d7f36d426b9e6a165d1c509974af026c9192b36aa7277e4d9949a32fb2

Observation 5624c9c5-cef8-4112-969c-fec3d4d8411b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.546311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.546311Z digest=sha256:856097c929d58dbac05e225f2e0ce72b0f5773d625849ba672e03d3f511cd348

Observation 4acbca54-faff-4b52-9264-555fd99b6fe8 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.612071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.612071Z digest=sha256:fbb7c14c7ab120c10d3becd7f4583b087a404cea4e75f5c87a6bed3b7326ad53

Observation ec8f81c2-c71e-48f4-8be7-a5f134b74053 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.704363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.704363Z digest=sha256:2f0b7078b71eb14ff00e2ec1c7af0b8752d7aa97e823119f7c36cf0c4e99a2c3

Observation 4c107b09-0a26-4ad1-b454-e6d62fdca359 · outbound

This paper cites Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.774601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.774601Z digest=sha256:a35c68ff73888e51af61b016f019040df0f9dfff05cbccba62be0dc5992217e1

Observation b5b6da0e-5c72-4e63-9fd5-00f508c0bc21 · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.868672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.868672Z digest=sha256:c9b0f80b5f0a584008caa2b3c635683475c315cd3ac1f77b99286b9bcb8c3f02

Observation 3e115bb4-420d-4c02-b88a-a10199752927 · outbound

This paper cites Inftythink: Breaking the length limits of long-context reasoning in large language models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Inftythink: Breaking the length limits of long-context reasoning in large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:59.964450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:59.964450Z digest=sha256:7c53f9a2bc3275b8faa9bcbfbefead7f97c1e24fff3a7a369d7a89c0bb296869

Observation 19b1f7ff-1647-42a6-8aef-00e65e33485f · outbound

This paper cites Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:00.072659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:00.072659Z digest=sha256:db47ca6f613ef9d7f3ca4e177efcd7d10bc8f8d5100b4bc8bbfcc67ae762014b

Observation d697b007-2cf2-4794-be4f-2283440f5ce2 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.328473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.328473Z digest=sha256:982e6116b2f9060ff182e098edc31199434db28809f86a56018a0e85976c612a

Observation 1673fa13-dfda-421b-8440-b4887d3289ba · outbound

This paper cites OpenAI o1 System Card.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs OpenAI o1 System Card

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.821730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.821730Z digest=sha256:43443b6b2a7d979cb364b713e31d04cf04656feb0b09626960474159acf7869f

Observation bd220da0-e9ad-43e9-b6fd-e1996fb0e4de · outbound

This paper cites Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.523176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.523176Z digest=sha256:3d8b867974dfe64da7de8c88498003fe0f5a5aac9b1d3fc079fa3d6c71d4b8c8

Observation f67615ef-5366-4947-a05b-f5a7327e10e0 · outbound

This paper cites SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.917129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.917129Z digest=sha256:1d4c752914ba6f37e04f728a5e0f0c1b3f2d1b55fb2dbd519f5ee086c3a82774

Observation 127a6973-4793-4e5f-bb46-1190510a15f9 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Evaluating Large Language Models Trained on Code

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.197865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.197865Z digest=sha256:cf4e1f11ed895a714e5d39680101d275e7906e99be3cf7402841346b6de2e0ae

Pith citing papers

No inbound Pith citation observations are available.