Pith. sign in

Paper Citation Record · LEDGER

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

As of 15 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 4 inbound Pith citation observations for arXiv:2507.14683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14683 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:55:42.622845Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:56:39.504430Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.510220Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f41dd2f5-c890-404e-8a61-91943603b5a6 · outbound

This paper cites Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.505183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.505183Z digest=sha256:bc1eaa028af3819c4ea50e2ee1506acf4e0b9b3f6a873390a4762b4011fc735e

Observation bf4c0288-0f85-4589-9c67-3dd807430791 · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization OpenThoughts: Data Recipes for Reasoning Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.523415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.523415Z digest=sha256:cd562f458ef308db5c122eaebcdc7cc9d9778f51580f80bf9a241ea3b659b348

Observation e3323e53-83b4-4e52-a678-a60a1bac7677 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.528683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.528683Z digest=sha256:0602127ccfa5ff698c752b030424b6b985f3f0bf1d8dc81926964e3957a6d25f

Observation def4558f-e3ab-4f03-9a65-2dc11a981de7 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.534276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.534276Z digest=sha256:6483838bd93030df222d5fa3491bea26c59c621e1d09189924b70fe905947ec0

Observation 851a1dcd-ee08-4c1c-bd10-434df931e55c · outbound

This paper cites AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.539416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.539416Z digest=sha256:4877c849c287e85d69035bca717d226278e526896ff93c9917a7ba8ad987aa9d

Observation 95b5c68a-747b-40cc-a7fe-7e7cbf65108b · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.546685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.546685Z digest=sha256:c6454b74ceca36dccf73e85b7286d622ac84104e461dff6b7e7d2f7eed76fd8b

Observation fe408ce1-a139-472b-86a2-cd9a40d73c78 · outbound

This paper cites Let's Verify Step by Step.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Let's Verify Step by Step

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.551665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.551665Z digest=sha256:43bb38dcb2ce61c2d0d017651a64f959950ba130ac94c9f4c5e1427847bb289b

Observation 9b541b4f-9a42-489e-8a94-581485138278 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.556087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.556087Z digest=sha256:40961dbe4da3c8d3801581cd65348381ebe17143474b8bc00c95bf3e4a721937

Observation 55adf272-4e07-4ef0-975f-904dc0245558 · outbound

This paper cites Magistral.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Magistral

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.560228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.560228Z digest=sha256:f230b962028cc895f6c4d41e7c566e6708e653fcf8e8832f9ed0977335a746fc

Observation 26475905-cb27-4011-9119-0f0a525756e2 · outbound

This paper cites Training language models to follow instructions with human feedback.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Training language models to follow instructions with human feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.564588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.564588Z digest=sha256:51afba52732be8e71d323d96e502d02bd8b99e976ecf42e1bd78d0d6d00a37af

Observation bed715db-3950-435a-ba11-91f47d4e0237 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.572359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.572359Z digest=sha256:17ac2613366372ac392cb04ab28fb10ef973cea82866da30676e175fe59b6d9d

Observation 41314aea-683b-4c79-9ca2-95e98c39808c · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Gemma: Open Models Based on Gemini Research and Technology

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.577131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.577131Z digest=sha256:97e41b8ac2c9893fdc81ca648a53d8e65327ad290c3546311f337ff10e6e7a8e

Observation b00e935a-2c58-46b2-ab08-e8702f93d437 · outbound

This paper cites DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.580195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.580195Z digest=sha256:ed71c0c5776f08781d49137c30f35f895af698e36f641b8d4da6ddedfe0cbade

Observation aea26196-b430-4ef0-bc43-bcd375bec63c · outbound

This paper cites Measuring short-form factuality in large language models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Measuring short-form factuality in large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.587682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.587682Z digest=sha256:fcb0e2a378179d0a4b43e38169e58ee83b4b884cf7946775eceb0b5a7a9d1c0a

Observation 45047020-db86-4237-9b92-81b3f553d448 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.592502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.592502Z digest=sha256:0642b299c2b256df7c157525ed7802c6f1f9cd80e0c51ce7ed5043d9bc443feb

Observation 0ef95cc2-b12f-4a6b-8a84-7b3cb7a23793 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.595925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.595925Z digest=sha256:b55067d025ac3dab4cd0e1546213fe39bb9b467ab6f20b3aa76b6fd4984d59a6

Observation d3e2fe4e-2353-4b39-b05f-cf96aee9919c · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.604562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.604562Z digest=sha256:5a29ef91a96d3d506ebf0c2243d397becb620bcae6279de205a2e54500a711e8

Observation 7917bd7d-861b-4f80-9ded-b4b348468bba · outbound

This paper cites HARP: A challenging human-annotated math reasoning benchmark.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization HARP: A challenging human-annotated math reasoning benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.608109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.608109Z digest=sha256:878c9839faf13788188c0f81c09421cb5d9350b9dc64f271a5743b78c086be6a

Observation 0815638b-df74-481c-8fa8-bdd6e337cb8d · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.611272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.611272Z digest=sha256:bf07b7530e9f2049900f6521fe2469042fc1f972a950437a9abad421aed14743

Observation ea541bba-8763-401e-8d46-369a528fb236 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.614460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.614460Z digest=sha256:72b8e59afc67c55d32b5d32cc9626e187c9eae39f6fb89ddeaab24b20eff8683

Observation 959f1a18-0010-458e-a947-05e298a39acd · outbound

This paper cites 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.618381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.618381Z digest=sha256:2e6e4a4c92fbc90a16841edc265ad239382079e886db16800973e668ea667cfe

Observation a1a3ce8a-cfa6-48b5-8796-7fc38837fe9e · outbound

This paper cites 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization 1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.622845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.622845Z digest=sha256:c2f0fd5d76f111a2c6fdfd48ab25efb05ac10c55a655eb17a8d4600f9e1a1ee8

Observation b0e8d495-3b3e-4664-b5a9-ae421033963b · outbound

This paper cites PlanGenLLMs: A Modern Survey of LLM Planning Capabilities.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization PlanGenLLMs: A Modern Survey of LLM Planning Capabilities

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.583833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.583833Z digest=sha256:97c6e4456019484bcfe56bbfe2ac1fd04c87b3f43cdd76d358c6f1aa55e55673

Observation 0c79cd03-d79e-4acf-a928-ee92ef571748 · outbound

This paper cites Proximal Policy Optimization Algorithms.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Proximal Policy Optimization Algorithms

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.568123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.568123Z digest=sha256:11886a45386e9d83e2c0021a648a7bd0ae88cb97cad8c8effafbf965523a61f0

Observation 5fe026e6-8a8d-4273-9404-f310a2f38811 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.600378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.600378Z digest=sha256:cd6879e0dc275a5124241ae1e66d6ab083413f81d06a65ef3544e71e2317e26b

Observation 95c82986-bede-402c-81e8-de0a59f8ff2d · outbound

This paper cites The Llama 3 Herd of Models.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.512830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.512830Z digest=sha256:a515d71ab4981c867eb3d7db4d7d060fe120ed189190fae39181ce344e2c3bbf

Observation b5ca1b6c-54fa-4e40-a847-b45d0f7aee59 · outbound

This paper cites Reasoning Does Not Necessarily Improve Role-Playing Ability.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization Reasoning Does Not Necessarily Improve Role-Playing Ability

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.518689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.518689Z digest=sha256:f69e549338097efb64495c9352f35a9fc15604e9258598785c0c014abebeb4c8

Pith citing papers

Observation 5cb70732-447f-4e2c-aab2-faa64c78da79 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Reference 287

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.767561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:33db9be2f76b626e696567f942655f894049d600a216d1d980ee410f1268ea58

Observation 9285862f-99e2-4b13-bd67-3bfdfc80eda2 · inbound

Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs cites this paper.

Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:06:00.557948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T16:14:38.371700Z digest=sha256:1c6ae9ecd710a1635ce6484243c3e08ddb7a2b1763ac21487132b89485e15c94

Observation c00b2957-89d4-47bb-b0ee-42962a5b5236 · inbound

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution cites this paper.

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-30T00:14:04.540542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T22:56:39.504430Z digest=sha256:812b4a496fc7e47ac00ebd0206dff1165437f59f8116eee079bb62beb070d9b3

Observation 05ef2408-2264-4ead-839b-6940bffe0e8f · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Reference 288

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.511537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:3d7d687df3605cf1d0eef12b1bdfc074994f48933f08111cf0e6168030382e5e