Pith. sign in

Paper Citation Record · LEDGER

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning

As of 20 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.15706.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15706 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:29:35.826782Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c61dd3a6-0f87-4fba-bed9-78da6f6c6430 · outbound

This paper cites Qwen Technical Report.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.069984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.069984Z digest=sha256:812b88be7ae3dab50dabf5ff25f68dbb802447377dbb285e620d24b4081d0cc7

Observation a9ad4ebe-4bfa-4155-a22e-7471c5277dd7 · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning ORPO: Monolithic Preference Optimization without Reference Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.631728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.631728Z digest=sha256:589139fc315f85382a4760b4821279cf81097474510d12e76dffb6732b9588c4

Observation 3d4f9a7d-1972-40c9-b23c-eae567766c12 · outbound

This paper cites MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.957668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.957668Z digest=sha256:34943bde97eb3a309deec14c6e03075750343e922f63d9d6e17576f4b276783a

Observation 946a8099-ad50-43ce-ada0-929c11b49985 · outbound

This paper cites Orca-Math: Unlocking the potential of SLMs in Grade School Math.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Orca-Math: Unlocking the potential of SLMs in Grade School Math

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:35.252687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:35.252687Z digest=sha256:b4e5855b01bd0fa818b7de9b04cdffd689a176294141f4db4ccbbed9d53f366b

Observation 2d94a3e4-526c-4e4e-a214-63c1098c235c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning LLaMA: Open and Efficient Foundation Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:35.566083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:35.566083Z digest=sha256:cab02d070d736d4c1c6b169509ef05cb4f8dc80dc40675e82bb56ca97081591d

Observation 3b4612a2-3425-4484-9dae-1f571c030c6a · outbound

This paper cites MathPile: A Billion-Token-Scale Pretraining Corpus for Math.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning MathPile: A Billion-Token-Scale Pretraining Corpus for Math

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:35.680196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:35.680196Z digest=sha256:9cc16f3e4964ea1ccd5b01e344a4c62cc8ba0a66786fbb5554b70b1fc067a3f2

Observation f6560daf-019a-4a7a-85c5-15fd3c253c84 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Automatic Chain of Thought Prompting in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:35.826782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:35.826782Z digest=sha256:5f2b3d6e961fcf75e8fdc0217d45d502dd106706e3f18b5819e830a065006faf

Observation 9a397e16-b92e-468e-8660-2dbfc3d1c1e6 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Training Verifiers to Solve Math Word Problems

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.220523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.220523Z digest=sha256:61aaab2f49f9c662eb159be557b8820c71dbdf325d4a74cc91487c82698935e4

Observation ada7517d-270b-424d-bb77-6823d4c674a6 · outbound

This paper cites MathScale: Scaling Instruction Tuning for Mathematical Reasoning.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning MathScale: Scaling Instruction Tuning for Mathematical Reasoning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:35.374285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:35.374285Z digest=sha256:d65134b28ed9495ffff306cd3b1f59740bbb239ceefd5445bdf320c666a60f48

Observation afa863ba-4dbf-4806-9fc1-8135714b49ef · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning KTO: Model Alignment as Prospect Theoretic Optimization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.363833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.363833Z digest=sha256:f051d45857a476fc951c662b4e7345280cf4707c44e839488ccb7b3615d9d60f

Observation b40b769e-4c06-48e5-9d9d-e10fc34fc93c · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.795071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.795071Z digest=sha256:b10c83d9156b396861fd0a5c277b18af9101147ee4c557a2747e78107107de36

Observation 83c80812-fb4f-4f46-8b67-046aef42f866 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:34.481780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:34.481780Z digest=sha256:c0e7d5bff6fa76d95ace9f2b72d9a7d986ceb33781d3515c541c06cfe3c2caf0

Observation 930ec7a3-8608-4099-8aa6-f78e94c68117 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:35.084817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:35.084817Z digest=sha256:3cebee75e7b24870b133ab011ec88799a25b247db7ca3e983003c1ca39d6e460

Pith citing papers

No inbound Pith citation observations are available.