Pith. sign in

Paper Citation Record · LEDGER

Mars-PO: Multi-Agent Reasoning System Preference Optimization

As of 13 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 3 inbound Pith citation observations for arXiv:2411.19039.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19039 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:39:36.971073Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:51:18.786861Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:47:25.851906Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44978676-6343-4ace-ac84-4735b30733d9 · outbound

This paper cites online" 'onlinestring :=.

Mars-PO: Multi-Agent Reasoning System Preference Optimization online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.844729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.844729Z digest=sha256:a765fe2e51eee4cacd165b9d0c08f5659a23e9d49678ae07cccf0e4520a05fc8

Observation fe1978fb-381b-4230-a4f3-3abbd91516c3 · outbound

This paper cites write newline.

Mars-PO: Multi-Agent Reasoning System Preference Optimization write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.849582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.849582Z digest=sha256:52dfb5076d7327e212dd3cc60d841757500790b9697fe57a3fbbfb1ca8d877ca

Observation b69d1bd1-275c-4047-8019-3f99d87201af · outbound

This paper cites Llemma: An Open Language Model For Mathematics.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Llemma: An Open Language Model For Mathematics

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.853722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.853722Z digest=sha256:17c1877bc51f4294165d2a028831d214cf8389707b5ab80d96088bc1a549ff4f

Observation a2e3ab6e-0316-465f-9de5-6442ec872262 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.858175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.858175Z digest=sha256:dc75b87186be68731ed7e03aa2bd03179af1e7c714918de7f9725f719c0b213e

Observation 4c76c731-d74f-4767-b21a-cd4ce34a847c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.862423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.862423Z digest=sha256:56ac7a306fd38bbadc34030b074cca284a7491507174bd646daf34312b91e528

Observation 5bc2d494-1f8e-4a1e-bdd4-cccab19ecc91 · outbound

This paper cites an unresolved cited work.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.866310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.866310Z digest=sha256:139e5e3acdc654c2accb56e02e19961ac99767e31ed03ac8d7276163a155dd06

Observation 90534f5e-f1eb-4eda-bd20-e0875d954a79 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Mars-PO: Multi-Agent Reasoning System Preference Optimization ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.871101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.871101Z digest=sha256:77db2b53d4272ead7a1cf9784ad2812e7289d7ae73fd389dacd0e2b5babe6022

Observation bdafeecc-3b85-4da4-a92d-de7193ada97d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Measuring Mathematical Problem Solving With the MATH Dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.875004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.875004Z digest=sha256:ad0c17fb7097b0103e0f8fe5a35d49f04abd2925d0ff5e1ca07aa5cee38429e2

Observation da187040-5038-43de-a5f9-d2fffd4447d7 · outbound

This paper cites Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.879146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.879146Z digest=sha256:6e9e4464c33b7b1c03f4f965bb5728643ffce41ff8b0005fe5caced2a429bcae

Observation 90bd06fd-699c-4981-a096-0ddeb21a5be2 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.883872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.883872Z digest=sha256:c1b5f2ffd81a71013bcefab046f3bdd001a426671788ec85819e17710eac1064

Observation aea2ab23-3e97-4f52-8a6c-2a852797e6cc · outbound

This paper cites MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs.

Mars-PO: Multi-Agent Reasoning System Preference Optimization MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.887300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.887300Z digest=sha256:df2e9ebf318f7da6ec40a0bf0e2ff2160ffac8f1669ff6dc57e3c249fd518a7c

Observation 092308c8-8fbf-432d-a907-950982865d6b · outbound

This paper cites Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.892177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.892177Z digest=sha256:0b9bcc9f3b4fb2c89a16611645ce84092568fa114c159dc250dc6c2e30fad24a

Observation dbba573a-fd9e-4537-a313-a3d9919ee766 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Mars-PO: Multi-Agent Reasoning System Preference Optimization WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.896228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.896228Z digest=sha256:d48a1b868f261c8d4abce279d25bc977a3983faa6764044b6ddcc70fc2fc1085

Observation 5196a5bb-39d1-4d20-88ca-3d17f19d2bbb · outbound

This paper cites Iterative Reasoning Preference Optimization.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Iterative Reasoning Preference Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.900256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.900256Z digest=sha256:a0925bf81d6a9987c365a146331d0bed40584fe2d648d54526aebd2654f4b507

Observation 3c7ce10d-34d5-4ea7-a7b7-f48b5feb2f5a · outbound

This paper cites an unresolved cited work.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:39:37.256819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:39:36.904514Z digest=sha256:2aa3f7de046d99ffa97f13433de1e7e54fc7339425b1b4b2a27f994903e9945e

Observation 862f3825-faaa-4363-ac4d-35ecea58e809 · outbound

This paper cites an unresolved cited work.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.908574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.908574Z digest=sha256:47e7f66362e609ecac180a0e2b3f15172aca85291a1dee60324929fdf6b24d10

Observation 0c753632-aed5-457f-9e7c-86c04c976b40 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Proximal Policy Optimization Algorithms

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.912463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.912463Z digest=sha256:921beeb9338d8b3adef3adb0e5283fb80b9b991102dc9af292f08af90028da5a

Observation 2fca0c12-0183-414e-8c32-758177f048b7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Mars-PO: Multi-Agent Reasoning System Preference Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.917079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.917079Z digest=sha256:768e7a8f121bd68c7748c61e3afdaec7f89584faf9f0f56518035fec11cddf35

Observation 7338620b-1acc-47d1-be50-6ee3f5322206 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mars-PO: Multi-Agent Reasoning System Preference Optimization LLaMA: Open and Efficient Foundation Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.921843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.921843Z digest=sha256:b3c151e212ce30e14a28bc410c34370225be6588e13313c406fb5f73e0c70aed

Observation 3b8ff6cf-443f-466f-ade9-ed8745ad3009 · outbound

This paper cites MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning.

Mars-PO: Multi-Agent Reasoning System Preference Optimization MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.927288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.927288Z digest=sha256:b8030c234bd9d3da829bdc246eb8775bf696cdd13d5c196f634594fa2a04bf68

Observation e5b6f45e-af7c-439c-9150-1561a4ece9ba · outbound

This paper cites MathPile: A Billion-Token-Scale Pretraining Corpus for Math.

Mars-PO: Multi-Agent Reasoning System Preference Optimization MathPile: A Billion-Token-Scale Pretraining Corpus for Math

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.932384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.932384Z digest=sha256:af0577cbfd1d273fe19a67c294413ce5f3da7cf17560ba42bea128b5165affef

Observation f3b16405-76e5-4512-84d3-52bc351c91b2 · outbound

This paper cites an unresolved cited work.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.936765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.936765Z digest=sha256:e32bd31b2d51cd45523ccba6202f2c488dc7ecfdb0901b31b5a8178a46db8495

Observation 721c0176-7414-437d-b0a2-011c714ec572 · outbound

This paper cites ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline.

Mars-PO: Multi-Agent Reasoning System Preference Optimization ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.940937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.940937Z digest=sha256:1719d1a022c314eff25be2683e2e11f01e2f90df65fa99f0a85d1f32b177c797

Observation 7adcefdb-fd75-4394-af8b-c99317d899a1 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.945956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.945956Z digest=sha256:ec3a0b3d4854219ebe621c71740321954ec19fc160e8ab04bb25bb423eb58485

Observation a1da0a09-133b-419e-a065-991a3d39e13e · outbound

This paper cites GPT Can Solve Mathematical Problems Without a Calculator.

Mars-PO: Multi-Agent Reasoning System Preference Optimization GPT Can Solve Mathematical Problems Without a Calculator

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.950309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.950309Z digest=sha256:7f15b5b0638f57752a84e27cef6d2e13f431ba6c9505cbe02fac0ea04cbd46c8

Observation 761051ce-13c4-49a1-9b8f-20f48b49cc5b · outbound

This paper cites an unresolved cited work.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.954267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.954267Z digest=sha256:9acfd0af741d7fcf835b07477bdd74748a6e9c3783fb4eca44dd45c8c82b3b60

Observation 305984be-5c6d-4764-8f89-5b14cf9c8b2d · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Mars-PO: Multi-Agent Reasoning System Preference Optimization MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.958091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.958091Z digest=sha256:44ad4f42dd3a44c567a49b84ce8d435de0ac4c2e0e1b8366dbabcf59a664d0eb

Observation dda0896e-d2c9-4f86-b7a2-0b527704616a · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.961902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.961902Z digest=sha256:d7a310ec68a7497a4d91cfa93eee41efea5fb5b2b0b5df14c8a57ca24a1d8c61

Observation 2df16948-1a9b-4da4-a342-2654562421b1 · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Mars-PO: Multi-Agent Reasoning System Preference Optimization MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.966610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.966610Z digest=sha256:15504089d44b2af29b14af71feb41236d896c2ee41e672f87d87bed3a5ec3c18

Observation ec6e52b0-60f0-4bfc-a642-8abeb6402a57 · outbound

This paper cites Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.971073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.971073Z digest=sha256:1053f7ea3eb771666a260ce8d9ef9d7de148d15ce1ab2054e9df61915fcb4bdc

Pith citing papers

Observation e0fde64d-8c8e-47f3-8ec0-1608e7cb85ed · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Mars-PO: Multi-Agent Reasoning System Preference Optimization

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:25.853804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T18:39:44.696961Z digest=sha256:077e943c40580e9ef725a94063fcf5dcd1ff182a62956804b90a40eddf74f226

Observation fad21fef-7abb-4ae9-bb49-ef5246c26335 · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Mars-PO: Multi-Agent Reasoning System Preference Optimization

Reference 160

Resolution
unresolved
no resolver link, observed 2026-08-02T12:05:17.904146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:05:17.904146Z digest=sha256:17f103f8026f2d18b741ec929773991629725ed5968b2bda5f59c259f7361a1a

Observation 89acc868-36dd-4f4b-a1b6-66445ed40372 · inbound

Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization cites this paper.

Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization Mars-PO: Multi-Agent Reasoning System Preference Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T00:51:18.786861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:51:18.786861Z digest=sha256:64f39f0ba6a5eb67ec0f230065d6ca774111b97ade487a1d9b7b637468d917d7