Pith. sign in

Paper Citation Record · LEDGER

Mars-PO: Multi-Agent Reasoning System Preference Optimization

As of 13 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 3 inbound Pith citation observations for arXiv:2411.19039.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19039 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:39:36.971073Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:51:18.786861Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:47:25.851906Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44978676-6343-4ace-ac84-4735b30733d9 · outbound

This paper cites online" 'onlinestring :=.

Mars-PO: Multi-Agent Reasoning System Preference Optimization online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.844729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.844729Z digest=sha256:5c929b437949aaa75af829c47fe70f3722667274e84dc89a363cb5234127430c

Observation fe1978fb-381b-4230-a4f3-3abbd91516c3 · outbound

This paper cites write newline.

Mars-PO: Multi-Agent Reasoning System Preference Optimization write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.849582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.849582Z digest=sha256:c60ad2a10ed6eab0f54b2b85869912a8e92a57fcc18611b4c2afcaf674709386

Observation b69d1bd1-275c-4047-8019-3f99d87201af · outbound

This paper cites Llemma: An Open Language Model For Mathematics.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Llemma: An Open Language Model For Mathematics

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.853722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.853722Z digest=sha256:5fb7aef441a4eb543e75ae4d1e730a7193e8ecbae74769c32fa0cba9275c27f3

Observation a2e3ab6e-0316-465f-9de5-6442ec872262 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.858175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.858175Z digest=sha256:4e796541ace2b66dae9987f0ce4789e6be426a5193231839df4c8b5de73bc066

Observation 4c76c731-d74f-4767-b21a-cd4ce34a847c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.862423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.862423Z digest=sha256:12f90d3801ac77129f45902771fc3241485a971b6140d56b086a6812a7e20b63

Observation 5bc2d494-1f8e-4a1e-bdd4-cccab19ecc91 · outbound

This paper cites an unresolved cited work.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.866310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.866310Z digest=sha256:f2c89903bd159520d255b565e590b26431099f0c468508d16993e8acaf3a6970

Observation 90534f5e-f1eb-4eda-bd20-e0875d954a79 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Mars-PO: Multi-Agent Reasoning System Preference Optimization ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.871101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.871101Z digest=sha256:8fe9ff983231e0ff5b63ffb2d87c1943991acfb6dc2412775c559586e0b004b8

Observation bdafeecc-3b85-4da4-a92d-de7193ada97d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Measuring Mathematical Problem Solving With the MATH Dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.875004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.875004Z digest=sha256:8ec2422648d878d2ae24628c2be0f62a92170af43257f7908dcda194d8f2d8b7

Observation da187040-5038-43de-a5f9-d2fffd4447d7 · outbound

This paper cites Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.879146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.879146Z digest=sha256:82c12a86d7b0d47f664f874932e4bf02c02fd77466a49fdb3ac67987ca834800

Observation 90bd06fd-699c-4981-a096-0ddeb21a5be2 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.883872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.883872Z digest=sha256:f043bafb283853dfb96854d1957a51620bda1226acda7f8b702598244274ebb8

Observation aea2ab23-3e97-4f52-8a6c-2a852797e6cc · outbound

This paper cites MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs.

Mars-PO: Multi-Agent Reasoning System Preference Optimization MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.887300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.887300Z digest=sha256:bf92e5b9a8415af6364b90690bcf2141c6e80cba287a89ab277f115b41d466e2

Observation 092308c8-8fbf-432d-a907-950982865d6b · outbound

This paper cites Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.892177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.892177Z digest=sha256:e9ff5eeed3de9a4ca24b64758d60aca4389c9def5d07f0fc68d6f1f9874440ee

Observation dbba573a-fd9e-4537-a313-a3d9919ee766 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Mars-PO: Multi-Agent Reasoning System Preference Optimization WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.896228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.896228Z digest=sha256:c91cf6b2ee8ab3f44c946ea539d93604d00a2c4a980974cdc5dcb866e4ac9c46

Observation 5196a5bb-39d1-4d20-88ca-3d17f19d2bbb · outbound

This paper cites Iterative Reasoning Preference Optimization.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Iterative Reasoning Preference Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.900256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.900256Z digest=sha256:7b7a2e778b9c0a5890e3208bb299babdafa0b6333dacb775fe17a952485938e7

Observation 3c7ce10d-34d5-4ea7-a7b7-f48b5feb2f5a · outbound

This paper cites an unresolved cited work.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:39:37.256819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:39:36.904514Z digest=sha256:697d16fb4c7cb4516dd692480386be2f436bf09b3ae6cbfe9adc1ad1e8daba9c

Observation 862f3825-faaa-4363-ac4d-35ecea58e809 · outbound

This paper cites an unresolved cited work.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.908574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.908574Z digest=sha256:5cff4b720555ef3632eaefe9a93d83c5fd7be61e18dc842b0d4ee1205db3cd18

Observation 0c753632-aed5-457f-9e7c-86c04c976b40 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Proximal Policy Optimization Algorithms

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.912463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.912463Z digest=sha256:6349665593d26e7920d187e8b869da4e237c592b6d883450ad8eeb15a862b43e

Observation 2fca0c12-0183-414e-8c32-758177f048b7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Mars-PO: Multi-Agent Reasoning System Preference Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.917079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.917079Z digest=sha256:3f7e4ac028b231c39b0af61faaadad03bca000077ae55114cbeeeb4f0cf371fe

Observation 7338620b-1acc-47d1-be50-6ee3f5322206 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mars-PO: Multi-Agent Reasoning System Preference Optimization LLaMA: Open and Efficient Foundation Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.921843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.921843Z digest=sha256:490ae7b8e9fccd21e238c182c3ef1c38084ca5744e0ebed14f725c9078748903

Observation 3b8ff6cf-443f-466f-ade9-ed8745ad3009 · outbound

This paper cites MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning.

Mars-PO: Multi-Agent Reasoning System Preference Optimization MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.927288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.927288Z digest=sha256:7c7663cfd31c4561da0d951e3224fd40b1ca42aff531117b1926f6f02b84472b

Observation e5b6f45e-af7c-439c-9150-1561a4ece9ba · outbound

This paper cites MathPile: A Billion-Token-Scale Pretraining Corpus for Math.

Mars-PO: Multi-Agent Reasoning System Preference Optimization MathPile: A Billion-Token-Scale Pretraining Corpus for Math

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.932384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.932384Z digest=sha256:246a35275e9873dfa5e4a7c96fd9c763673115f1eeae0a6b3503ae5057cd43c1

Observation f3b16405-76e5-4512-84d3-52bc351c91b2 · outbound

This paper cites an unresolved cited work.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.936765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.936765Z digest=sha256:043cf56928ba2bbf98ce145a631f8c3ebc8e3db1d16653fa2870eca318374d1a

Observation 721c0176-7414-437d-b0a2-011c714ec572 · outbound

This paper cites ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline.

Mars-PO: Multi-Agent Reasoning System Preference Optimization ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.940937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.940937Z digest=sha256:0a3d846e5b4a5e3219dede2c81d1eb855277e3f2fab696269ffb9a9cb8d7d981

Observation 7adcefdb-fd75-4394-af8b-c99317d899a1 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.945956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.945956Z digest=sha256:68f34497f0e1d6b6f17108d84d9b39f65b5bcb6e4a48d86ee9bcff944632513a

Observation a1da0a09-133b-419e-a065-991a3d39e13e · outbound

This paper cites GPT Can Solve Mathematical Problems Without a Calculator.

Mars-PO: Multi-Agent Reasoning System Preference Optimization GPT Can Solve Mathematical Problems Without a Calculator

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.950309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.950309Z digest=sha256:716f44c50446988281044a3594c0ab7ea344c6adc04ca4b12130e010ba118581

Observation 761051ce-13c4-49a1-9b8f-20f48b49cc5b · outbound

This paper cites an unresolved cited work.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.954267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.954267Z digest=sha256:75c01e87d93f865a66efcb00f363112d1639da225879609ee918f26e95df7986

Observation 305984be-5c6d-4764-8f89-5b14cf9c8b2d · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Mars-PO: Multi-Agent Reasoning System Preference Optimization MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.958091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.958091Z digest=sha256:e4061142e0394a9667cd9a06dcb477e53187f0e46fdee8d03ff20090e0b7bb02

Observation dda0896e-d2c9-4f86-b7a2-0b527704616a · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.961902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.961902Z digest=sha256:e9dcaf87383b9aa0111982220ac3162d9a41b3b207dfc372775bb61bd380185e

Observation 2df16948-1a9b-4da4-a342-2654562421b1 · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Mars-PO: Multi-Agent Reasoning System Preference Optimization MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.966610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.966610Z digest=sha256:ff28ceed74a6b30a5ba3a8b84a2b015a7a265164585d43dfe6af8765b4618d0a

Observation ec6e52b0-60f0-4bfc-a642-8abeb6402a57 · outbound

This paper cites Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.971073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.971073Z digest=sha256:996e7ae13e3926ed4c919fb1083302a44c2e01f2889e56534e232623af463974

Pith citing papers

Observation e0fde64d-8c8e-47f3-8ec0-1608e7cb85ed · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Mars-PO: Multi-Agent Reasoning System Preference Optimization

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:25.853804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T18:39:44.696961Z digest=sha256:5a393b178511aef943a6b7d75dc9cd0a55a8179fe95322d073d17091c82ceade

Observation fad21fef-7abb-4ae9-bb49-ef5246c26335 · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Mars-PO: Multi-Agent Reasoning System Preference Optimization

Reference 160

Resolution
unresolved
no resolver link, observed 2026-08-02T12:05:17.904146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:05:17.904146Z digest=sha256:73d91fd7b84baf894f02ba7388d9d527eb5a76d06463f0b628b4ebf44495bf08

Observation 89acc868-36dd-4f4b-a1b6-66445ed40372 · inbound

Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization cites this paper.

Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization Mars-PO: Multi-Agent Reasoning System Preference Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T00:51:18.786861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:51:18.786861Z digest=sha256:d5924e1e99bfa6c644d7dd1f47c07c7d2a24da0e172d08194848802f81997d2c