Pith. sign in

Paper Citation Record · LEDGER

Segment-Aligned Policy Optimization for Multi-Modal Reasoning

As of 10 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2605.01327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.01327 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T14:44:31.160543Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact14
  • verified fuzzy2
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57500d56-dcb2-4521-9ae0-83162d278553 · outbound

This paper cites Qwen2.5-VL Technical Report.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.841335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:81cccfdf1852a99b5666a892d53487c186754aac36a327d428903553e30d6695

Observation c3192f1e-9139-4f3a-a8a9-fdae2d14ec60 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.612219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:a60da2cb92f2414d0053df08e57cce8eee6e32c5a1b7c51e9cbae4d63ba2ad71

Observation fad68340-1666-48f6-87af-96c5bdc877ea · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.852419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:b16ff40e18b8517d476c1045dd24d4d3501d1175496c03b11f167f3a4a1a5570

Observation df3831d5-8104-445d-8732-2eb5f1bc24f0 · outbound

This paper cites Segment policy optimization: Ef- fective segment-level credit assignment in rl for large language models.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Segment policy optimization: Ef- fective segment-level credit assignment in rl for large language models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:51:07.834936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:1cbdedac262253c0c7448867e78416e790f71d10e7d0d2118bc8bfeb29abb6fc

Observation e6e99dd3-3679-4fcf-826d-619c11f07f40 · outbound

This paper cites Rectifying LLM Thought from Lens of Optimization.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Rectifying LLM Thought from Lens of Optimization

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.826343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:ce51f4b8ddf0353069d10d1ff678fce7c7a5e537462fa1482af359de09f95273

Observation bc31d786-8a70-4b1c-8076-a10e29f79c4e · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:51:08.012769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:8f892daf54234683a17379f0667be98a663f6d5581a0376ca3859d89dd5da9c3

Observation 8c7bb46e-d9c8-4dd6-9a0d-e94ff0e7a0e0 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.951371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:202ea2be68941f70415e26503f4bbbc564a909d73434be9de6d8e1df86db0b13

Observation dd2360ba-092e-4baa-9a4b-5c808e68b6a5 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:00:08.752459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:8374be78f7a191c93c7ccf9f3ac6b5704b7db0ca235561b3adb01c10f55ef102

Observation fff18f11-d21d-43f3-bc31-21a0e6f38277 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:04:44.937279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:ba4be78724368deb4f29d15dd64c4762631a6f04e9d0bfc46e12dd26d43442be

Observation b91e4203-3f45-4d49-a053-a7a36d8defcf · outbound

This paper cites Proximal Policy Optimization Algorithms.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Proximal Policy Optimization Algorithms

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.878376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:62203535033788797a4fd771c8ba6d5f50737b4a9da2813b1150284523eea251

Observation 488272d9-e371-4749-a8f5-851bd352c3b3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.904346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:d6123ad7abcb3f7bb8ae7ab7b39e256158e78eff56ae736abd2a714b8327e01c

Observation 7f8b7a38-3b6b-4ba0-80a8-754214783e0c · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.868222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:b70d6b476a0387fc6c50745b926a9a90528435640db724d88d62de34a8a5b618

Observation 83deb363-64c7-4970-958d-fc6bd40d1a98 · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:51:07.861678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:8ddc8441f9cf2c372763157dde9822d114d9c127e048bfce82dfb16027466c0b

Observation d78e37a4-71cf-47eb-93eb-2689954b05e7 · outbound

This paper cites Sutton, Doina Precup, and Satinder Singh.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Sutton, Doina Precup, and Satinder Singh

Reference 14

Resolution
metadata mismatch
doi, observed 2026-05-09T22:18:59.682350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:05a529b5ec14b6d4a80738f91d85dc81e05658f56fcb89ec10ddab772a42818f

Observation ee970ee1-02c0-41a9-88f0-c77b1c286bc6 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:47:26.275168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:45c5103e6760237df1cd5a5fa4a8066d797440e3ef8b2d375344576fbbd14491

Observation 910260f2-c59d-4abd-ad8c-d6f7827ee00f · outbound

This paper cites Single-stream policy optimization.arXiv preprint arXiv:2509.13232.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Single-stream policy optimization.arXiv preprint arXiv:2509.13232

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:51:07.964146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:04fc2c28a2add953cc84389451d031ded1d1b2872382770bb3dbd78e225448a1

Observation 2fe4635e-d5d1-4b5f-99aa-f632a0fb50c4 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:51:07.895095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:6cea1521347242396ae0de5e04e7a3527726795280bb307bd2a3fc4e6f0ecce2

Observation 6b88b11e-4781-4203-9539-f00e70edb1cd · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:b45e03114f719747537a4c8c88ffa0f011ecfc54c75416fbcacb19a2a8b61143

Observation ee07488f-8eb9-4377-b331-e03faf20c0ca · outbound

This paper cites Stolfo, A., Balachandran, V ., Yousefi, S., Horvitz, E., and Nushi, B.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Stolfo, A., Balachandran, V ., Yousefi, S., Horvitz, E., and Nushi, B

Reference 19

Resolution
metadata mismatch
doi, observed 2026-05-09T22:18:59.686726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:94f8f5493cd877aadbdfc07eb0ee7cde162ea92b1b474f6fd96071155f2fe7b5

Observation 2d067ac6-72f4-4e55-9123-4af0e02a0bf3 · outbound

This paper cites emnlp-main.668/.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning emnlp-main.668/

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:47:16.637643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:9d80d16451985ed60cfa829a967fc74a21b495606dc92efe526af3f5e59f266b

Observation dd39c7b7-2c8d-45cc-a046-8e1008d9c602 · outbound

This paper cites Group Sequence Policy Optimization.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Group Sequence Policy Optimization

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:07.932042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:488f25ccf037dbf870d91af70a1dd9b80610179c1e29ebe0b20c4a6589a55dcd

Observation 574323eb-0efb-4a20-8e7b-e5dc35346274 · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:51:07.917445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:42e6be8a5f947f6b1e052479f4d0eef4455f4065ef222445a8b671f7c7d0f98d

Observation 69ba39f0-073a-4adb-ad41-becb0ca9c2d7 · outbound

This paper cites Benchmark Settings.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Benchmark Settings

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:47:16.636649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:943158b23851107667c28ed7b7c2f78187a42e8e382c4543d5f9032f3000303b

Observation ccb2e44b-165b-4b55-9977-63aad079ea22 · outbound

This paper cites an unresolved cited work.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-26T01:47:16.646518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:1294cd34a01498715567df95bd96d329e5e7e8fb81bdb6fcde6daea3ef0763f0

Observation 46acefa1-ba61-4f9c-828d-f82f9f4b4520 · outbound

This paper cites an unresolved cited work.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-26T01:47:16.642592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:c63da3bf7626a5c3231364dfc7c0f7f0b1747ff97d0431351816122df5fad3e3

Pith citing papers

No inbound Pith citation observations are available.