Pith. sign in

Paper Citation Record · LEDGER

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model

As of 9 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 2 inbound Pith citation observations for arXiv:2505.19406.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19406 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:19:04.977166Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T15:11:57.228949Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T14:29:53.301476Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2015ac20-33ef-4a2b-983f-c547a9ca023d · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:01.915052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:01.915052Z digest=sha256:183526defd27447b35d2dd5b484102391061c8e58b4921c1afc9a61ab9c8e9d8

Observation b1e26433-bd32-401a-a42f-42851105daa1 · outbound

This paper cites Virgo: A Preliminary Exploration on Reproducing o1-like MLLM.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.318580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.318580Z digest=sha256:f29a1b1f27451d6065f6732422e216df9f8bcb74a43efb4eda19e711cd0a1717

Observation 18acaf29-64ab-4962-bfb0-4fb940ccd2e8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.504324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.504324Z digest=sha256:823c877a040253dbfa620b0e0e58541adec61447ed701eaf08dcece3ce4f2b28

Observation 82ab3d3d-13f9-4382-8cf0-d5cfdb58dcb1 · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.650097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.650097Z digest=sha256:90de551a06c20562bfdecfe1cdb6c07b7643691396d64565460b9c2d96d37a0f

Observation 4e063a37-5cf6-4970-b35f-34de18d07322 · outbound

This paper cites T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.805674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.805674Z digest=sha256:509907b442722c1fe6c6d5d590da974b7925aafde5799a8341112cb6411a42ec

Observation bdc4a137-b272-4730-8576-cbc00ffd0c81 · outbound

This paper cites Boosting MLLM Reasoning with Text-Debiased Hint-GRPO.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.964530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.964530Z digest=sha256:3189070a6f475d09efb1d9b72ba24401d1b6d4333dfa15367db9a0864f911a65

Observation 0a2542e1-2b09-4150-a480-e600b0aea9d6 · outbound

This paper cites OpenAI o1 System Card.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model OpenAI o1 System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.067077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.067077Z digest=sha256:4a97d925fa27409a228725649bf8d93c1c870162b57090d9bc0ae05e58a76f21

Observation c20c6894-9fcc-4cb0-a18d-f9177110cd8f · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.201102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.201102Z digest=sha256:bd7532b353254308cf7ca64e839459f72168dbbe458eb5fca66fd95361b9f109

Observation 5bfc92f5-42a7-4f04-9283-7df32d795d5b · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.442626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.442626Z digest=sha256:a46e3097ff45f85417dab5f482ab0191449239fd7521b82b2dae6e8937aa314d

Observation a380a69c-12d4-4f35-a5f8-1577f7eb76b2 · outbound

This paper cites Metaspatial: Reinforcing 3d spatial reasoning in vlms for the metaverse.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Metaspatial: Reinforcing 3d spatial reasoning in vlms for the metaverse

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.584308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.584308Z digest=sha256:882ca85a2a20e29a5a080c8b9c8d2141a7a2c6bea2a2bb4710f787b735c22f6d

Observation b9799e86-71b6-418d-afac-5e6ab7f019b4 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.674719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.674719Z digest=sha256:5020bca38bac64864cb5b667967ef20aae52a2db1f1d05e18b54920ff9e7eb7b

Observation 6a944832-4ef3-4194-9ba9-a99067b7f3c5 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.823999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.823999Z digest=sha256:4cf76b40cc8e65b88a34b0824ae09b715fa5b0f980f9c261c72244450f1203f0

Observation 9de96a95-7ef6-4650-b085-1905bfe6dfd2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.941349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.941349Z digest=sha256:72aaa786d61108fcaad3d047607b2726442444f4c359d33504c45d99c0acbfd2

Observation 2c5e2e8c-d1f1-45c9-808b-d94d2edc3773 · outbound

This paper cites Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.095763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.095763Z digest=sha256:cba61fb8606c80565f0987fa77684eb60a06d9c801a56c73cda80dcfd0043ee9

Observation 05057945-d3bb-4106-8c93-394e0f6d9656 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.218387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.218387Z digest=sha256:6a3332cf96b0e6d3af7af4e7e4f69bea201d87221043e9834a91441d5452efbd

Observation 47598c7c-5742-4d49-b975-a4a009bdb9c0 · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.371188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.371188Z digest=sha256:852b0738cd63abbfc16b0140edcdc2309433ef324bd76b11d4e9a74e70accf07

Observation 9c7d8425-f4b5-4bbf-9fc8-2eb76722dfad · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.527707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.527707Z digest=sha256:3bde6067631bed4fc859b80f6b3c9bcae4d2190796832c3ddd0f762c5a33c202

Observation 2f88aef6-8697-47bb-a8ff-b2997eefa401 · outbound

This paper cites Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.657935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.657935Z digest=sha256:87beb955db6055a598cd17f3fb91b59862ba67854135db42b29c8490087ef1bd

Observation cd2034a0-62da-402e-93e3-4bcac76dd10f · outbound

This paper cites Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.798512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.798512Z digest=sha256:58f517f7da7c73f4586aaf5cde3f906cc9346b711899dba918970eea917807bc

Observation 8b14d4b8-a2bd-44ab-8590-7b87811c4b6e · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.977166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.977166Z digest=sha256:810ab9e721c11d4a03cafecd5d10663fc4e06c2299f8e051d4135c6a77af211e

Observation 71bf01ed-c9d8-45de-a256-2f291490d0f3 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.335586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.335586Z digest=sha256:17f400e17610bf2812edf4dac9663e98ae30b3a97030eac4386f773fc099dad9

Observation 86d89c6b-f0d1-4f19-b05b-551e409a7a4b · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.165225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.165225Z digest=sha256:9558c1f7efdab75d558c47eee7bf32808a09be6f61b781bfe31b302b5bea88e5

Observation f295fbe9-90b6-48b8-ad51-5d90ac34222a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:01.977150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:01.977150Z digest=sha256:73630d2a0d833adfd27042a957f979b5ec6a01e7691166defe5abf90ead30889

Pith citing papers

Observation f2b08c53-5e61-4ef5-888e-dc17d57686b4 · inbound

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning cites this paper.

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:29:53.304356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T03:51:51.827622Z digest=sha256:be6ce40c430cfec936e52e45de828c73171c21a3eeaa8aeb7e0810622b8f2997

Observation 031bfc5f-2dd0-4126-9ef8-6cc40ce19550 · inbound

Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval cites this paper.

Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:07.304743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T15:11:57.228949Z digest=sha256:f11955e7c6a4e4ad0f44e81781d970216415c77b0ecabae0eaa92b15ab6d39c0