Pith. sign in

Paper Citation Record · LEDGER

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model

As of 18 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 2 inbound Pith citation observations for arXiv:2505.19406.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19406 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:19:04.977166Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T15:11:57.228949Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T14:29:53.301476Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2015ac20-33ef-4a2b-983f-c547a9ca023d · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:01.915052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:01.915052Z digest=sha256:39be754552ab4755375f12552a2cd309bced8db51b82fa8ea54af13dfb548877

Observation b1e26433-bd32-401a-a42f-42851105daa1 · outbound

This paper cites Virgo: A Preliminary Exploration on Reproducing o1-like MLLM.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.318580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.318580Z digest=sha256:59344d0c4421e71ff6bd008bb8c1a4d3da06f0998bc894b7dc7c2f52a3c66392

Observation 18acaf29-64ab-4962-bfb0-4fb940ccd2e8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.504324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.504324Z digest=sha256:851e51ae7567af53dfd19c077a2c947fa7cfda77c3893df83357030d895b829b

Observation 82ab3d3d-13f9-4382-8cf0-d5cfdb58dcb1 · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.650097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.650097Z digest=sha256:8ee8c80bd19add8e029386dd74716546c39339b1fdd3154c81908c0344d9435c

Observation 4e063a37-5cf6-4970-b35f-34de18d07322 · outbound

This paper cites T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.805674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.805674Z digest=sha256:396345bd88d19ceb5d5b93b060cc409af67d93e72498187bf68fdfdcf915c3d8

Observation bdc4a137-b272-4730-8576-cbc00ffd0c81 · outbound

This paper cites Boosting MLLM Reasoning with Text-Debiased Hint-GRPO.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.964530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.964530Z digest=sha256:cb42a29ef7ff95a4d0b27e46d95649ffaeb4e57f138880a4dde1af1adac3c856

Observation 0a2542e1-2b09-4150-a480-e600b0aea9d6 · outbound

This paper cites OpenAI o1 System Card.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model OpenAI o1 System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.067077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.067077Z digest=sha256:c7a3bc3eca55b0f8bd5296ec2584f8a5760d21517513ccddfcd6b5a65ab69c0c

Observation c20c6894-9fcc-4cb0-a18d-f9177110cd8f · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.201102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.201102Z digest=sha256:c2ae0755df65c27581d28a6a1f967f30babb25e84a12d9556db7386c858dadda

Observation 5bfc92f5-42a7-4f04-9283-7df32d795d5b · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.442626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.442626Z digest=sha256:ff1c00b7e5bfc6264216ea795779b37cc08bfb7a525ad2c2f03cd5414f86a467

Observation a380a69c-12d4-4f35-a5f8-1577f7eb76b2 · outbound

This paper cites Metaspatial: Reinforcing 3d spatial reasoning in vlms for the metaverse.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Metaspatial: Reinforcing 3d spatial reasoning in vlms for the metaverse

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.584308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.584308Z digest=sha256:fad4a3c617f001e2e5ed5e60f45884bfe38f0acaabfdc3f35aced10c1570b7c2

Observation b9799e86-71b6-418d-afac-5e6ab7f019b4 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.674719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.674719Z digest=sha256:3f46e6e74ea43163ae002db54a282a11e6fab3237c4189f4a3742729d4ca4966

Observation 6a944832-4ef3-4194-9ba9-a99067b7f3c5 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.823999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.823999Z digest=sha256:d3003dd07f5c300c3c0bdd552e310e479c4cef1c6d5ab945d185f54340fdfd27

Observation 9de96a95-7ef6-4650-b085-1905bfe6dfd2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.941349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.941349Z digest=sha256:57a8d1a80bc5419e731c6fa1eb8923e49478f2e4bb32abd16f15fbe0878116e1

Observation 2c5e2e8c-d1f1-45c9-808b-d94d2edc3773 · outbound

This paper cites Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.095763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.095763Z digest=sha256:08956f13b1bb8338a065752f4f65bd5af90ae89d6e3208d714cebabf193b1d1b

Observation 05057945-d3bb-4106-8c93-394e0f6d9656 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.218387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.218387Z digest=sha256:fbbbdbf7e19f15647ba13a5f92a6aeb661f3df1644fbc657b11e90ddc66db107

Observation 47598c7c-5742-4d49-b975-a4a009bdb9c0 · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.371188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.371188Z digest=sha256:4206d70bd6954cb614f14f6a5a8f89b422419c35fdd5e0886d38bc376c498050

Observation 9c7d8425-f4b5-4bbf-9fc8-2eb76722dfad · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.527707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.527707Z digest=sha256:7501025ddccf1ed6932d8de2263794dceb5ae868da5de81fe1900d3479f1f205

Observation 2f88aef6-8697-47bb-a8ff-b2997eefa401 · outbound

This paper cites Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.657935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.657935Z digest=sha256:ab91d31c1ceabf09183c6cd0094c46945aa75b4f15a2971e4adc32af45ffaf00

Observation cd2034a0-62da-402e-93e3-4bcac76dd10f · outbound

This paper cites Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.798512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.798512Z digest=sha256:3d2d17572cc3b1ebbe759a7ca31e1b648268ac7bc2d471ae0622964beec1f506

Observation 8b14d4b8-a2bd-44ab-8590-7b87811c4b6e · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.977166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.977166Z digest=sha256:213ccaf08d701c5d8e1e081a3f35e429b35e774d7e0a64fd65319f0f180040ce

Observation 71bf01ed-c9d8-45de-a256-2f291490d0f3 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:03.335586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:03.335586Z digest=sha256:7a4d4a8d3ba7475029101044dcb66a36ada95712c6d63f5cac288f65b0faa7b2

Observation 86d89c6b-f0d1-4f19-b05b-551e409a7a4b · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.165225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.165225Z digest=sha256:36251408f189fd5b5df92869ecb23dc92d035eb9c0a53e736e0f48895707350a

Observation f295fbe9-90b6-48b8-ad51-5d90ac34222a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:01.977150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:01.977150Z digest=sha256:e72879bc81bac6214b191d0f3730cc6aea3f0b7075d2fa77aaafb133fc938c86

Pith citing papers

Observation f2b08c53-5e61-4ef5-888e-dc17d57686b4 · inbound

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning cites this paper.

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:29:53.304356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T03:51:51.827622Z digest=sha256:450f7f1e8dae39abfd2cef6093dddccec8267bd93b8934d52324ce3becaa5cfc

Observation 031bfc5f-2dd0-4126-9ef8-6cc40ce19550 · inbound

Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval cites this paper.

Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:07.304743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-02T15:11:57.228949Z digest=sha256:f39848c4331636501e10e815e61e7d73ebd3122cd5d32812e38de0ef7c4a5c98