Pith. sign in

Paper Citation Record · LEDGER

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2411.14432.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14432 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:30.841430Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:18:57.819941Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d6f69275-87bb-439a-a51b-a089d5d6d31d · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 162

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:23.531557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:8eb0f91a7873712034dacb65bc150dffd080188fe7657f5d8ab657cc0142bc42

Observation 8a505bfb-9f47-4ae9-b5b0-807212b8c0d2 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.752430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:431b5c5d26267eaf3bae1ab072cef84d0b7f538ed8273247f247c841342fae96

Observation 0b8427e1-700d-4a23-a967-7abf67faf2d8 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.821657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:b18a3164c09e816eaf4908a31b05b0d27f8f290ea71539e2f81596450fa63d3d

Observation 66edf0f7-809f-4720-b248-3123724d02d7 · inbound

Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models cites this paper.

Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:30.841430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:30.841430Z digest=sha256:4a1c0ffea28f8160629327a3bfbe823d8b6054b0659026b0a43e3b92e3424342

Observation c67c16d5-3ab7-46a2-917c-ff48f0d0e502 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.861983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.861983Z digest=sha256:93cb36baf59ed0eaabe1a8d627466f6fc0dfe00ada52fc6b9020fb62c0b21aa9

Observation e501b4aa-ab4b-48e9-bb12-9dc7b57355a7 · inbound

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing cites this paper.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.148574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.148574Z digest=sha256:1a33749ee970974e96b003b4ecdddf4377da2b11e7b92ac65ca554f403fc9164

Observation b86b18bf-b91f-452c-bf36-547c9fa8a0c5 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:16.974601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:16.974601Z digest=sha256:4787b07360102e6cce99197304ddbd3c0297483a7b8377aee1e463ae7a7ea84c

Observation b430029a-c757-45c2-a86b-c954d6c1023f · inbound

ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding cites this paper.

ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:54.394713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:54.394713Z digest=sha256:96ee543376e23da9c3414defc4a4ac8a52e1a517bffedac7ec61c416310b752a

Observation ae7909c2-2307-423c-b177-905206a5db85 · inbound

More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models cites this paper.

More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:31.645915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:31.645915Z digest=sha256:b3f3423146028a2bd06b63a5ba9fab35d0e24605d5ddde6055f052847281fa50

Observation 33e31527-5aaa-426e-b768-4588cb091ccd · inbound

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start cites this paper.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:53.376458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:53.376458Z digest=sha256:79e8b97d9f97a9dcdf589d4b04d791d561301cb7275a3b0d35beb368cf118449

Observation 8461ec5f-050a-4ead-b5ed-13ef51f6d045 · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:05:52.139807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:36facc89f3bba651548c891234bb0f0b5a750ca5f4d5788b1dc885acb49ed665

Observation d4247e12-4d1b-40f1-b03b-85d20ca21f7d · inbound

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning cites this paper.

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:47.438815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:47.438815Z digest=sha256:9e9bd651e9be7e4aa3690e37ada62af449203c0d6e1ca778b033b4778caecc77

Observation 67242505-2a02-4a27-9a10-cf79db3fc8f0 · inbound

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs cites this paper.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.484085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.484085Z digest=sha256:3388d06a4066b60e525c86f9e880f28bfc4a7a69111813d50d6cb16d653f47ce

Observation 5b7b1725-dd49-47a6-83dc-b953cace3bf9 · inbound

Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning cites this paper.

Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:37:15.034292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T10:34:48.849524Z digest=sha256:c689e4e5b9558e443c204a26bd2f9215cde9949c9bfdce951d8aff0624f18df9

Observation b21cc11a-7cb7-4090-b1d7-b43f9e8a02c6 · inbound

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning cites this paper.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.088125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.088125Z digest=sha256:53ed120684b386bec6de34f0ba2fd99a78109085c57b47014dae02c30980f450

Observation 6d67ff3a-1535-42b7-a981-6617d1b4012e · inbound

AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions cites this paper.

AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:27.248876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:27.248876Z digest=sha256:2b9168a0d09c777187c742936ef35467c84ec46689b008cd756004b31c9fa1b3

Observation a718607d-b096-44db-8248-2b437ce2e913 · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:28.953249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:28.953249Z digest=sha256:1f51c3d6eeb4f89c61fb80f1db501dfb061747a79b429a1bf389c1f4158f1e36

Observation e3a3aa44-7bc6-4e3b-b410-5bbf810d6588 · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.773345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.773345Z digest=sha256:89d9eaacd7303a0d837427be08df3631804134c1db5f0d4fdcb0f8c232fc56e8

Observation 932f6aae-56d4-4782-a543-4a0fdd953aae · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.105593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:9fabbfd9a4cd08db962f105b35e0a9f12c4b8cb108c663f90dc9cf38ba847ce6

Observation 1f66b485-f4fb-4722-afef-bb7155abd312 · inbound

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning cites this paper.

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:33:56.797064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:33:56.797064Z digest=sha256:1c47626e3f1293d9e71d07049e19f004c5cf7f97a3f251e705be68ecec0f4ac3

Observation 7811dae3-8c3b-4a69-ac98-5634d2a7ae5b · inbound

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding cites this paper.

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.423977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.423977Z digest=sha256:fd66d7d053ad19a6d896062bd1aa40940b9f4f55f7469424aede3b436dfab99c

Observation 52161b2b-6cfa-4777-94b1-0ba830cb8561 · inbound

SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning cites this paper.

SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:12:01.803711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:12:01.803711Z digest=sha256:81a3ba2144cb58a47181080015cffe5a333b056c64709061ebc3ed7ed3e5f735

Observation 10790b76-89d7-4bca-917b-6c638921155e · inbound

GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment Design cites this paper.

GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment Design Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T18:05:36.802960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:05:36.802960Z digest=sha256:4e8d28bdddb7f6ff493863cdb9aa7d60d661e4519ba354c307b0b6cb9f325432

Observation d62c896b-9f3a-4713-8d08-1799de8b65e5 · inbound

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization cites this paper.

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:11:32.319995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:11:32.319995Z digest=sha256:6498e19c41fc33113c9b47b35547f67dd04c482f7365596fb67c582b83a6c725

Observation 412a07ed-253d-4bc4-b971-b7bcb89e1096 · inbound

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning cites this paper.

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T17:46:56.202091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:46:56.202091Z digest=sha256:86999a1b997d10343c121533907e5f972bd337ad313f31971820998b1ccaad56

Observation 845b5c78-7297-445e-8109-9ba35228f073 · inbound

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM cites this paper.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.478593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:74f2203b3b4da8e8137613a3287ec00c8e32333156427e0e659388020a8498ca

Observation 3b544053-1907-4c6a-89cb-7715de179913 · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:57.823334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:36aa59d79874f3170e296773dbe3bf1831723f8be81cc767795796f625f580f6

Observation 70c9f362-4838-494c-b649-7b29b135f8fd · inbound

MIRROR: Learning from the Other View for Multi-Modal Reasoning cites this paper.

MIRROR: Learning from the Other View for Multi-Modal Reasoning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T07:15:04.612610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:15:04.612610Z digest=sha256:49b2c35109f5a344948cf8e8b8d3cd0ed870234d3e79e6c95048f958bdaba7fe