Pith. sign in

Paper Citation Record · LEDGER

Grounded Chain-of-Thought for Multimodal Large Language Models

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2503.12799.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.12799 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:34:36.765233Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f2d0cd6a-a6de-4a82-aeb9-aee55ccb2dbc · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.451307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:5e9fd9127a252778715b97bae3a940d92ccbb1e6cfb2f3d408b5e2fcefb29306

Observation 2ea0c3d6-bc3a-43b0-b421-951261aefd76 · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:36.765233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:36.765233Z digest=sha256:35dc69f9c7a9a7b17b5bad382c33c0be6e5be361ba0ae023c2667a6c76ba377b

Observation 33aafc2c-50ed-4f2a-8985-4d80ee31dcc9 · inbound

Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation cites this paper.

Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:57.113550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:57.113550Z digest=sha256:329f5b9bfea022980c87fe061fd0f6758be5a7513a76d65e8fba7b999dc18e55

Observation f2774bf3-af99-407d-af29-21f65c5c471d · inbound

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture cites this paper.

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:24.078659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:24.078659Z digest=sha256:d2c135cd537dc975b4c2dfd61e9fe3678be5cac4533c415c82a0af1f4c71fa1b

Observation 51891bc4-500e-4224-9a61-e537c1f404e1 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:02.881423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:02.881423Z digest=sha256:8f2b3a3d6329a629cc402802130560db0d122dd137234551fe4892f82afa3217

Observation 17b940df-2077-4fa0-ac82-5b632675a6f1 · inbound

Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning cites this paper.

Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:36:25.208637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T13:33:12.508639Z digest=sha256:955511087d3423f004bf671004404689dea57341b210e8aa5ab918a3f670987c

Observation 59ee0f4e-8e78-40bf-9884-17d6dddf6adb · inbound

Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images cites this paper.

Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:16:14.393486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T10:13:08.112072Z digest=sha256:beff33b504b3160a3a55d3fb1cf847260a0b35d81d8def7be29d6ad6ddf289e5

Observation 0c874784-5226-47e3-b452-a4842a504355 · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:33.489899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:33.489899Z digest=sha256:ccffd5ec4102cdb68c293950705a665a72426d452290ac9ca9d742784cd30678

Observation 8edea675-28c1-49e9-8a8c-c67c7668c18f · inbound

Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention cites this paper.

Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T17:15:17.698711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:15:17.698711Z digest=sha256:401aa7966ab742acd87d587dc74307a95e2e6af58529a42b3eb544ff4b4cc2ce

Observation 4d3f4706-9dff-476e-8931-6d15122eac37 · inbound

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning cites this paper.

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:11:22.875299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T05:07:39.571227Z digest=sha256:90a5d6522b6d0fbc627df31607acbf0f7871f13e9976960095007e54b99099aa

Observation ab30d83a-a8e2-433e-87cb-8c060130f687 · inbound

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning cites this paper.

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:47:31.637135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T07:43:50.633779Z digest=sha256:f1c4e925a8eb156f6a68a38b216456a22b17042e826cd3c5f349e73745dda587

Observation 4a50d179-76ff-46a8-bce4-4ffcff692ba5 · inbound

CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning cites this paper.

CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:13:44.773975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T21:10:56.533649Z digest=sha256:3c2253cec1f2f5316841e56328bb7bc32defe2147e23ced579e3854c6f57fe38

Observation 05970f97-2491-4dfa-be92-e26319b89bde · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.527660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:e880211bf8a1eb804caf4566b2e98dd1476a767fdc3e629d2666cf385f26d0ad

Observation c014bab5-6996-423b-b930-9bf9c6509f2e · inbound

Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models cites this paper.

Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.063163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T05:37:23.299943Z digest=sha256:8feb0b73e5b017c222971d0f6ab09177536d269b0fc0756e8b1c2d5324456103

Observation b4dbc6f8-3b17-445e-92ae-457a0d065b00 · inbound

ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs cites this paper.

ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:23:59.845931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T22:23:38.178195Z digest=sha256:678bc5ceab69955ed529db7a240e445fb3a4aa28bd32a9761217b836880a5617

Observation 85f17c59-249d-408c-acc2-0d3fe13fe54e · inbound

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning cites this paper.

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.642947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T13:41:44.049230Z digest=sha256:b0f378460df92ee143dab2e84b3fe85c0de83917c6374af7364221e4b879a241

Observation c3d27d2b-00ff-46ed-bcdf-653d8dd7df31 · inbound

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning cites this paper.

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:43:29.020880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T13:38:01.819121Z digest=sha256:92fe67ab1bd1985c604a7f789daa87427a71700a2e320c786080065fdd3cbaa2

Observation 65dd6ab2-5237-4ff1-9f72-0d32f2290ac7 · inbound

Grounded 3D-Aware Spatial Vision-Language Modeling cites this paper.

Grounded 3D-Aware Spatial Vision-Language Modeling Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:13:15.379490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T08:08:36.012761Z digest=sha256:b51dd594fc3f68b9706bcd25ab0151f86dab73a46accf76ca0ca463339634cd4

Observation d9c0e8e4-f5d4-4704-9801-13007a9e5f77 · inbound

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning cites this paper.

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:07.212605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-25T21:23:10.051805Z digest=sha256:038b5b0a52290109a86b134078f9b2555a388eedb61627223242ba1e966cfcd8

Observation d07fa59e-6f8d-4f08-aef1-664a84ff5b83 · inbound

Visual Access Boundaries in Vision-Language Model Reasoning cites this paper.

Visual Access Boundaries in Vision-Language Model Reasoning Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T06:24:28.750079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:24:28.750079Z digest=sha256:3205f871081121ccac8677be36f34616bb945077cfbeb5a10517406bab341b29

Observation 5c1400db-a6ca-4bb0-afa2-ed604e0e375d · inbound

OPLD: On-Policy Latent Distillation for Multimodal Reasoning cites this paper.

OPLD: On-Policy Latent Distillation for Multimodal Reasoning Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T16:13:17.383178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T16:13:17.383178Z digest=sha256:f23b4f60aafb1d895126c83ae359fd59d8b68a5541074409b87c551e1d49c08e

Observation 23fadc11-0a1a-41f7-9d4d-9088468e9795 · inbound

Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs cites this paper.

Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs Grounded Chain-of-Thought for Multimodal Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:11.202535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:11.202535Z digest=sha256:921144ea663871455c039b749d9e501185ba45af91b5841c87e253876da0566f