Pith. sign in

Paper Citation Record · LEDGER

Visual Prompting in Multimodal Large Language Models: A Survey

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2409.15310.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.15310 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:39:00.033360Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T16:03:08.102407Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 35ccaa1c-edcb-4a3e-9b49-be7bb19f70a5 · inbound

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey cites this paper.

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey Visual Prompting in Multimodal Large Language Models: A Survey

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:00.033360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:39:00.033360Z digest=sha256:9ab604e1f0cf817ad9b4ce32382f6c888469d84bac7bbc161f1bb0df78a595de

Observation a6cf2500-24d8-4980-90e6-625bbe754fe9 · inbound

Incorporating Token Usage into Prompting Strategy Evaluation cites this paper.

Incorporating Token Usage into Prompting Strategy Evaluation Visual Prompting in Multimodal Large Language Models: A Survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:01.080099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:34:01.080099Z digest=sha256:cd87e91148d7ca4c7c502e366764cb132dcf5bd1c6d7407f99939d5fab997a57

Observation 22a4348b-32aa-45d7-840f-981ad9e234c0 · inbound

MLLMs are Deeply Affected by Modality Bias cites this paper.

MLLMs are Deeply Affected by Modality Bias Visual Prompting in Multimodal Large Language Models: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:22.096111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:22.096111Z digest=sha256:b25b81d0df350c481d09ae31ac8a1a1962c9bb2fe341b47979293094e6757847

Observation 0ac4b01f-4cd7-4810-8161-76edb9375d4f · inbound

LPOI: Listwise Preference Optimization for Vision Language Models cites this paper.

LPOI: Listwise Preference Optimization for Vision Language Models Visual Prompting in Multimodal Large Language Models: A Survey

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:57.523050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:57.523050Z digest=sha256:dac00331c427e454fd49a492962ae552c8f14fc145cbe735b6ba807e107db8a2

Observation a40d6278-33db-482d-a494-cb1094d91504 · inbound

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models cites this paper.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models Visual Prompting in Multimodal Large Language Models: A Survey

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:41.257873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:41.257873Z digest=sha256:401832d86f25b9c2487b1c217df2e1d088730be32db4a128b84f7e14254f1d5a

Observation 7aaf53df-d37c-4d31-b5c1-7996665ae2f2 · inbound

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models cites this paper.

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models Visual Prompting in Multimodal Large Language Models: A Survey

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:15.716207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:36:15.716207Z digest=sha256:285e9ecfa5e64e179be1eec49ddb4c2bb58a71546f0530fdf689ec6e3a487ba1

Observation ee4b47a9-2a6b-4935-b17b-3a2656142bf0 · inbound

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs cites this paper.

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs Visual Prompting in Multimodal Large Language Models: A Survey

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:47.733190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:47.733190Z digest=sha256:4a9cacf4a9ac0176a5f59ed78677b79d67db55b1561ebcbd5b365eb8e0b471e2

Observation 8f1ad2cc-cb70-4d95-a6eb-141ee985027d · inbound

PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior cites this paper.

PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior Visual Prompting in Multimodal Large Language Models: A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:55.827401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:55.827401Z digest=sha256:0f69230dce48ed80d93391339265e47c6796e9a2221166c48b645868a89bd3dc

Observation d7e906a8-9839-4d82-9328-c67c32ea5419 · inbound

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? cites this paper.

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? Visual Prompting in Multimodal Large Language Models: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T23:35:24.104443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:35:24.104443Z digest=sha256:70b17158f3a981e22f3cc2feae8e44543f67f29cb499184483f9fa9e57185f90

Observation 3e86d794-ff2a-413c-b997-c1cb4c560696 · inbound

Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design cites this paper.

Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design Visual Prompting in Multimodal Large Language Models: A Survey

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T21:01:26.976295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:01:26.976295Z digest=sha256:380843c002145918408c4285f1cd1c8e140be8d6d2b15e66ad6a5828ef8b37cd

Observation 45f422a6-bf94-4aba-a52a-2f60719e77aa · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Visual Prompting in Multimodal Large Language Models: A Survey

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.850843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.850843Z digest=sha256:2b5cce1a376c523c9077361461e43505c3dc94b0f5e2337a9d568edd970014e1

Observation be421962-57d0-4d6b-a57b-c015f11c8160 · inbound

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs cites this paper.

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs Visual Prompting in Multimodal Large Language Models: A Survey

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:35:57.458742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:04:05.157103Z digest=sha256:7ac9a434543b2bbbbb749fff1bd91a185ec4725380a03708f543cdb83b8624d1

Observation 6005aaec-6a6e-487f-9189-3de52d2a65b5 · inbound

STORM: End-to-End Referring Multi-Object Tracking in Videos cites this paper.

STORM: End-to-End Referring Multi-Object Tracking in Videos Visual Prompting in Multimodal Large Language Models: A Survey

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:00.442212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:25:31.777907Z digest=sha256:3b2de4b2ae91d649d455c33ec4aaeca91e641edbe67650ae2b0b44ac96beaf7c

Observation 32c52558-2f0c-44c1-9586-f9c50a23884d · inbound

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification cites this paper.

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification Visual Prompting in Multimodal Large Language Models: A Survey

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:56:47.420763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T06:55:33.464951Z digest=sha256:7002c991350e4af3b47f7ec8cda16b96365895a49b0d4c8c3635930ee22c4f29

Observation 3f808a75-1b6c-497b-ab1a-dbc0d9006b24 · inbound

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification cites this paper.

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification Visual Prompting in Multimodal Large Language Models: A Survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T19:10:41.883296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:10:41.883296Z digest=sha256:129aa69c4fefe99699c694c1625632000da0003ea1a68f645bf6b3b71fcadf78

Observation c2e6a976-bb1d-443f-89b0-e8f8e9a0a29f · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Visual Prompting in Multimodal Large Language Models: A Survey

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.439868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:cc448bc9c3b7a42957c1c42e7de987e14e87761c6b832480005135e534e131aa

Observation 5a3cfb1b-a25e-4961-b5b8-646f10e721cf · inbound

Skill-CMIB: Multimodal Agent Skill for Consistent Action via Conditional Multimodal Information Bottleneck cites this paper.

Skill-CMIB: Multimodal Agent Skill for Consistent Action via Conditional Multimodal Information Bottleneck Visual Prompting in Multimodal Large Language Models: A Survey

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:01:28.145905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T01:24:59.721212Z digest=sha256:62d31c4d87304250ec9577e15ce7ba5adfde06fb561c11c76fcebce7f0d5ab77

Observation 1a7e3a96-1854-41d2-bc00-f77617cb93cc · inbound

OLIVIA: Online Learning via Inference-time Action Adaptation for Decision Making in LLM ReAct Agents cites this paper.

OLIVIA: Online Learning via Inference-time Action Adaptation for Decision Making in LLM ReAct Agents Visual Prompting in Multimodal Large Language Models: A Survey

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:27:07.225834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T02:24:42.530103Z digest=sha256:2d61466c7d99a34613490e9df87e702aa6b8752e78415179f1c0ad8f87dea9fd

Observation 9d51bcf7-f6d3-43cc-9d9d-296b3341e32c · inbound

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding cites this paper.

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding Visual Prompting in Multimodal Large Language Models: A Survey

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:03:08.104159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:02:53.887605Z digest=sha256:74d66ca9a34abf499174b5015729ed13171554e5fb792be7f21854c35f1cf1d0

Observation e0630495-5989-46a3-9cb4-bb265b6a0e85 · inbound

Plover: Steering GUI Agents through Plan-Centric Interaction cites this paper.

Plover: Steering GUI Agents through Plan-Centric Interaction Visual Prompting in Multimodal Large Language Models: A Survey

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T23:57:13.850939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:57:13.850939Z digest=sha256:ce388f20edec18832415aedcef0a8edc32a0704480bd58a6a70c1f05495ac236

Observation c9cbef5a-a706-4849-baec-6e32fb03b999 · inbound

Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs cites this paper.

Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs Visual Prompting in Multimodal Large Language Models: A Survey

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T06:18:16.279690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:18:16.279690Z digest=sha256:a80dd1385e50bc8d1ea3ce3f0ba62a972880cc36eb971075c087413dace9d11d