Pith. sign in

Paper Citation Record · LEDGER

V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2312.14135.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.14135 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:44:24.179591Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c4752a49-937d-4f37-8801-70f4bc348d93 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 208

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.239118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:296a44d1dab7e92d8a645c13e06145a29c8c5154aa1322958b02a628366991b8

Observation 5d2ab2cf-19a8-45c1-8e0d-6a1e5c44ce3f · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.236393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:08707d2a85aef15c24df3a2ba83d0f6dc40130aea7a0108ebc0b5ea5eeabc4ae

Observation 49c95fd6-1bec-413f-b9c5-b35a268e6263 · inbound

Improved GUI Grounding via Iterative Narrowing cites this paper.

Improved GUI Grounding via Iterative Narrowing V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T18:44:24.179591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:44:24.179591Z digest=sha256:9c6de630db50fdba60e4906363785a99202bacb916b5043852add819942d1e44

Observation 1d2ef04f-9865-412b-9b35-1307004e607b · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.692676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:bb8f14ee93833adeb6f2dd96145c85e28a4e930564fc8c2745a3a1f977e0c0d8

Observation 5684eb9b-6dd6-4071-a5a2-23d709b86867 · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:05:52.086561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:69dac90296407a42cf3c841f98909f20f0d44f1122819cd7a89a7d33f9b9a11a

Observation 71f3e515-41de-4d7a-b0bf-ad75ebc445e9 · inbound

Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research cites this paper.

Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:19.034618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:19.034618Z digest=sha256:478118fb0f95a14030b06ace6330f72b554f28fa84bdc6bc072c6b545e163f88

Observation 7599ee8d-24d9-4414-ba22-ad3aa2d74bf5 · inbound

MiMo-VL Technical Report cites this paper.

MiMo-VL Technical Report V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.241662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.241662Z digest=sha256:324d856e6a40c058943a184c572f4b1222033498284233b11b861deba56c593f

Observation 858d4b93-a6c5-4dbe-9a08-d68c203ae2d6 · inbound

Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations cites this paper.

Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:23.231324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:23.231324Z digest=sha256:e49f98e319864e7d0ac380be8ec454e9b875f8b197ea5d50a318a7705640a3eb

Observation e3c446d9-6703-4203-825c-4baea7897aba · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.885634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.885634Z digest=sha256:63f47cbfb886ddaa20186277b4c4aadda335142b418e54d34dea4bcd1e4e0772

Observation e64e0d55-fe91-403e-8dc5-24610c1afd0b · inbound

Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning cites this paper.

Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:36:25.246023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-18T13:33:12.508639Z digest=sha256:90d9225e36bceef30f73f9465cc9127131ffc43acdd036cd13ad05e5e60e7cd6

Observation a052efc3-2ef7-4196-8175-27680ac2b9d1 · inbound

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling cites this paper.

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:14:22.807207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T21:11:47.804851Z digest=sha256:7b8483e15714c49456ed380ff85b2ac7f5a3968782358808be6fd7a26995f22b

Observation e7edd0ef-5bc5-42cd-8a8d-baaa308ee84a · inbound

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling cites this paper.

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:58.310244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:42:58.310244Z digest=sha256:943381c79769bdd27cf2f0bd93549866c4ec72570ca25c3c6377a7cf4981d6e6

Observation 0bead9c2-52c2-44ea-9125-0979bab87762 · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:57.799831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:57.799831Z digest=sha256:8dc3fe9ffe12a767ab3803937cd32fd4e75aa49563c75eddc1194ff0f8b047b4

Observation 978dc477-32ba-4afb-8e4c-f600cbfa318b · inbound

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding cites this paper.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.165356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.165356Z digest=sha256:78888dd954f5f86c4973110e023649435e0696ed3e33aa4644966008067effec

Observation 92a875fa-dc60-4fdb-91e4-ed1ceac05f94 · inbound

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension cites this paper.

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T03:30:33.046213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T03:27:53.694506Z digest=sha256:3c89968919fde8a62a2a24f67fd0e069ac525691a3b8919a05be3e1f17f83d53

Observation 1c20257b-e448-47cf-bad4-28b558dd828a · inbound

DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model cites this paper.

DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:05:04.972295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T12:04:08.487678Z digest=sha256:a0ab3d0d4fd642740407283d0031c907e8435c676e4ea2573a0ebf225fe431d9

Observation 1e1dfc9c-74da-4b71-a0dc-5347b54124fa · inbound

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning cites this paper.

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T19:34:58.789459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T19:34:58.789459Z digest=sha256:afde0ad01c8ac92add2671a3b4fc47243d28c6be2172704c125ca57ed558ed1e

Observation cc98f7b9-99a6-49ed-a866-90b9366f1cf1 · inbound

LanteRn: Latent Visual Structured Reasoning cites this paper.

LanteRn: Latent Visual Structured Reasoning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T17:26:40.239147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:26:40.239147Z digest=sha256:8df124d4a38ee565ef20ae81755116c54aa7295e7f5921b0dee0ba76a77c01bc

Observation 40ed16f7-83f2-455c-a1fc-4d451d3d4b24 · inbound

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models cites this paper.

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:51.617866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T18:46:26.869644Z digest=sha256:af923af8d040d821e00b552970fecd5aaee7b41455233f95f1b402471044750c

Observation 7eb14478-c1d8-4d8c-9be8-ebd448c8b100 · inbound

Multimodal Latent Reasoning via Predictive Embeddings cites this paper.

Multimodal Latent Reasoning via Predictive Embeddings V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:46:52.537656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T17:26:27.683269Z digest=sha256:32b286c1112294fe60a1a64d15d8f9b6952398ad201772037556c8c39e5d31bf

Observation d99eb290-4d03-4448-8e02-44676820aed7 · inbound

Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models cites this paper.

Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:10.605697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T17:14:11.941977Z digest=sha256:fb76f395d7ca5775072361e131be55cd26470559e0547fd9d8deff99d5bb1a9c

Observation 83b8a0dd-9add-4223-8527-8a5659dfeeff · inbound

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding cites this paper.

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:51:18.794113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T17:25:31.097385Z digest=sha256:8e90cfcf3227b8bac69a8e33eb2bb55a4dcf1ed8fe4bc241ba6b530fa55a87ee

Observation d3c74dd9-a762-4006-bd93-04f769f35297 · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:25:58.822046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:03:15.222571Z digest=sha256:6ea6ed1e32891cb908dc91c8eb75f429fecaca13f1afe2c193d681eb86aa6142

Observation 983d9ea9-53f7-41d6-abae-b8d9662b0c16 · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-12T22:48:45.647588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:48:45.647588Z digest=sha256:94b44cb90cf4d8e09608afdbfc520c140d383273e68c78025620e604baa4db05

Observation 814f6ec7-7150-4d8d-9bf2-18e47e7f8447 · inbound

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation cites this paper.

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 213

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:43:49.502913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T05:10:44.608959Z digest=sha256:8fd89da9a2dd8ea1b1651e383a645178f01facf4404608057fce399e3fbc6b5b

Observation ea694479-978c-4604-9b46-9c67c71b600d · inbound

SketchVLM: Vision language models can annotate images to explain thoughts and guide users cites this paper.

SketchVLM: Vision language models can annotate images to explain thoughts and guide users V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-09T21:33:27.942692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-09T21:32:08.503584Z digest=sha256:42448bf86fb94735b1321ee04267622e537f8f02439bdbb559bccc3a70bd1251

Observation d44ab193-c2c8-4f94-9ecf-b6b415bbd7b6 · inbound

Improving Vision-language Models with Perception-centric Process Reward Models cites this paper.

Improving Vision-language Models with Perception-centric Process Reward Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.052295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T04:33:36.634359Z digest=sha256:c7373d83282cabf304e1b1f224cadc8a9fcb19062ee7c1bf0252c9c6934f94e6

Observation 4d1a14b4-c409-4d27-8502-8d53df71a759 · inbound

GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning cites this paper.

GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:54.650794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T02:03:52.566413Z digest=sha256:371d42d53bf8eb852e2dab0761962bd65fe97ad2f7e7b227dc116c18c9dba81e

Observation ccad97df-3074-475e-b73f-7de3b8ff1d3e · inbound

What's Holding Back Latent Visual Reasoning? cites this paper.

What's Holding Back Latent Visual Reasoning? V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:03:15.473753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T11:59:14.134917Z digest=sha256:f037b7934430ce9570c191aed8a69678a24401212d4a94ffe31b1e6539d6397c

Observation e89677b3-b613-488c-9467-a4394c408300 · inbound

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth cites this paper.

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:53:13.123588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T10:52:32.238241Z digest=sha256:9d26d19b956b05fe1caa3253dfea6ebb3c35c7ce637da128f114b3db997915e0

Observation a867bb2b-1d00-4b86-9775-0b198c98424d · inbound

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth cites this paper.

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T16:34:34.452430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:34:34.452430Z digest=sha256:0c20e3ada2e0be4100a9d7f1d295f90f2924e70bfa75aa73e21613601a8606ca

Observation d5408f8a-eebe-4ad9-a88a-24da1b98dd3f · inbound

Self-Prophetic Decoding to Unlock Visual Search in LVLMs cites this paper.

Self-Prophetic Decoding to Unlock Visual Search in LVLMs V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:03:26.884105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T12:53:40.783281Z digest=sha256:2f6ef409638b00514926f25344eec3519312c1f48b623c3c27f59fe2f7a70a2d

Observation 22009e46-56ed-46c6-ab6a-b1cf6c9a5d9a · inbound

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents cites this paper.

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:28.939546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-28T10:38:41.735204Z digest=sha256:0a9290d85b5abebadcd309f22638158ece289c04ea5e16135acf3fd6da2ad5aa

Observation 79c455c0-6e7c-4123-848c-947c636d6fd2 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.951269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:ea88df4bf09488c2a9415b4e2810d998bee84837219a3f9c7d2930ad2cf6b5b6

Observation 02a06411-4ac5-46f5-9891-ef32db2a3b91 · inbound

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning cites this paper.

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:49:02.271066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-26T21:47:51.437284Z digest=sha256:05885c50c7b575ec93a6211db55691b62e7e9401c6cf8770d6bed61a103679b0

Observation 8497a1cf-e8ba-4d0d-a171-6b6a1eee502d · inbound

Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG cites this paper.

Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:49:38.554498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-26T12:54:06.319322Z digest=sha256:0509502b0a501bcd0cbf267cf5a021f42f4e68bd9ee043dc5d420848cf59038d

Observation 3b1d8caa-5deb-486b-8c29-6611a5ad14b9 · inbound

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse cites this paper.

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:09:48.891429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-26T07:13:19.593093Z digest=sha256:c5b12d09f6fc82d339a03f1b9abcf0c78088b96d2bb0f93c4475ee43c347c547

Observation e17be238-b3bf-460a-8f63-3713f499ddcd · inbound

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding cites this paper.

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T16:48:36.519393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:48:36.519393Z digest=sha256:29b390900ece530b86e05745f8a1453c2c80a317613db023443f53abc1249e7b

Observation 54c1de77-47be-4752-be6a-b354061bc3c3 · inbound

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation cites this paper.

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-31T02:56:58.290582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T02:56:58.290582Z digest=sha256:bf2619da6508ef3e1a194c8de087005bcbac0d1661b82fa45b6d8848186bbf9d

Observation 3c555d85-f35f-4299-851a-8fcf9e14a703 · inbound

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs cites this paper.

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T02:46:49.732380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:46:49.732380Z digest=sha256:dc03a6bd871c3317c7148f3455af5ab1d5eaa0e204ad5242fc956e6c780883e0

Observation dada7984-74ec-4417-be64-008b1daadaac · inbound

AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models cites this paper.

AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:20.618767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:17:20.618767Z digest=sha256:ea77bbf29eb87c6e048cdea121168efb17d492199e10f46bde7a12baa5e52a71

Observation c1358137-74bd-4884-b10f-9687c7390116 · inbound

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning cites this paper.

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:14.372654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:14.372654Z digest=sha256:3cde27fc7ba27e626710bc1bbe078aa1c232b43107466d9abe19f54809ac2e96