Pith. sign in

Paper Citation Record · LEDGER

Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2410.21220.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.21220 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:57:54.119975Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T15:27:04.356488Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 366a2314-cb85-48de-a10c-ec5e863ee7ff · inbound

Agent-based Condition Monitoring Assistance with Multimodal Industrial Database Retrieval Augmented Generation cites this paper.

Agent-based Condition Monitoring Assistance with Multimodal Industrial Database Retrieval Augmented Generation Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:54.119975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:57:54.119975Z digest=sha256:b382a817e2b0c4162cacfede241c4630c5a68eb973f9da78c22745aa6445e350

Observation fdd8921a-fe0a-40e5-b704-e73f448640eb · inbound

MMSearch-R1: Incentivizing LMMs to Search cites this paper.

MMSearch-R1: Incentivizing LMMs to Search Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:27:04.359334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:27:04.228144Z digest=sha256:7f3ea54dca6251b7580b027bb79e805701155e49acdf20b9f9265e3702f9aed2

Observation fdafa1c5-c530-4513-85cb-d8b059f3caab · inbound

Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models cites this paper.

Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T16:15:28.012224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:15:28.012224Z digest=sha256:3163cdd6b4df96482d8a35c8450df6890803179fe85ed9ec1a0c98035f298331

Observation 87c3752c-9c26-4dbd-a70f-c555aeab7736 · inbound

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents cites this paper.

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:07.876274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T03:21:30.732925Z digest=sha256:167d3591b9bd2b224c500d1397b53a52d08d10f79d4e86872dd599dcfdd8f3af

Observation eb930070-8b28-49f4-8e0b-cc3f4616b474 · inbound

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards cites this paper.

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:10.100388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T01:12:17.469552Z digest=sha256:9d72eaa416dea29c4e43f01259b331ed13d1c01aff5958d4892224649b50e84a

Observation 69f8f00b-ff32-477a-a004-e9721ce81336 · inbound

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG cites this paper.

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-02T10:20:56.272120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:20:56.272120Z digest=sha256:a8cb61fa545338ecc76f8bd3e2cdb6e1ac2ef8e7831720d2ba60b89b4107e0a0