Pith. sign in

Paper Citation Record · LEDGER

Vision-Language Models for Vision Tasks: A Survey

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2304.00685.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.00685 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:54:57.423394Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:59:53.217043Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c73c9ba7-c9ea-4e43-9e37-a0a8a497c1a9 · inbound

Humanoid World Models: Open World Foundation Models for Humanoid Robotics cites this paper.

Humanoid World Models: Open World Foundation Models for Humanoid Robotics Vision-Language Models for Vision Tasks: A Survey

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:57.423394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:57.423394Z digest=sha256:de94af6f8cd1359a114910dd5709fc44ca4d98edae65a3608776be7be261c49d

Observation ff4d8d66-ad65-4f5b-af4d-afbdd0c8c57c · inbound

Mobile GUI Agents under Real-world Threats: Are We There Yet? cites this paper.

Mobile GUI Agents under Real-world Threats: Are We There Yet? Vision-Language Models for Vision Tasks: A Survey

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:02:07.906593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T06:57:25.257775Z digest=sha256:be6c8f72cdcc10add30ba77c01f2592fc1f23931fe199fdafc97e8a2dc48c8d9

Observation 3db49778-a5be-4ef0-a1c8-f7351c95ef90 · inbound

BlueGlass: A Framework for Composite AI Safety cites this paper.

BlueGlass: A Framework for Composite AI Safety Vision-Language Models for Vision Tasks: A Survey

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:22.192112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:46:22.192112Z digest=sha256:60bd91600b05955204ffded67309c1bc0e5ec314328169ee7159c1899f45023c

Observation 1e779dc7-e67d-466b-a7b0-912ccd56bd2c · inbound

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems cites this paper.

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems Vision-Language Models for Vision Tasks: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:43:10.857101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:43:10.857101Z digest=sha256:b5bfc6d5c8a3155ea03291ec4e85f813fa7da528af593f3c3dc68fd42358fd65

Observation fc1f0fa0-6955-4ad4-8235-c919c200d8d5 · inbound

Re:Verse -- Can Your VLM Read a Manga? cites this paper.

Re:Verse -- Can Your VLM Read a Manga? Vision-Language Models for Vision Tasks: A Survey

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:16:54.116669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T23:14:25.088197Z digest=sha256:44a2f29b459d680c681c6865b6a58ba8118c7023906f04fdb555cecac960d39c

Observation 29a3e9b4-a2d3-46c1-80dc-c3f24e36a74e · inbound

From Image Captioning to Visual Storytelling cites this paper.

From Image Captioning to Visual Storytelling Vision-Language Models for Vision Tasks: A Survey

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T10:27:39.491216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:27:39.491216Z digest=sha256:5a209242b79f2542907e29aa66359a814d209752a305ed188f9d9d4a067ad9dc

Observation 7bcb7bce-92b6-4178-9dc3-1075095b3b0c · inbound

When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective cites this paper.

When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective Vision-Language Models for Vision Tasks: A Survey

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T16:33:29.108092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:33:29.108092Z digest=sha256:74029dd0d3ec165a4d1ee06ce2f64a6d87801eea53a0e7e489c625450a6c4038

Observation 06b7ccd8-d487-4846-ac50-b8239a562185 · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning Vision-Language Models for Vision Tasks: A Survey

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:42.375221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:42.375221Z digest=sha256:0f3a903543261c6c3dad07d960d4e993487e6f5fb97de6f08937b44751e735d0

Observation a59938c4-8385-45a5-abbc-10bcb40fce8f · inbound

Jointly Learning Predicates and Actions Enables Zero-Shot Skill Composition cites this paper.

Jointly Learning Predicates and Actions Enables Zero-Shot Skill Composition Vision-Language Models for Vision Tasks: A Survey

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:13:58.585221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T05:10:49.673130Z digest=sha256:670581a8ee024fda9c2df43724f227a5a8c23a092afcadc159c0455ea37eabf6

Observation 1d9724cc-0f04-4d66-a058-affd20e900ef · inbound

Robust Onion: Peeling Open Vocab Object Detectors Under Noise cites this paper.

Robust Onion: Peeling Open Vocab Object Detectors Under Noise Vision-Language Models for Vision Tasks: A Survey

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:59:53.218477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:33:25.373918Z digest=sha256:6355c319cb84b9822fe67e25d33d3151cc91a94ca3032c042f1e19974c0d6dc5

Observation 1b779d3c-1858-4f7b-82bd-9cb8f7a7c7cb · inbound

Robust Onion: Peeling Open Vocab Object Detectors Under Noise cites this paper.

Robust Onion: Peeling Open Vocab Object Detectors Under Noise Vision-Language Models for Vision Tasks: A Survey

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:15:50.080203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T00:42:36.605869Z digest=sha256:e002f8ced3ebb0a8ea9da8c1775b58f890b16d547642f8454adadf63f6421006

Observation 7ef60ee1-85ce-467f-a5df-44e7adf32d77 · inbound

Autonomous VR-Based Risk Detection for Situational Awareness in Dangerous Settings cites this paper.

Autonomous VR-Based Risk Detection for Situational Awareness in Dangerous Settings Vision-Language Models for Vision Tasks: A Survey

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T20:36:26.660978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:36:26.660978Z digest=sha256:7357e281ad3200a7b1f614613709daa7a39d1a9105fd1f423f5c1ae893a04e8c

Observation d85a81fc-c181-46ff-9174-878c230ffae3 · inbound

One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models cites this paper.

One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models Vision-Language Models for Vision Tasks: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T13:26:30.037563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:26:30.037563Z digest=sha256:09a4102cd9dc4518c81639a26a0f1bdc05597a7d5ddfb4f474701581f9faf3a8