Pith. sign in

Paper Citation Record · LEDGER

ClipCap: CLIP Prefix for Image Captioning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 53 inbound Pith citation observations for arXiv:2111.09734.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.09734 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 53 of 53 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T15:04:39.924340Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:09:44.426723Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9e39cb2a-dfe9-4514-bac3-8f235c4f40e0 · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language ClipCap: CLIP Prefix for Image Captioning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.718168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:a2a3af10ca723135786df156c5b9c6be358d3cec26ad3b1fd5467096849ef704

Observation 8cfeab8e-d3b3-4a90-9f76-16506dde437b · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning ClipCap: CLIP Prefix for Image Captioning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.476566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:975fd74e029c45cfce69c08adb0b0925a7bfa94b2712c21f221bbdd4861ac8a6

Observation 7f11e737-20da-46c7-9a1c-c8852ac93363 · inbound

LAION-5B: An open large-scale dataset for training next generation image-text models cites this paper.

LAION-5B: An open large-scale dataset for training next generation image-text models ClipCap: CLIP Prefix for Image Captioning

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:22:17.333000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T14:22:16.968028Z digest=sha256:0056200230683c7fc285c8f7e49640888c2b97f1e9a07fea436a571ecc49035c

Observation 1ef5c014-1562-48d7-b252-38993a4e3396 · inbound

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention cites this paper.

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention ClipCap: CLIP Prefix for Image Captioning

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:07:42.528525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T23:07:42.245641Z digest=sha256:80092180fd1891ffe0c0db3fb0916536b566d89bebfdb2efe07e00277b76e03f

Observation f7f1d37a-c0ba-42f8-a149-1b8b39a8f713 · inbound

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model cites this paper.

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model ClipCap: CLIP Prefix for Image Captioning

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T08:41:04.887711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T08:41:04.743886Z digest=sha256:d73faa79d3642fd8d193592230ffc0e298745fbcab4bca22cd3ad62b29b488bd

Observation c9e2d1eb-3be7-4be7-96fe-9b7e798c957e · inbound

EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers cites this paper.

EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers ClipCap: CLIP Prefix for Image Captioning

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:11:49.649827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T06:11:49.475825Z digest=sha256:c6b2a8871e06c9a0bce1c07e365e794a212b868908f6adbd6e4ba5342c800741

Observation e0b950ca-9483-42a5-b8e3-d2a945e9fbc9 · inbound

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective cites this paper.

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective ClipCap: CLIP Prefix for Image Captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:39.924340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:39.924340Z digest=sha256:dd18d8da1ae3f56cf218eb456cab51bb2f71bfe30baca9eb6bd29cee1dbfef0a

Observation 5236dadf-4852-4991-8817-d122aa7d8f9b · inbound

ADIFF: Explaining audio difference using natural language cites this paper.

ADIFF: Explaining audio difference using natural language ClipCap: CLIP Prefix for Image Captioning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T22:40:26.071258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:40:26.071258Z digest=sha256:24589da6ef6c6a9323c91449743ba53183c65a6fbbdc3a4da7dde4fc91e532d9

Observation 7656f08c-4e8f-485a-a9d6-9c2b5ed4700a · inbound

Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search cites this paper.

Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search ClipCap: CLIP Prefix for Image Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T17:27:25.774500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:27:25.774500Z digest=sha256:d4a4fb88cc276354fbab6fcf5ecaad8289b9ce89a49a439da03e28d43bf8c4a3

Observation 6624c8e7-75f7-4c57-bb29-6bfab70825db · inbound

SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps cites this paper.

SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps ClipCap: CLIP Prefix for Image Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:16:36.087824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:16:36.087824Z digest=sha256:acffda1bf90dca2a19172ea22488f2ce7ff64186d59f407fe53a8c3dec9c75e3

Observation b33833fc-4b2a-4ccf-82ac-1e8cdf30cc3c · inbound

Enhanced Multimodal Aspect-Based Sentiment Analysis by LLM-Generated Rationales cites this paper.

Enhanced Multimodal Aspect-Based Sentiment Analysis by LLM-Generated Rationales ClipCap: CLIP Prefix for Image Captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:18.994193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:18.994193Z digest=sha256:feac597ada9e75d4561f9a56285dd0000531413e3f5175a0f3c8a85c315bbc7d

Observation c22dbbde-1000-446b-a047-623d6876de8c · inbound

Optimizing fMRI Data Acquisition for Decoding Natural Speech with Limited Participants cites this paper.

Optimizing fMRI Data Acquisition for Decoding Natural Speech with Limited Participants ClipCap: CLIP Prefix for Image Captioning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:19.824796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:40:19.824796Z digest=sha256:1031772f94562ece95a4a9e3722c5acd815d52809efa6197643060d11b56d8f4

Observation ab95742f-af91-4837-a8be-418c7492f131 · inbound

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model cites this paper.

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model ClipCap: CLIP Prefix for Image Captioning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:44.090056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:44.090056Z digest=sha256:5f23caaba1af924e0e16e1b3742a389556e36919aacda0b6e945fbac112a3261

Observation 22bc1e94-042a-402c-aa38-1f117cddb53e · inbound

Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models cites this paper.

Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models ClipCap: CLIP Prefix for Image Captioning

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.351321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:35:36.351321Z digest=sha256:a2f23986ad0f1aa8987f07cbd7da917ec37b7c1f79faf5e72779f903b8dcc7d7

Observation 01ceb64c-abd4-4229-8c0d-a17cd390032c · inbound

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering cites this paper.

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering ClipCap: CLIP Prefix for Image Captioning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:10.602373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:10.602373Z digest=sha256:b649faf4ff0dcfd2bbc433fbdaf0002e06098a7e3ac1bdf13e6df238bc9ff426

Observation fe79f438-6836-4404-9e98-d2e7b2503694 · inbound

Diffusion-based Cumulative Adversarial Purification for Vision Language Models cites this paper.

Diffusion-based Cumulative Adversarial Purification for Vision Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:59:18.927640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:59:18.927640Z digest=sha256:8aed347f717ddb42e8742d6aa948fc8338271828deb83a89f499ac68823d1ffc

Observation f8c28283-b524-4c6a-bf1d-4adbcaf03595 · inbound

CoLMbo: Speaker Language Model for Descriptive Profiling cites this paper.

CoLMbo: Speaker Language Model for Descriptive Profiling ClipCap: CLIP Prefix for Image Captioning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:15.589604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:15.589604Z digest=sha256:3d633b5aff1920b320d6bfb948d9c4143f1509f416483de45ec591dccef2302c

Observation cf1e06e5-7ae9-47e5-bee9-f8ce5f3120ed · inbound

InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning cites this paper.

InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning ClipCap: CLIP Prefix for Image Captioning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:58:07.909306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:58:07.909306Z digest=sha256:fc3eff19895e9acecfc9d73530338cb10554f786bc0729a1f1b5e5fcc8c395ed

Observation 3bb5433a-a261-423b-a2dc-691adfbaa877 · inbound

CLIP-HandID: Vision-Language Model for Hand-Based Person Identification cites this paper.

CLIP-HandID: Vision-Language Model for Hand-Based Person Identification ClipCap: CLIP Prefix for Image Captioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:56:01.612867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:56:01.612867Z digest=sha256:defcedfe3d94f5226a842252d4fe8e3aaaad8ba5757a552bdf08467c7f224e7c

Observation 59700003-ea25-4573-9220-5fa415e5f49c · inbound

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge cites this paper.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge ClipCap: CLIP Prefix for Image Captioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.828926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.828926Z digest=sha256:fe626f5ad010e00def0b3c3e9604b94356460ae6adda6466326df97ce76c8a5b

Observation 46177889-2581-457d-bd97-0e86f0a016ed · inbound

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition cites this paper.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ClipCap: CLIP Prefix for Image Captioning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:55.557921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:55.557921Z digest=sha256:1554996e44604561f190ae4db274bd9a14a2303fca4ae53955e3ff3cde2ef5d5

Observation ca8ccce3-3a66-418b-96d1-4e4cd7f48a02 · inbound

On the rankability of visual embeddings cites this paper.

On the rankability of visual embeddings ClipCap: CLIP Prefix for Image Captioning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.511161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.511161Z digest=sha256:89b05099011a786ca7635d9c3532428f0efc3a68dbbeec842c7ddcfaebce83c5

Observation 5772af9d-23eb-449e-94e4-c3813760953a · inbound

GLAD: Generalizable Tuning for Vision-Language Models cites this paper.

GLAD: Generalizable Tuning for Vision-Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:21.189934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:21.189934Z digest=sha256:cdb7bfd2e6bbc4ff5a0654d1efc0274f823f9b4cc30d6b8a8c57ddfdd1b0e884

Observation bc89fd79-83dc-4729-b9f0-7d938b076d9e · inbound

HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals cites this paper.

HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals ClipCap: CLIP Prefix for Image Captioning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:32:42.735751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:32:42.735751Z digest=sha256:22c54088e6064daccfdf057ff5eb6b1d13b875399c8733ca9b9b5edd16b41a01

Observation 31ab0b8a-1df3-43a3-a031-d3671c0d56f2 · inbound

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings cites this paper.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings ClipCap: CLIP Prefix for Image Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.416252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.416252Z digest=sha256:12a94da8d2e95d704036c7cc7dac0d0fa5bda84ec3c19238b83ddedd6957813f

Observation e640f75c-908b-4092-86a0-39c9f77051fa · inbound

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models cites this paper.

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:51:16.018283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:51:16.018283Z digest=sha256:ed4bb610e82b0a71307b6d2c413953419ccba71bbb4f53e77a08831af207d11e

Observation 6ce12810-cef8-4ad1-84c7-e4a6e8381ad2 · inbound

RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification cites this paper.

RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification ClipCap: CLIP Prefix for Image Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T01:04:10.551418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:04:10.551418Z digest=sha256:c88259abd6445f8e47621a4cdcccb505be575ef7aaee924f249aa80bd6734022

Observation 37b79eda-e591-486e-aa15-4fabec394308 · inbound

From Image Captioning to Visual Storytelling cites this paper.

From Image Captioning to Visual Storytelling ClipCap: CLIP Prefix for Image Captioning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T10:27:39.338214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:27:39.338214Z digest=sha256:e823e0b488fbcf11ca7d635ad934d803248e6b121e01d15f092580ffce32b23e

Observation 572e5683-939d-454a-90ad-ac72b0abbdc5 · inbound

Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval cites this paper.

Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval ClipCap: CLIP Prefix for Image Captioning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T17:25:01.008591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:25:01.008591Z digest=sha256:c5f463b7f16ca268d35e229442b5d8410950ccf8ab1a5b1ca2374d17936a8230

Observation 5d63d557-c52c-4e6c-9495-328a74914cf5 · inbound

Sample-efficient Integration of New Modalities into Large Language Models cites this paper.

Sample-efficient Integration of New Modalities into Large Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.509894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.509894Z digest=sha256:963ca3c1a382d2f8a0858a06ecb4852851b63b5a5899515654d64bedc8a48197

Observation cdb27136-dcf8-43a4-98ed-2f7bc45059c5 · inbound

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models cites this paper.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models ClipCap: CLIP Prefix for Image Captioning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:40.077127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:40.077127Z digest=sha256:a99df999f496ead07e208a8b366d2e4331e08355be30c149998d76a724c2372c

Observation 2b145175-6648-4219-93e1-146d1a146dd3 · inbound

Unpacking Hateful Memes: Presupposed Context and False Claims cites this paper.

Unpacking Hateful Memes: Presupposed Context and False Claims ClipCap: CLIP Prefix for Image Captioning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T10:26:22.120641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:26:22.120641Z digest=sha256:bfbdaceb8f057283dadc802d5b07fea8c746d9ac4694d118eeffb5352f76d99b

Observation b975f0e4-ca19-4fb7-8c9c-0fd08cb67a97 · inbound

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models cites this paper.

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models ClipCap: CLIP Prefix for Image Captioning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T18:25:04.998120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:25:04.998120Z digest=sha256:476cbaf2ae4f02afd906179c54a3fa56476a4768abdcc58b5ae7fd204e73a413

Observation c1124a2e-2683-494b-8f0b-106e84704a6b · inbound

Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain cites this paper.

Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain ClipCap: CLIP Prefix for Image Captioning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:27:59.324632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T14:26:08.188950Z digest=sha256:4415a47c9f79def0dcaf66cf7079e74001bb39e2836e16f9a18b73402e42138d

Observation b0ce17dd-d1ad-4f4c-bd35-81170c9a0278 · inbound

LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs cites this paper.

LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs ClipCap: CLIP Prefix for Image Captioning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T06:06:52.217522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:06:52.217522Z digest=sha256:94d160edf899ec9c852502d9f36a8d555e44cdc61accc5f8b7e8787a3f987dff

Observation 9df69bec-bde3-4b33-8a69-849040417ee2 · inbound

Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks cites this paper.

Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks ClipCap: CLIP Prefix for Image Captioning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:55:52.655740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T20:26:42.128350Z digest=sha256:b1afe210dd27bd8e74f60d7be34a296ab3ccb2db4f5409155753eec287108d2a

Observation 5f07988e-3ad6-4978-80fe-4f081982dca2 · inbound

Beyond Semantics: Uncovering the Physics of Fakes via Universal Physical Descriptors for Cross-Modal Synthetic Detection cites this paper.

Beyond Semantics: Uncovering the Physics of Fakes via Universal Physical Descriptors for Cross-Modal Synthetic Detection ClipCap: CLIP Prefix for Image Captioning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:53.240217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:52:34.752473Z digest=sha256:5973f6628effd1c52d9e1ffab643d84998921dcc236fa8f53c765a0ec4fc60e5

Observation 71069924-22ff-4aa0-9434-e119450329e7 · inbound

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval cites this paper.

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval ClipCap: CLIP Prefix for Image Captioning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:50.060959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:31:53.371412Z digest=sha256:766164e73f2b9a02f37cd3c665e5bf519409965c6fcfe91ef4a14baccd50a760

Observation 7220993c-f54f-4276-a5d2-fa3928c388d9 · inbound

UIPress: Bringing Optical Token Compression to UI-to-Code Generation cites this paper.

UIPress: Bringing Optical Token Compression to UI-to-Code Generation ClipCap: CLIP Prefix for Image Captioning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:00:59.421777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:21:32.024105Z digest=sha256:ee467bc3caa43e5e1e5a38676cb968c5421c9d8cc2f1e08ec490db2bf24e4568

Observation 8baee550-4afd-4f0e-93bb-73d6a391dc69 · inbound

Semantic Manipulation Localization cites this paper.

Semantic Manipulation Localization ClipCap: CLIP Prefix for Image Captioning

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:21:01.497760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:32:30.345594Z digest=sha256:6a67ec982e20d1abd45c51f0246c6c9f2a545db99dd8edbf214860da7b476044

Observation 0f5c2bc7-85a8-4cc2-9b2d-d5129afbf14b · inbound

Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment cites this paper.

Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment ClipCap: CLIP Prefix for Image Captioning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:21:27.865530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:19:31.570783Z digest=sha256:edc51b8916efb61454d594659b550e3beb8f473da931359998b65f8408add926

Observation 6448c78a-b35d-4ece-9770-8503072ef485 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:22:56.290381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:911a490221cf48031cdbda5ebdb4e16f30b7ce4fb28d0567d351d4315dafad59

Observation c8556bc0-b861-478f-a3f6-fd1aa2d5a44e · inbound

CB-SLICE: Concept-Based Interpretable Error Slice Discovery cites this paper.

CB-SLICE: Concept-Based Interpretable Error Slice Discovery ClipCap: CLIP Prefix for Image Captioning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:33:15.862374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:24:05.284777Z digest=sha256:2d14247306280fe9fc49562dced09a7540ee549efdc91af21d64f1d2d22d0840

Observation 7d56f8ed-0f49-49f4-9f85-a227732d0549 · inbound

BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning cites this paper.

BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning ClipCap: CLIP Prefix for Image Captioning

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:56:11.285729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T21:52:23.150188Z digest=sha256:cf19fa0c2f060bdc8e13e92783e8a8d83b29ba2f7aff7ceb119c9446f615043b

Observation f8dcd612-4f0c-4991-9f92-1439fb961543 · inbound

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models cites this paper.

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.791066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T13:04:14.886733Z digest=sha256:6052a194f8257b7700cda9d1ed7c68cb3497d939777a89232a77af9c698b61db

Observation 654c879a-756d-4bd0-8cac-797032805010 · inbound

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning cites this paper.

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning ClipCap: CLIP Prefix for Image Captioning

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:07:48.402050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T10:28:11.440915Z digest=sha256:89c93652bdf90e2f41fa728eac642ef4482a578baf1abe92ed1aaff74fd90fe2

Observation a362ee84-0f95-42af-9045-4b243e1564fa · inbound

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality cites this paper.

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality ClipCap: CLIP Prefix for Image Captioning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:28:31.474526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T07:03:50.311891Z digest=sha256:ab0fa0904d1e75965135fb7d5f3ac840bc2bdb2b467380c10f0d5bb4a162f641

Observation 0d8381db-8362-4d59-a50f-42218480667e · inbound

Revisiting LLM Adaptation for 3D CT Report Generation: A Study of Scaling and Diagnostic Priors cites this paper.

Revisiting LLM Adaptation for 3D CT Report Generation: A Study of Scaling and Diagnostic Priors ClipCap: CLIP Prefix for Image Captioning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:45.673868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T03:21:31.374575Z digest=sha256:3fcf24772a568543e4ebe0a6d20a2e0ef95110ea78c99ed17d4a2c6a11b6e8ad

Observation bafcd518-7a9b-4e9f-bb70-5fe86f8df660 · inbound

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations cites this paper.

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations ClipCap: CLIP Prefix for Image Captioning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:44.428197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T09:09:23.918802Z digest=sha256:178bfec72623d626848621ebd9118c70359b05ffb4c3a7cc41a98d0ab8fe34c0

Observation a5b94a51-91e3-4fd7-a8e7-d2225fc2f004 · inbound

Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection cites this paper.

Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection ClipCap: CLIP Prefix for Image Captioning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T12:36:44.253836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:36:44.253836Z digest=sha256:9ef5045d90bcd86a3fc575d390dbd3c1af021fdbd1639795f7396344a0e2cc67

Observation 8a086778-493c-4d97-8136-8dd41944fa15 · inbound

REPREC: Representation Driven Parameter-Efficient Recommendation System cites this paper.

REPREC: Representation Driven Parameter-Efficient Recommendation System ClipCap: CLIP Prefix for Image Captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T04:11:19.003885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:11:19.003885Z digest=sha256:4d5465ffa600724785bc9d6ee22f067b282b824bbfc50e09df121094b07332d9

Observation 5627ffab-96fe-4073-beb5-8f53e77b3fa6 · inbound

Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning cites this paper.

Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning ClipCap: CLIP Prefix for Image Captioning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T15:58:39.520616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:58:39.520616Z digest=sha256:4bac5c2ffa77aa3d9d5185ca9d503124406ce936687006933a8fe100c0e8e731

Observation 4410c2a3-1780-45aa-a9df-09d1e6aa492e · inbound

Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis cites this paper.

Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis ClipCap: CLIP Prefix for Image Captioning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T00:56:16.585209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:56:16.585209Z digest=sha256:46282bb5b877fb6b80e176e05a4c2b14615f4e15c5c01ea834dcf7d6996a3882