Pith. sign in

Paper Citation Record · LEDGER

ClipCap: CLIP Prefix for Image Captioning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:2111.09734.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.09734 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:27:25.774500Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:09:44.426723Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9e39cb2a-dfe9-4514-bac3-8f235c4f40e0 · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language ClipCap: CLIP Prefix for Image Captioning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.718168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:75d61144d1abfc3fbeb7ac2c7fef9f393fefa52f8422db039d28ce9a7b225a8c

Observation 8cfeab8e-d3b3-4a90-9f76-16506dde437b · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning ClipCap: CLIP Prefix for Image Captioning

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.476566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:016fac76fe24f03ac62aade5c27378c3f8fac2e8f28d200abc8cdae5573875c2

Observation 7f11e737-20da-46c7-9a1c-c8852ac93363 · inbound

LAION-5B: An open large-scale dataset for training next generation image-text models cites this paper.

LAION-5B: An open large-scale dataset for training next generation image-text models ClipCap: CLIP Prefix for Image Captioning

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:22:17.333000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T14:22:16.968028Z digest=sha256:7407ffbecc3b54c7da9c4d7f11557390b77b127ea40ecff51543ce09ee27ddbc

Observation 1ef5c014-1562-48d7-b252-38993a4e3396 · inbound

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention cites this paper.

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention ClipCap: CLIP Prefix for Image Captioning

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:07:42.528525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T23:07:42.245641Z digest=sha256:14073124eb1dfc1381aae0c8d2841609b9ebb8bb9f3096ee67e3a5bdf41da601

Observation f7f1d37a-c0ba-42f8-a149-1b8b39a8f713 · inbound

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model cites this paper.

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model ClipCap: CLIP Prefix for Image Captioning

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T08:41:04.887711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T08:41:04.743886Z digest=sha256:9cc3afbb03fe7bce1424745875fb4a5eea202aae815ca70aa91515655b46f5f4

Observation c9e2d1eb-3be7-4be7-96fe-9b7e798c957e · inbound

EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers cites this paper.

EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers ClipCap: CLIP Prefix for Image Captioning

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:11:49.649827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T06:11:49.475825Z digest=sha256:1c05da74b01591e026e551db3e722e351e0c11f44e3eb012360db4a0b276800c

Observation 7656f08c-4e8f-485a-a9d6-9c2b5ed4700a · inbound

Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search cites this paper.

Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search ClipCap: CLIP Prefix for Image Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T17:27:25.774500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:27:25.774500Z digest=sha256:d4a4fb88cc276354fbab6fcf5ecaad8289b9ce89a49a439da03e28d43bf8c4a3

Observation 6624c8e7-75f7-4c57-bb29-6bfab70825db · inbound

SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps cites this paper.

SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps ClipCap: CLIP Prefix for Image Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:16:36.087824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:16:36.087824Z digest=sha256:acffda1bf90dca2a19172ea22488f2ce7ff64186d59f407fe53a8c3dec9c75e3

Observation b33833fc-4b2a-4ccf-82ac-1e8cdf30cc3c · inbound

Enhanced Multimodal Aspect-Based Sentiment Analysis by LLM-Generated Rationales cites this paper.

Enhanced Multimodal Aspect-Based Sentiment Analysis by LLM-Generated Rationales ClipCap: CLIP Prefix for Image Captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:18.994193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:18.994193Z digest=sha256:feac597ada9e75d4561f9a56285dd0000531413e3f5175a0f3c8a85c315bbc7d

Observation c22dbbde-1000-446b-a047-623d6876de8c · inbound

Optimizing fMRI Data Acquisition for Decoding Natural Speech with Limited Participants cites this paper.

Optimizing fMRI Data Acquisition for Decoding Natural Speech with Limited Participants ClipCap: CLIP Prefix for Image Captioning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:19.824796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:40:19.824796Z digest=sha256:544615c90ec82e8f3b923e678adab7c3e00bd31c16bf9988e265bd22d77719a8

Observation ab95742f-af91-4837-a8be-418c7492f131 · inbound

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model cites this paper.

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model ClipCap: CLIP Prefix for Image Captioning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:44.090056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:44.090056Z digest=sha256:5f23caaba1af924e0e16e1b3742a389556e36919aacda0b6e945fbac112a3261

Observation 22bc1e94-042a-402c-aa38-1f117cddb53e · inbound

Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models cites this paper.

Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models ClipCap: CLIP Prefix for Image Captioning

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:36.351321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:35:36.351321Z digest=sha256:34759265c24a6ec1272d6513b9e70cda710b45dc075a5b537346f0a318fd3733

Observation 01ceb64c-abd4-4229-8c0d-a17cd390032c · inbound

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering cites this paper.

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering ClipCap: CLIP Prefix for Image Captioning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:10.602373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:10.602373Z digest=sha256:cdce98ee9c072c91d815200fda59a039fa71a9232dc6c0b3464c81b740b20cb6

Observation fe79f438-6836-4404-9e98-d2e7b2503694 · inbound

Diffusion-based Cumulative Adversarial Purification for Vision Language Models cites this paper.

Diffusion-based Cumulative Adversarial Purification for Vision Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:59:18.927640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:59:18.927640Z digest=sha256:8aed347f717ddb42e8742d6aa948fc8338271828deb83a89f499ac68823d1ffc

Observation f8c28283-b524-4c6a-bf1d-4adbcaf03595 · inbound

CoLMbo: Speaker Language Model for Descriptive Profiling cites this paper.

CoLMbo: Speaker Language Model for Descriptive Profiling ClipCap: CLIP Prefix for Image Captioning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:15.589604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:15.589604Z digest=sha256:3d633b5aff1920b320d6bfb948d9c4143f1509f416483de45ec591dccef2302c

Observation cf1e06e5-7ae9-47e5-bee9-f8ce5f3120ed · inbound

InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning cites this paper.

InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning ClipCap: CLIP Prefix for Image Captioning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:58:07.909306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:58:07.909306Z digest=sha256:fc3eff19895e9acecfc9d73530338cb10554f786bc0729a1f1b5e5fcc8c395ed

Observation 3bb5433a-a261-423b-a2dc-691adfbaa877 · inbound

CLIP-HandID: Vision-Language Model for Hand-Based Person Identification cites this paper.

CLIP-HandID: Vision-Language Model for Hand-Based Person Identification ClipCap: CLIP Prefix for Image Captioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:56:01.612867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:56:01.612867Z digest=sha256:defcedfe3d94f5226a842252d4fe8e3aaaad8ba5757a552bdf08467c7f224e7c

Observation 59700003-ea25-4573-9220-5fa415e5f49c · inbound

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge cites this paper.

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge ClipCap: CLIP Prefix for Image Captioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:41.828926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:41.828926Z digest=sha256:fe626f5ad010e00def0b3c3e9604b94356460ae6adda6466326df97ce76c8a5b

Observation 46177889-2581-457d-bd97-0e86f0a016ed · inbound

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition cites this paper.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ClipCap: CLIP Prefix for Image Captioning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:55.557921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:55.557921Z digest=sha256:1554996e44604561f190ae4db274bd9a14a2303fca4ae53955e3ff3cde2ef5d5

Observation ca8ccce3-3a66-418b-96d1-4e4cd7f48a02 · inbound

On the rankability of visual embeddings cites this paper.

On the rankability of visual embeddings ClipCap: CLIP Prefix for Image Captioning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.511161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.511161Z digest=sha256:89b05099011a786ca7635d9c3532428f0efc3a68dbbeec842c7ddcfaebce83c5

Observation 5772af9d-23eb-449e-94e4-c3813760953a · inbound

GLAD: Generalizable Tuning for Vision-Language Models cites this paper.

GLAD: Generalizable Tuning for Vision-Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:21.189934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:21.189934Z digest=sha256:cdb7bfd2e6bbc4ff5a0654d1efc0274f823f9b4cc30d6b8a8c57ddfdd1b0e884

Observation bc89fd79-83dc-4729-b9f0-7d938b076d9e · inbound

HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals cites this paper.

HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals ClipCap: CLIP Prefix for Image Captioning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:32:42.735751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:32:42.735751Z digest=sha256:c38db86741cf2089232bf07cad4311dbadbc0be74c6cf11c4a21d48c88caf064

Observation 31ab0b8a-1df3-43a3-a031-d3671c0d56f2 · inbound

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings cites this paper.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings ClipCap: CLIP Prefix for Image Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.416252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.416252Z digest=sha256:12a94da8d2e95d704036c7cc7dac0d0fa5bda84ec3c19238b83ddedd6957813f

Observation e640f75c-908b-4092-86a0-39c9f77051fa · inbound

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models cites this paper.

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:51:16.018283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:51:16.018283Z digest=sha256:ed4bb610e82b0a71307b6d2c413953419ccba71bbb4f53e77a08831af207d11e

Observation 6ce12810-cef8-4ad1-84c7-e4a6e8381ad2 · inbound

RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification cites this paper.

RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification ClipCap: CLIP Prefix for Image Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T01:04:10.551418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:04:10.551418Z digest=sha256:924d31caa7d218b4e449011f2fbccd50ad39a9155c7dfb58a3893ada4a190c30

Observation 37b79eda-e591-486e-aa15-4fabec394308 · inbound

From Image Captioning to Visual Storytelling cites this paper.

From Image Captioning to Visual Storytelling ClipCap: CLIP Prefix for Image Captioning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T10:27:39.338214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:27:39.338214Z digest=sha256:e823e0b488fbcf11ca7d635ad934d803248e6b121e01d15f092580ffce32b23e

Observation 572e5683-939d-454a-90ad-ac72b0abbdc5 · inbound

Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval cites this paper.

Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image Retrieval ClipCap: CLIP Prefix for Image Captioning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T17:25:01.008591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:25:01.008591Z digest=sha256:c5f463b7f16ca268d35e229442b5d8410950ccf8ab1a5b1ca2374d17936a8230

Observation 5d63d557-c52c-4e6c-9495-328a74914cf5 · inbound

Sample-efficient Integration of New Modalities into Large Language Models cites this paper.

Sample-efficient Integration of New Modalities into Large Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T06:00:27.509894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T06:00:27.509894Z digest=sha256:52e81d99f153f7f218de37a211868a3101f60484c79470178f1de6056892354c

Observation cdb27136-dcf8-43a4-98ed-2f7bc45059c5 · inbound

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models cites this paper.

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models ClipCap: CLIP Prefix for Image Captioning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T04:53:40.077127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:53:40.077127Z digest=sha256:a99df999f496ead07e208a8b366d2e4331e08355be30c149998d76a724c2372c

Observation 2b145175-6648-4219-93e1-146d1a146dd3 · inbound

Unpacking Hateful Memes: Presupposed Context and False Claims cites this paper.

Unpacking Hateful Memes: Presupposed Context and False Claims ClipCap: CLIP Prefix for Image Captioning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T10:26:22.120641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:26:22.120641Z digest=sha256:3df0b552b7c877feba52e5ea39e7bd3506afe272bf3e856f12c05e91281e7c85

Observation b975f0e4-ca19-4fb7-8c9c-0fd08cb67a97 · inbound

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models cites this paper.

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models ClipCap: CLIP Prefix for Image Captioning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T18:25:04.998120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:25:04.998120Z digest=sha256:476cbaf2ae4f02afd906179c54a3fa56476a4768abdcc58b5ae7fd204e73a413

Observation c1124a2e-2683-494b-8f0b-106e84704a6b · inbound

Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain cites this paper.

Multi-Modal LLM based Image Captioning in ICT: Bridging the Gap Between General and Industry Domain ClipCap: CLIP Prefix for Image Captioning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:27:59.324632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T14:26:08.188950Z digest=sha256:83ca9e223de8d9a77d538dd9300bc42b5a6e3be38b3a619b4d088d2504e2debd

Observation b0ce17dd-d1ad-4f4c-bd35-81170c9a0278 · inbound

LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs cites this paper.

LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs ClipCap: CLIP Prefix for Image Captioning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T06:06:52.217522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:06:52.217522Z digest=sha256:94d160edf899ec9c852502d9f36a8d555e44cdc61accc5f8b7e8787a3f987dff

Observation 9df69bec-bde3-4b33-8a69-849040417ee2 · inbound

Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks cites this paper.

Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks ClipCap: CLIP Prefix for Image Captioning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:55:52.655740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T20:26:42.128350Z digest=sha256:c1665fe6cdfd2185cba94eddb43b1fbdb5f5f5000da3bd51420054913aabd8f0

Observation 5f07988e-3ad6-4978-80fe-4f081982dca2 · inbound

Beyond Semantics: Uncovering the Physics of Fakes via Universal Physical Descriptors for Cross-Modal Synthetic Detection cites this paper.

Beyond Semantics: Uncovering the Physics of Fakes via Universal Physical Descriptors for Cross-Modal Synthetic Detection ClipCap: CLIP Prefix for Image Captioning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:53.240217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:52:34.752473Z digest=sha256:112fe6bef9516e5de20fc5d3e067df5bbcde6cfb7360563e1ae23f2206096f07

Observation 71069924-22ff-4aa0-9434-e119450329e7 · inbound

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval cites this paper.

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval ClipCap: CLIP Prefix for Image Captioning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:50.060959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:31:53.371412Z digest=sha256:fccc107e0de4cdad41df605a07ce5b9cca56af5c5143877a09ac4fc8f0efd01b

Observation 7220993c-f54f-4276-a5d2-fa3928c388d9 · inbound

UIPress: Bringing Optical Token Compression to UI-to-Code Generation cites this paper.

UIPress: Bringing Optical Token Compression to UI-to-Code Generation ClipCap: CLIP Prefix for Image Captioning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:00:59.421777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:21:32.024105Z digest=sha256:b9dbb58fb606f70011beff6cf4637db6f6a0f057cf569e892392057a4ff84fad

Observation 8baee550-4afd-4f0e-93bb-73d6a391dc69 · inbound

Semantic Manipulation Localization cites this paper.

Semantic Manipulation Localization ClipCap: CLIP Prefix for Image Captioning

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:21:01.497760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:32:30.345594Z digest=sha256:ee3d50409249c171bd0b427450951eb1679fb148915dfcf475a1acc74cfb7347

Observation 0f5c2bc7-85a8-4cc2-9b2d-d5129afbf14b · inbound

Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment cites this paper.

Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment ClipCap: CLIP Prefix for Image Captioning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:21:27.865530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:19:31.570783Z digest=sha256:51def8c5a7da81ac5b0212932b487cac1515864532e9cf4a20a0284ac5b2c63e

Observation 6448c78a-b35d-4ece-9770-8503072ef485 · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:22:56.290381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:73782274643d7ff09e5d519091ee62c8659de11c19326da927cef926cde1c515

Observation c8556bc0-b861-478f-a3f6-fd1aa2d5a44e · inbound

CB-SLICE: Concept-Based Interpretable Error Slice Discovery cites this paper.

CB-SLICE: Concept-Based Interpretable Error Slice Discovery ClipCap: CLIP Prefix for Image Captioning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:33:15.862374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:24:05.284777Z digest=sha256:7647a21df282e462ca449cd665ea6d17d0280fcee7229f8262a883866ba8545e

Observation 7d56f8ed-0f49-49f4-9f85-a227732d0549 · inbound

BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning cites this paper.

BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning ClipCap: CLIP Prefix for Image Captioning

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:56:11.285729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T21:52:23.150188Z digest=sha256:ccb48faa4af9700bc8bf23d6b9ac08307beb62e7a5febbcf8fbe3c55ff1925b1

Observation f8dcd612-4f0c-4991-9f92-1439fb961543 · inbound

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models cites this paper.

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models ClipCap: CLIP Prefix for Image Captioning

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.791066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:04:14.886733Z digest=sha256:a00398bf90e7cdd4a693175856a307c03f69014e5804f01474c90815d53fc1f5

Observation 654c879a-756d-4bd0-8cac-797032805010 · inbound

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning cites this paper.

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning ClipCap: CLIP Prefix for Image Captioning

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:07:48.402050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T10:28:11.440915Z digest=sha256:6f125f9a6eef3d745f4d84965ac022dfe941115c503b604a47d606914a78b247

Observation a362ee84-0f95-42af-9045-4b243e1564fa · inbound

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality cites this paper.

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality ClipCap: CLIP Prefix for Image Captioning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:28:31.474526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T07:03:50.311891Z digest=sha256:4db020a8c81f7479729069e95642a43a69bb04ef03ba29e8d001db69f14246b6

Observation 0d8381db-8362-4d59-a50f-42218480667e · inbound

Revisiting LLM Adaptation for 3D CT Report Generation: A Study of Scaling and Diagnostic Priors cites this paper.

Revisiting LLM Adaptation for 3D CT Report Generation: A Study of Scaling and Diagnostic Priors ClipCap: CLIP Prefix for Image Captioning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:45.673868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T03:21:31.374575Z digest=sha256:917b65bef90b9be58ab96c544135d225612311395cab5c5add762f02c7021072

Observation bafcd518-7a9b-4e9f-bb70-5fe86f8df660 · inbound

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations cites this paper.

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations ClipCap: CLIP Prefix for Image Captioning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:44.428197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T09:09:23.918802Z digest=sha256:aac504d76d25568b2aefcffa1845feeb86863663d5dc99d87a7fc61e9924a7c5

Observation a5b94a51-91e3-4fd7-a8e7-d2225fc2f004 · inbound

Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection cites this paper.

Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection ClipCap: CLIP Prefix for Image Captioning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T12:36:44.253836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:36:44.253836Z digest=sha256:9ef5045d90bcd86a3fc575d390dbd3c1af021fdbd1639795f7396344a0e2cc67

Observation 8a086778-493c-4d97-8136-8dd41944fa15 · inbound

REPREC: Representation Driven Parameter-Efficient Recommendation System cites this paper.

REPREC: Representation Driven Parameter-Efficient Recommendation System ClipCap: CLIP Prefix for Image Captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T04:11:19.003885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:11:19.003885Z digest=sha256:0ea222d294c74567e98f0a2600a2b8cf194cf0836298fbfb8d2891496dd88b2c

Observation 5627ffab-96fe-4073-beb5-8f53e77b3fa6 · inbound

Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning cites this paper.

Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning ClipCap: CLIP Prefix for Image Captioning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T15:58:39.520616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:58:39.520616Z digest=sha256:4bac5c2ffa77aa3d9d5185ca9d503124406ce936687006933a8fe100c0e8e731

Observation 4410c2a3-1780-45aa-a9df-09d1e6aa492e · inbound

Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis cites this paper.

Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis ClipCap: CLIP Prefix for Image Captioning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T00:56:16.585209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:56:16.585209Z digest=sha256:cbe2cf485a65aa34677a2720d37fe33472daf30b7f1b905c53ad448fb3482908