Pith. sign in

Paper Citation Record · LEDGER

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching

As of 10 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2506.23502.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23502 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:46:02.319525Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T04:46:34.946640Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T04:49:02.937990Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy52
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f4e51e50-4dc3-49c3-ae30-6e3617a91fff · outbound

This paper cites Incorporating geo-diverse knowledge into prompting for increased geographical robustness in object recognition.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Incorporating geo-diverse knowledge into prompting for increased geographical robustness in object recognition

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:11.326339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:58.204051Z digest=sha256:942f6033a1dc1de117a84c562f55a5679a2b22061b0fe0318cb109c89aa0e5c8

Observation c74f03bd-4c91-43f0-991b-d0ac68c1b878 · outbound

This paper cites Learning the best pooling strategy for visual semantic embedding.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Learning the best pooling strategy for visual semantic embedding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:11.079672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:58.252978Z digest=sha256:093c4a067464e26842f063c42fa6cec4bcefbe29e7da776cf1028621bad654c6

Observation 9652da36-73e0-41f2-a322-e5c10925e01e · outbound

This paper cites Large language models are visual reasoning coordinators.NeurIPS, pages 10–16, 2024.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Large language models are visual reasoning coordinators.NeurIPS, pages 10–16, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:10.883779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:58.322328Z digest=sha256:6c062b1a56342ef7deff3d43c4325538437e8e928efd2f20db9843fb027e9ff2

Observation ea30dcd6-f499-4062-91a9-6da3efc4fa07 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:58.361450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:45:58.361450Z digest=sha256:21427ee53077e81fd357ae392f82af34e70b77ecc79eeb5088f12cc31539ecfa

Observation 77d9e395-835c-4803-a5cc-711754ad2412 · outbound

This paper cites Cross-modal graph matching network for image- text retrieval.TOMM, 18(4):1–23, 2022.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Cross-modal graph matching network for image- text retrieval.TOMM, 18(4):1–23, 2022

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:10.698382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:58.450649Z digest=sha256:3079d0b7280f7719c6d98ec985085398fc3ce9d2a960cd11b178c1fe32c7ff25

Observation b2a6fbbc-57a4-43f6-a773-5ec29fdd347f · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Reproducible scal- ing laws for contrastive language-image learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:10.582949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:58.532573Z digest=sha256:f4eb4eb56deaf5e7ba8976529199d74e982181b8736aa68d1c80ea53fa20b575

Observation f8701963-6001-43a5-a0c7-04f1a90cf92e · outbound

This paper cites Sim- ilarity reasoning and filtration for image-text matching.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Sim- ilarity reasoning and filtration for image-text matching

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:10.437714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:58.566838Z digest=sha256:b1ec5e9d181d234974ae430a1315479f5bed6994f71e05e4f0d07b3d4431c06a

Observation 351997d4-cfa6-4159-8360-6514f64e8570 · outbound

This paper cites Fleet, Jamie Ryan Kiros, and Sanja Fidler.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Fleet, Jamie Ryan Kiros, and Sanja Fidler

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:10.289170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:58.636191Z digest=sha256:3858f7b29e2e029f44e60dbe4b7d8b7c458b83798655dc57dd1c49f41ffce0a2

Observation 095d7716-b2e3-4634-bb8a-5ea1bdb37e31 · outbound

This paper cites Learning semantic relationship among instances for image- text matching.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Learning semantic relationship among instances for image- text matching

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:10.097973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:58.705862Z digest=sha256:783c5172b9403cf863dff0b42a793eb8746abd6dbe29aeb6c33c29a07208fb3c

Observation e664c977-85da-4a1a-a165-cdcda54adc90 · outbound

This paper cites Pyramidclip: Hierarchi- cal feature alignment for vision-language model pretraining.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Pyramidclip: Hierarchi- cal feature alignment for vision-language model pretraining

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:09.926777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:58.790852Z digest=sha256:6d57ce46402ccc70a034b3aea1dbe66d38307d635db4b1816feda3a12fb257d8

Observation dee985aa-e72f-41c5-9cae-5522cf82d76d · outbound

This paper cites Cross-modal semantic enhanced interaction for image-sentence retrieval.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Cross-modal semantic enhanced interaction for image-sentence retrieval

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:09.708912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:58.898987Z digest=sha256:c41a3073df88ad328cc0a31b49d6c6515d0a792fd455a58d5eb94c72e340943d

Observation b42cf75e-bdf4-481f-bd3a-d4dd5c018bdc · outbound

This paper cites Hiclip: Contrastive language-image pre- training with hierarchy-aware attention.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Hiclip: Contrastive language-image pre- training with hierarchy-aware attention

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:09.527887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:58.985185Z digest=sha256:42346d32f59be3e81105301bc761a3c2921fb0c6a8bc70b41ff8d8552434925c

Observation b1934b8d-98d5-4a57-bcd2-32ae4b06f868 · outbound

This paper cites From im- ages to textual prompts: Zero-shot visual question answering with frozen large language models.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching From im- ages to textual prompts: Zero-shot visual question answering with frozen large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:09.381531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:59.056867Z digest=sha256:40109344d1a8d3720454238d675d6dfa08fcca258ad013307ba971e4ac3747e0

Observation e5f97260-9444-42ce-a1fb-844f1d922875 · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Visual program- ming: Compositional visual reasoning without training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:09.245934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:59.127145Z digest=sha256:e47814811d45d2dd3fa8505879b2b26623315728c5946cc4a7aadd7d9d6510b9

Observation 3231f3df-5b52-4677-a74f-d088b652f9e4 · outbound

This paper cites Structure-clip: Towards scene graph knowledge to enhance multi-modal structured representa- tions.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Structure-clip: Towards scene graph knowledge to enhance multi-modal structured representa- tions

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:09.059318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:59.196134Z digest=sha256:10a04fc88ae823624b31075be5dc7bcafaaa9167f9f184d5025b2c6184419698

Observation 2309f4f4-fa2c-41ed-b325-866efec1c4a0 · outbound

This paper cites Fineclip: Self-distilled region-based clip for better fine-grained under- standing.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Fineclip: Self-distilled region-based clip for better fine-grained under- standing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:08.957054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:59.266266Z digest=sha256:07824958e82ab5b488080667d9c40f9b9e9f33f3700bee064bc5f6db2f229a73

Observation 0c33dae3-0e7e-4537-b57b-8138ce59554b · outbound

This paper cites Knowledge-aware prompt tun- ing for generalizable vision-language models.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Knowledge-aware prompt tun- ing for generalizable vision-language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:08.801890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:59.389988Z digest=sha256:e81d39cfa3da98e0ffe15b0c3f8768e47c5bc03c5ecbedd24c59af9d915ccc71

Observation 8ae628f9-cb49-4137-9e6e-ec1174499752 · outbound

This paper cites Maple: Multi-modal prompt learning.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Maple: Multi-modal prompt learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:08.621773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:59.431964Z digest=sha256:0b87bd0eca22bfb1345a96ef586716c45921f997070b9e13c24fd6080accff3c

Observation 37eeb6a9-2b61-4b18-b572-09e88fe4e44e · outbound

This paper cites Stacked cross attention for image-text matching.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Stacked cross attention for image-text matching

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:08.493167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:59.503349Z digest=sha256:0160bdec80d75f0bfa51d0ebee75c349f1f61cf66191406c25baf139f8792345

Observation 66a2902f-4696-492e-a048-c47e39ade055 · outbound

This paper cites Action-aware em- bedding enhancement for image-text retrieval.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Action-aware em- bedding enhancement for image-text retrieval

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:08.391733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:59.578340Z digest=sha256:e62231cd03a0f7719ca9ca29f36903c358dc76511315cedcfbed012902fd05dc

Observation 2e7a3d5f-46e1-487c-a0af-9296832a10e9 · outbound

This paper cites Learning background prompts to dis- cover implicit knowledge for open vocabulary object detec- tion.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Learning background prompts to dis- cover implicit knowledge for open vocabulary object detec- tion

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:08.296908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:59.664580Z digest=sha256:81b2b374f0a16f7114779fd47b72bcf4e5bacd675dcee97e05d218673a8aa6ce

Observation 54680ab0-30df-4504-99de-ff27fa414cb2 · outbound

This paper cites Cross- modal alternating learning with task-aware representations for continual learning.TMM, 2023.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Cross- modal alternating learning with task-aware representations for continual learning.TMM, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:08.134264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:59.724477Z digest=sha256:3a28b173d98c48dbef4e673cbb9b25fb790a3d0f6964232673c8cb965b287d83

Observation 1abbc369-b536-41d1-8579-bc551a8d61ea · outbound

This paper cites Image-text bidirectional learning network based cross-modal retrieval.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Image-text bidirectional learning network based cross-modal retrieval

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:07.992070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:59.808291Z digest=sha256:7353e5126608421d3a1f565ce374c7ff6a29572be0923d2d533eb9ac3631bcc2

Observation 85adc5a4-ee2a-49da-af69-1ad9ebcfc1f7 · outbound

This paper cites Learning customized visual models with retrieval-augmented knowledge.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Learning customized visual models with retrieval-augmented knowledge

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:07.878180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:59.897619Z digest=sha256:5e10fcdd8c32ba4e7aed9f30954a3756dc17f2bb4b6fd2e7807ab03194c7d08c

Observation 0cce8655-cbc3-4550-86c7-124630b1c125 · outbound

This paper cites Multi-modal attribute prompting for vision-language models.TCSVT, 2024.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Multi-modal attribute prompting for vision-language models.TCSVT, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:07.818852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:45:59.975439Z digest=sha256:96486a0ef15aa216d1987e3e9fc64a2e081569d72616312989c220a46d5b62b2

Observation 71ccba72-23ac-486c-93f6-c97ae5125f4b · outbound

This paper cites Fine-grained visual– text prompt-driven self-training for open-vocabulary object detection.TNNLS, pages 1–11, 2023.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Fine-grained visual– text prompt-driven self-training for open-vocabulary object detection.TNNLS, pages 1–11, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:07.698542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:00.077299Z digest=sha256:4c05363d441c9739444330c725a09819a7c6722978025c98d9367445d31eb5ee

Observation 216fa119-48a1-4cea-943f-eef01ebd306f · outbound

This paper cites Visual classification via description from large language models.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Visual classification via description from large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:07.545899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:00.144303Z digest=sha256:2030ea16b38b3004d9881a08362a5075bc978be368c2c1dda24ad427bea0f38f

Observation f7bfca97-fb79-4275-8ad4-e8578c7cb3fd · outbound

This paper cites SCHEMA: state changes matter for procedure planning in instructional videos.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching SCHEMA: state changes matter for procedure planning in instructional videos

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:07.369802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:00.199420Z digest=sha256:6f8bc477c2cee92ddef727eb18b14196d161ad19fa9f2867d77595c7442370bf

Observation 492369ea-3c73-4552-ae31-4dc7e546eac7 · outbound

This paper cites Teaching clip to count to ten.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Teaching clip to count to ten

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:07.239545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:00.285075Z digest=sha256:5f9e89484bfb0aad33676698dddec52172ce5fcd71fab9f2a7bb4f5c71f9f2d0

Observation 1a958523-11f2-4aac-9de5-3c11cf5486b2 · outbound

This paper cites Fine-grained image-text matching by cross-modal hard aligning network.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Fine-grained image-text matching by cross-modal hard aligning network

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:07.051157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:00.354666Z digest=sha256:ea3eda40fa2be4302beb4af4da3ac13cc5e97d0abc066bf42ca4c702eebae924

Observation 3d500885-10b2-4dfe-83df-900a3002a46d · outbound

This paper cites Dynamic modality interaction modeling for image-text retrieval.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Dynamic modality interaction modeling for image-text retrieval

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:06.880385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:00.412406Z digest=sha256:78d9a1a53dafe2c878be3302be32445a80dd6e6fe866627076d99e570f7884be

Observation 784facf2-5037-4c1f-8ffa-229826f8e16a · outbound

This paper cites Learning transferable visual models from natural language supervision.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Learning transferable visual models from natural language supervision

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:06.607635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:00.509155Z digest=sha256:1eb00b82448cb9d0264983e0ea2394f115e1954136adf6f583ce2ca79f676d60

Observation df139403-5b87-4590-aef3-278e5f8cfb7a · outbound

This paper cites Vlc-bert: Visual question answering with contextualized commonsense knowledge.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Vlc-bert: Visual question answering with contextualized commonsense knowledge

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:06.368510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:00.576317Z digest=sha256:0e5b2caaeb3deb48afc59dd2e8b6cedae50381357e44180553a88b6be52c504c

Observation 6fec0c54-161a-4dde-ad96-fec20fb6ce54 · outbound

This paper cites Language models are causal knowledge ex- tractors for zero-shot video question answering.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Language models are causal knowledge ex- tractors for zero-shot video question answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:06.147358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:00.647825Z digest=sha256:69bf8d2e571fdb1185fb3ef1c1f22528a7b4bef907e024519a34c48e7601553b

Observation 97d650ee-0f8e-49fa-b45a-415ff64f2dc8 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:46:00.728702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:46:00.728702Z digest=sha256:8d4da4c6d19050d963e4d4522e898b7a3a1355579a507bb4763bada4ed1a6974

Observation 306a3b8e-e638-4a3f-873f-ffc8c008123b · outbound

This paper cites Compound text-guided prompt tuning via image-adaptive cues.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Compound text-guided prompt tuning via image-adaptive cues

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:05.899299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:00.814075Z digest=sha256:72f3db2a7f69bd1823a24c3c0fb7fd3cd18dc00e0881b3686401bae262efb6a9

Observation 683ec262-ffed-4b1b-b9f8-1ba75477c251 · outbound

This paper cites Multi-granularity cross-modal align- ment for generalized medical visual representation learning.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Multi-granularity cross-modal align- ment for generalized medical visual representation learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:05.612801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:00.916739Z digest=sha256:e23bad6bdb0226b0e1f43e0dc5b57cd81274ff2495fc4fc3a45966ff8220161c

Observation e7505452-b647-4910-8541-9e75740df224 · outbound

This paper cites Consensus-aware visual-semantic embedding for image- text matching.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Consensus-aware visual-semantic embedding for image- text matching

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:05.284410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:01.008060Z digest=sha256:4a210d253dd4758a58b0d9885e4e56b8fbcd7252820e296a920d9c407e101861

Observation 106832fd-dd91-4d87-98f0-f50eb4dd2bb8 · outbound

This paper cites Vilt- clip: Video and language tuning clip with multimodal prompt learning and scenario-guided optimization.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Vilt- clip: Video and language tuning clip with multimodal prompt learning and scenario-guided optimization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:05.104822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:01.109203Z digest=sha256:dfe12d7f04be5fd17efd154a28fae5ea1f2f981d6e500a1773aa907477ed9329

Observation ce17a852-6b73-4383-837c-9ee1589a4735 · outbound

This paper cites Position-guided text prompt for vision-language pre- training.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Position-guided text prompt for vision-language pre- training

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:04.878349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:01.194101Z digest=sha256:8c0c601857e66615d876e908dfb0f1a49f43628e4a4b002315af6fe453a0b75e

Observation 4e85078f-a7d3-4f7d-9755-7bc0e018c68a · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching ActionCLIP: A New Paradigm for Video Action Recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:46:01.269123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:46:01.269123Z digest=sha256:80f6ac1c643e578cca5201d21e3a74786f70650f6368c47dbacd5ae76994e651

Observation 6e788e16-0286-4999-b93e-6f5467d83c86 · outbound

This paper cites Cross-modal scene graph matching for relationship-aware image-text retrieval.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Cross-modal scene graph matching for relationship-aware image-text retrieval

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:04.742902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:01.316475Z digest=sha256:a463d21d3931be32553228de9a3700bc61566c48c29a8ce8449ee9d1b8632a18

Observation 69cc5f4b-da85-4fc3-9195-af6880442ea6 · outbound

This paper cites Diffusion feedback helps clip see better.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Diffusion feedback helps clip see better

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:04.591379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:01.407864Z digest=sha256:0b3770f49bad04c64263b30a9f9da58dad3f344f4f65794933ac3e3939e8c16d

Observation c6b2526c-d082-4833-896e-56fb0d19f469 · outbound

This paper cites Balance act: Mitigat- ing hubness in cross-modal retrieval with query and gallery banks.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Balance act: Mitigat- ing hubness in cross-modal retrieval with query and gallery banks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:04.426824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:01.486247Z digest=sha256:b00fa0a98eea6d4deb2163fba6f9057a14a98f705cd052391ddf5d1a0d43cb60

Observation a085f20a-3d6a-4c66-95a4-c2e1911b43d8 · outbound

This paper cites Learning hierarchical prompt with structured linguistic knowledge for vision-language models.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Learning hierarchical prompt with structured linguistic knowledge for vision-language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:04.270094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:01.539766Z digest=sha256:9fb95744de28836e05d58832c2b4b8f4e85b4e657e0bebc2a7b47d8919425936

Observation 5b308134-caa9-4bed-893e-e343532fe2b1 · outbound

This paper cites Multi-view inter-modality representation with progressive fusion for image-text matching.Neurocomputing, 535:1–12, 2023.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Multi-view inter-modality representation with progressive fusion for image-text matching.Neurocomputing, 535:1–12, 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:04.098062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:01.607195Z digest=sha256:ce8cc39d52311b54a3f0f60290d5fb51d7aa8cb5a94f42a90def3adfcc861fff

Observation fcd561d1-4d2e-4a45-9c47-0713da403415 · outbound

This paper cites Saco loss: Sample-wise affinity con- sistency for vision-language pre-training.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Saco loss: Sample-wise affinity con- sistency for vision-language pre-training

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:03.952171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:01.685643Z digest=sha256:13eeaee55ae1a10eb2f5090bd4db281e6f1752204186509fe425994e44c11653

Observation 27802d92-26da-4d5c-a7a2-ada570d0ec48 · outbound

This paper cites Visual- language prompt tuning with knowledge-guided context op- timization.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Visual- language prompt tuning with knowledge-guided context op- timization

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:03.712784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:01.734901Z digest=sha256:db7923ba9e65d20c9a038432594cc5be08e1c11cc07b4c4444ad37ebd6236618

Observation 5c60afda-21e5-4646-838d-6aa3ccc5a0ee · outbound

This paper cites FILIP: fine-grained interactive language-image pre-training.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching FILIP: fine-grained interactive language-image pre-training

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:03.389697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:01.784837Z digest=sha256:231740c7fcc804866e86977979cc3a2c3d2de8314b670aece7fc07bc73ee31ed

Observation ffffa9ae-697c-4519-8c11-e33500c03a49 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.TACL, 2:67–78, 2014.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.TACL, 2:67–78, 2014

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:03.122250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:01.848329Z digest=sha256:ca9dcd9026f0197b66722059e4c81f6501e20342aa83a97bf05d53d0df7cf158

Observation e0c6aea0-a122-448b-a748-51b057bff30b · outbound

This paper cites Dual-path convolutional image-text embeddings with instance loss.TOMM, 16(2): 1–23, 2020.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Dual-path convolutional image-text embeddings with instance loss.TOMM, 16(2): 1–23, 2020

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:02.907213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:01.927793Z digest=sha256:0340447c0623b9efef6541440729440773d8643be345798a8b87b51951e75b23

Observation dde1ff41-5a67-4728-82f1-f784c595b202 · outbound

This paper cites Large Language Models are Good Prompt Learners for Low-Shot Image Classification.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Large Language Models are Good Prompt Learners for Low-Shot Image Classification

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:46:02.007187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:46:02.007187Z digest=sha256:1eadd466baedbd2b595b0908edc43fc3c3de55c316ae02615f7373e82596869c

Observation b13f4e3d-710c-4d46-b2bd-61e6ae75c0fe · outbound

This paper cites Regionclip: Region-based language- image pretraining.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Regionclip: Region-based language- image pretraining

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:02.807417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:02.083073Z digest=sha256:268315a58e699ff9d6ccc8df92ec72d47100a642c4c4efea0cda69223da6ad1d

Observation 15d59c97-d7a3-4b17-b0ac-37ee4fb5d7bf · outbound

This paper cites Relclip: Adapting language- image pretraining for visual relationship detection via rela- tional contrastive learning.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Relclip: Adapting language- image pretraining for visual relationship detection via rela- tional contrastive learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:02.679618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:02.167435Z digest=sha256:bfd29510846af68092c619b0de31b3e7f18db122ee0ee6a6f2cb2f0f8b843c14

Observation 757f26c7-3387-439a-ba94-4b7b9d7a6fe0 · outbound

This paper cites w/o ac- tion knowledge.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching w/o ac- tion knowledge

Reference 336

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:02.464676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:02.319525Z digest=sha256:33440434a96be8737f9a190547e1d7d919ab82f733bde509a68ac5c8179c9183

Observation 7d955b46-826e-4e4d-821c-5930fc88c276 · outbound

This paper cites Datasets Details Flickr30K[50] dataset contains 31,000 images collected from the Flickr website.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching Datasets Details Flickr30K[50] dataset contains 31,000 images collected from the Flickr website

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:46:02.578898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:46:02.237082Z digest=sha256:5296be9dabe4e42cc5881ce90c1ca57f678e2aaf708ad9b47017a26d69ffb698

Pith citing papers

Observation 99e15fe9-6af7-4601-9aff-15c5c1ba243f · inbound

Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection cites this paper.

Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:49:02.940987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T04:46:34.946640Z digest=sha256:ad5fdf4a896efd26a58dd9411979965ce4148ab95ffbdd5a70d6563d8b8fe13c