Pith. sign in

Paper Citation Record · LEDGER

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings

As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.17080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17080 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:02:08.473719Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved35
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e587879-ef23-4ec9-8f4a-c8eb9b7fd05e · outbound

This paper cites Henning, Karun Singh, Omkar Parkhi, and Fedor Borisyuk.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Henning, Karun Singh, Omkar Parkhi, and Fedor Borisyuk

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.355576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.355576Z digest=sha256:78c62b855032d9c0e37b7d4c264fa67b8750c3c857d26fefea1c88e5f135af79

Observation 844a9230-2331-4399-9665-0d3e25fdbdbc · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.359638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.359638Z digest=sha256:d87c8ddf1cb4e7cf1b8459e2fa049671d9331422cd1c181cf3d82319f9f1a137

Observation b64296a2-9663-44ac-b72f-94da67100eeb · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.270659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.362782Z digest=sha256:26db67131b86940c0d242bb3c4fb75102b0e7305efda194de100c239a6c5c611

Observation 765a8da2-d7fe-4482-8385-60761a367c32 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.365468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.365468Z digest=sha256:17ffb7e0b1eac6de9f03221afc15d23663dee84f94b81564631a959f0e01c9c5

Observation 1ee02fa7-f74e-4583-9dc4-81bf246a1003 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 5

Resolution
verified exact
doi, observed 2026-08-06T15:02:08.533236Z

Source-reported events for the cited work

correction dated 2023-01-23. Source: crossref record 10.1038/s41598-022-26364-y->10.1038/s41598-022-23052-9:correction, observed 2026-07-11T02:56:50.008426+00:00. This notice travels one citation hop only.

source=pdf_text observed=2026-08-06T15:02:08.370948Z digest=sha256:7b844e21680f5a92d0cc2ed8b6fbd99a7bad76d2c52dc3c5c3e5f9dcc73f1b04

Observation 341e46de-466d-49aa-9570-6f8171eb28c6 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.257971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.373920Z digest=sha256:04c523f3b461812a164c26c3e91f99b9a58477b13fe403a684e60ef85095e0c2

Observation 4175c460-0d17-4b84-8575-ef52509497aa · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.250094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.376798Z digest=sha256:771106b0e1e87f28245ee5e53c170435e84a7252553ed7b641fd874746b4414a

Observation 4e72cf30-9b17-4545-9688-aa8411e055e6 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.379262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.379262Z digest=sha256:759b6a39f32ac659da4a798ac29ffd2559ce374a20f1c016c85f6e1376262743

Observation cfb65044-be84-4fd5-b734-8c814278eb4f · outbound

This paper cites Huang, O.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Huang, O

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.382220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.382220Z digest=sha256:a62313bc5741a66e0529d7200a892a4bba6b343f01401ff20f99c4b7ec759385

Observation 1af7e53b-4716-4a9a-9792-d2c52d23a835 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.242063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.385017Z digest=sha256:e491cc8d7e09306b8eabfac16e49652edf7ffbe0fca22946708ed64c104e04d1

Observation d9f035e0-1131-4977-b355-036e288f4267 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 11

Resolution
verified exact
doi, observed 2026-08-06T15:02:08.520558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.387425Z digest=sha256:0703611c786a73dbdb98f4e068a096d45baf480f661d28312c45749b9373790c

Observation b4de798c-beb0-442a-9c44-1641e34bf37c · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.390026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.390026Z digest=sha256:ca4a3d9a0ec11a71d9af89f1ea7c0a5501a06c6f794d93f3221cbbe0f716c90f

Observation 70bd7a1a-276a-4ecb-8c2e-a43001c407c1 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.392478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.392478Z digest=sha256:f6ec42fe8a9c0eb74fe9f18fe6b745601f20ce313d6914c32aecee296fa0598f

Observation 49576e56-9fbd-43d9-bbda-b0a145568ab5 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.397586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.397586Z digest=sha256:d586ec6c3e58b0f4dabf2b52d85b74ff572e4e74fe99d6e725db9c8c8a715e66

Observation 4fbaae12-5680-4736-944d-c35b78eba9a0 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 15

Resolution
malformed identifier
no resolver link, observed 2026-08-06T15:02:08.399985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.399985Z digest=sha256:f874ba7ea7484494d45d264f302961e25843024c7c1da0a43563577c30ada434

Observation 82a568d5-e24d-470e-b562-a08f7a9e6716 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.228856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.402508Z digest=sha256:08d7afd594e122f2278f17db9ac0178aa13ed32dc5af7532d34a84d94851c552

Observation 8e825e7e-e9be-487c-92f5-f6d93eb75f15 · outbound

This paper cites Triple Modality Fusion: Aligning Visual, Textual, and Graph Data with Large Language Models for Multi-Behavior Recommendations.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Triple Modality Fusion: Aligning Visual, Textual, and Graph Data with Large Language Models for Multi-Behavior Recommendations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.404945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.404945Z digest=sha256:3375b9449a384177207d5da99e8f3aab07d35481b6207576d29cb762ffec53b1

Observation 5a95845c-8c68-49ca-a391-8a7b8dd4a6de · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.407832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.407832Z digest=sha256:237f418f9f4fe767b41aadfe7c42a3e9e1d1118977f6322d1ab35ce01cc53346

Observation 7adb5744-9b3f-41e3-a39b-a56e4afdee7a · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.220388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.410338Z digest=sha256:c1056c1e22120b162913aa7cfc3899d973104009f4ffc5e6e989f98cd8dbe240

Observation a2297b6e-6811-4110-abb9-1e3cbb392dbe · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 20

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T15:02:08.503053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.413385Z digest=sha256:b49314c7383e6ab4bee382332f9531066c21e1cf9c084bb2d2ba42ed38be8e4a

Observation 31ab0b8a-1df3-43a3-a031-d3671c0d56f2 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings ClipCap: CLIP Prefix for Image Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.416252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.416252Z digest=sha256:12a94da8d2e95d704036c7cc7dac0d0fa5bda84ec3c19238b83ddedd6957813f

Observation cd3a511c-8b4b-47d9-928a-3b31db4b6010 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.212382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.418876Z digest=sha256:2a81c50494811fc5eaafd2a73f27aafe49b59c3d5d3b402caccad35f7accdec6

Observation 8e350083-1d36-423e-8eb6-1849e69de9ba · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.204366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.421246Z digest=sha256:ef1614dd99900919ae405c57968cb4418cd977d5010ef103893b6b2432264d8f

Observation ca3c161f-50b8-44e6-994b-182804d4f47f · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.188234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.426399Z digest=sha256:d2029f4d78d3c1bd76662c4fb91c4bcf3a2669c25adeadb1d2e3302bcce84166

Observation 5760ac08-3b64-4215-9748-1f32df6eacae · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.428768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.428768Z digest=sha256:fd423cc6b13e2db526ac02e98388ac35a821a96d94302a90cd74061471788fe2

Observation ea5801f0-b793-425e-969d-ea6e774de381 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.175328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.434446Z digest=sha256:23713496bf91dcb55b7c9ec0eabf258ca2ed6cdc0bf32ccbb4ddc78c05a5bb1d

Observation ba62af85-ecad-41aa-bcef-c6615d55a0c3 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.439848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.439848Z digest=sha256:91a83b4883ca5dc19707ec4822e55e92c67685a12a1277b7304a9b5fe74b17da

Observation e8d1cd1c-7c72-4d48-84e8-2597df66a409 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.442379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.442379Z digest=sha256:86a678762b228ed0f13f2b2a0fc06e1cde8917df0eb933dd48001ead6eed1d9c

Observation 6b1aadf8-f6dd-4c60-8df0-cd4e244780d1 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.431508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.431508Z digest=sha256:aef26bebbf37fde9fb1dd826cb89a020d4a962e6f8ee2bf4e3fccf469a1290e1

Observation 9be1fbc0-3c67-4fb0-8c39-4cc115fd06f4 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.146325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.447332Z digest=sha256:344e94b80eeed58b2adbd742608b06097de824319461dd7b50d79beada4f4746

Observation 370b74f3-5ec6-4060-b0d2-3e56cbd5c5f5 · outbound

This paper cites In International Conference on Learning Representations (ICLR).

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings In International Conference on Learning Representations (ICLR)

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:09.167393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.437303Z digest=sha256:26a13e66f31b7b987c6b61db3e54f156147eba6210745afed97319ff7e98dc20

Observation a4c89211-4d1f-4078-bad7-72f1e5f6ec08 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.120199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.455153Z digest=sha256:cc17d6af864f6a4b91474c9b9db52ab53a9993b9187c8ac94116c9be36015f1f

Observation 41811caa-1aeb-4ed8-8b28-cb1af4b46947 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.460518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.460518Z digest=sha256:ae22ff8d465621bf582758b6922203bb6afed052be07649de5da338058c3bfa4

Observation 2534c8d2-e230-4d2b-9205-58838d3efa83 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.154214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.444911Z digest=sha256:76ddfeaa9813aa6eafe840f3d8cb2d8ab1cad2ee8ba5ab09e80b99557456b1e8

Observation 7e2841af-aa2e-4edd-88fb-cbf5fe245d5d · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.137817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.449944Z digest=sha256:726babbe9cf551b4095ad482253d0a5008929d5096e5d3648e3cd83a2b158172

Observation 54d10b83-08ad-4ff7-956d-39b7ff5e0e32 · outbound

This paper cites dress”, “rug.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings dress”, “rug

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:09.111978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.463065Z digest=sha256:60152b39d70214f1f8803ff72c04c99e84463d7d1bc06fe73688866776f08db6

Observation f20a7962-97a1-4ced-b2de-5bf712666180 · outbound

This paper cites productName.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings productName

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:09.103924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.465918Z digest=sha256:b0244d30aca2f3d7e2fcdfdb7931c2319d6a3197362bd857969295f4214e9a23

Observation 64f47fb2-427a-4ace-ba29-acfca3407879 · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.095624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.468567Z digest=sha256:65e99ac82ea1a13ff7b00a420c217d56875b175d72e44c95a88596a8ee4fe05c

Observation 643ea409-02d1-4c2e-9ab1-c1961ef3f94b · outbound

This paper cites an unresolved cited work.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:02:09.087403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.471120Z digest=sha256:4228ee6a323511194b8bb29a1a026d333c55a8876b5a2b0c2e26505eee476de4

Observation 89797b0e-cd56-42ec-9c02-a4aa148e0361 · outbound

This paper cites color”) and𝑣𝑖 is its value (e.g., “ blue.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings color”) and𝑣𝑖 is its value (e.g., “ blue

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:09.078869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.473719Z digest=sha256:a744f9bb89e3accc4a56ee140fd54a7ad5e7459562ebfa8380ed64ef5e7f48ab

Observation 67156b5c-067f-4e09-ac3d-33812b283b3b · outbound

This paper cites InProceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings InProceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:09.129394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.452541Z digest=sha256:9662be5a6dd62c74bdd179a63ab1416466ecebbf40b4d6dd99c11ff210bdd954

Observation 96eca364-7576-4559-95fa-b653762da436 · outbound

This paper cites In Computer Vision – ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXX (Glasgow, United Kingdom).

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings In Computer Vision – ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXX (Glasgow, United Kingdom)

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.395019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.395019Z digest=sha256:b0cd3ff114c98bcc21830d5758e9cd2174a56e1beac9cba3200d40c57a1db49d

Observation 43983f7f-f4e4-418c-a760-9863f596fdc4 · outbound

This paper cites In Proceedings of the International Conference on Machine Learning.

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings In Proceedings of the International Conference on Machine Learning

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:09.196435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:02:08.423824Z digest=sha256:63aca6a68c9756568a8e13eedca0fdd771308770feb899ad3ab09ce0154d2bd6

Observation 707ce20f-18e7-4cd0-8081-ec627485e9b9 · outbound

This paper cites In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.458000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.458000Z digest=sha256:1b848e076487516b2614cdd7da23859420c47502163a656c40276c394547edf7

Observation 98986562-b5ec-4066-a8e5-78309200ceca · outbound

This paper cites In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:08.368225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:08.368225Z digest=sha256:45a9d35140c92f86618069262c251ede48aeb7776ae845c75c1c68235bb320f9

Pith citing papers

No inbound Pith citation observations are available.