Pith. sign in

Paper Citation Record · LEDGER

Compositional Image Retrieval via Instruction-Aware Contrastive Learning

As of 12 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 4 inbound Pith citation observations for arXiv:2412.05756.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05756 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:28:39.587387Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:36.479575Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:49:57.653860Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact2
  • verified fuzzy26
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe4579b8-7c84-4040-99ec-ea33b2c5b376 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.061159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.061159Z digest=sha256:b5c0694e17fcd3d59627657c88909666798b029716afe37608ee4159818b3dc4

Observation 6b4bc66d-2da8-47f4-91b5-79d34442d652 · outbound

This paper cites GPT-4 Technical Report.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.082738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.082738Z digest=sha256:f453fea3be424c38dcba36ba280a432e696edd473e05b23c989cc331587e5271

Observation 0488177e-37a3-45e0-8c7e-9a0357350780 · outbound

This paper cites iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.125412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.125412Z digest=sha256:29a79f218490a5c762e1f92822e0d45365150dbbd6ebede04ec0c04aa459b8a9

Observation 8c5bf4f5-e147-47aa-980c-2ce4518e9465 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.149769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.149769Z digest=sha256:56e10ca1a22b680b68841645d9ddc9d0512088481d8e040f9ab765a33c1f460d

Observation d909b817-4055-4144-ad10-2425a21afc38 · outbound

This paper cites Zero-shot composed image retrieval with textual inversion, 2023.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Zero-shot composed image retrieval with textual inversion, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:43.980981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.155101Z digest=sha256:a90b7d1af1037bc8302df5289afd355880122a80bd118e5f3277768a5acef98b

Observation ec82d223-912b-4fe8-a361-38d1a665db99 · outbound

This paper cites Zero-shot composed image retrieval with textual inversion.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Zero-shot composed image retrieval with textual inversion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.159922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.159922Z digest=sha256:55967c5d7fd8c5f1f4f330d4f7efe3bb375a921749e7fdb9ea483cd75ad5abe1

Observation ea801c9e-2069-4fea-ba41-0357a6647fa6 · outbound

This paper cites Leveraging large language models for multimodal search.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Leveraging large language models for multimodal search

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:43.866825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.165273Z digest=sha256:8cc70026af6894965fa8c0f5aaa60a726e4fc18f0f6e7f7e0f31f2ce8c653f1a

Observation d9bfedf3-3047-4333-b19d-681570ffc15e · outbound

This paper cites Language Models are Few-Shot Learners.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Language Models are Few-Shot Learners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.170425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.170425Z digest=sha256:891c6021b695000f0d97862f27bc0ba70e23aa455dd24f2c3b3ed5f11cd1bacf

Observation d00ffdd2-f22e-4ffa-ab8c-02c2d6f1dda2 · outbound

This paper cites An Efficient Post-hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image Retrieval.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning An Efficient Post-hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image Retrieval

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:28:40.159886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.176255Z digest=sha256:8542941cfb487630c768c0d230042cd5ae1aff5925a31ed8cf6dd709c93ac086

Observation ba80f392-ff8a-4b30-97bb-038ed2131cd1 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.182222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.182222Z digest=sha256:76c66313660833047102fdcc0f290a7fb1538da487a321bccb8cf4510c59bdac

Observation a4d79d16-9be5-46af-b122-948929647e09 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning A simple framework for contrastive learning of visual representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.188509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.188509Z digest=sha256:8835a2c5729a69b2903d30e642c264861a5f958284a783e1cd874b6da1d7b52a

Observation d320a3c9-d50a-4b6a-b87c-5d377f1eae19 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:43.825169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.193647Z digest=sha256:3a349d049071b93aab83d21d3086fb382eda54ec957c9b609e81c1b4626f2243

Observation 4025ca91-fb83-4e40-959d-aac49698a04c · outbound

This paper cites Scaling instruction- finetuned language models.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Scaling instruction- finetuned language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.198371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.198371Z digest=sha256:9bcb1b59b7232f62909b8eb3be0aab34c20f0e5f1585fbe58e63405e23dcf350

Observation bae50a6f-0b0f-4db8-a8a2-c69365d3a179 · outbound

This paper cites Xtuner: A toolkit for efficiently fine-tuning llm.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Xtuner: A toolkit for efficiently fine-tuning llm

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:43.787846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.202797Z digest=sha256:89a7f5da855224d743892b4eca9a7859777836c1526feebcdfe050f7fc7a54e2

Observation ccc32b6a-e4a1-4fa8-8935-8c7536e4e9bd · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.207262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.207262Z digest=sha256:9de816435fa40a87a39931cb36480eba5d773bfa9c1318556509cfc6853dee2f

Observation e75f1fbf-4f80-4508-bae6-53d8f2d1c453 · outbound

This paper cites Image2Sentence based Asymmetrical Zero-shot Composed Image Retrieval.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Image2Sentence based Asymmetrical Zero-shot Composed Image Retrieval

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.212097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.212097Z digest=sha256:214b6da9272a8ad47728e4050137ad5abdd031572d8bb4d28ca4e68390c75c7a

Observation 5401473b-e092-41a8-beda-0baac29ebf6a · outbound

This paper cites Language-only efficient training of zero- shot composed image retrieval–appendix–.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Language-only efficient training of zero- shot composed image retrieval–appendix–

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:43.745152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.217294Z digest=sha256:25f3c9378c6bf23940bc7a83e580ade66815c0444924dcaad1bdc03125d83bdf

Observation 334a1f4e-2acb-4896-88ba-7c9bc268d5c5 · outbound

This paper cites CompoDiff: Versatile Composed Image Retrieval With Latent Diffusion.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning CompoDiff: Versatile Composed Image Retrieval With Latent Diffusion

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.268822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.268822Z digest=sha256:4a222a7c57e91e7a09da825f262c17b20ac959d6f952afb2da5892c229f707bc

Observation 14399c4d-679c-49c0-bc67-787fcd52882e · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning LoRA: Low-Rank Adaptation of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.390610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.390610Z digest=sha256:a843e5fe600f452e3660799bb68c8a82549ab50682315c835ae3ce1f1aaaddb1

Observation 24075eff-2837-44c3-bff7-dc206592f60b · outbound

This paper cites Spherical Linear Interpolation and Text-Anchoring for Zero-shot Composed Image Retrieval.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Spherical Linear Interpolation and Text-Anchoring for Zero-shot Composed Image Retrieval

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.494207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.494207Z digest=sha256:d2d44766a64be3bab4d4663bcbc483b1d21332e6e33ab8bc18fd25c7fdd58e48

Observation e4488461-2813-439f-b85d-2dd46ef5673f · outbound

This paper cites Visual delta generator with large multi-modal models for semi-supervised composed image retrieval.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Visual delta generator with large multi-modal models for semi-supervised composed image retrieval

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:43.666815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.544882Z digest=sha256:39b1fc150a46c593fb819470f2c5a6e427ff0779af45d2d5de333282e8e7ed1e

Observation da375af5-811b-47ec-aedc-551e13edd096 · outbound

This paper cites Scaling Sentence Embeddings with Large Language Models.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Scaling Sentence Embeddings with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.550195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.550195Z digest=sha256:5c65398d0b3defbf3d1e6de6f24995db8eec9aedbfe9b7fde2ea5267a15afdbf

Observation 0a486c30-e9e5-490c-a9af-1026c69cacb5 · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.554702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.554702Z digest=sha256:f428bcc21fc7473aec1a236dd89a24ea8b7f238398a4176581a868a990aa0dd3

Observation 15e232e7-0357-45ac-a97f-0406eca3b912 · outbound

This paper cites Vision-by-Language for Training-Free Compositional Image Retrieval.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Vision-by-Language for Training-Free Compositional Image Retrieval

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.560509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.560509Z digest=sha256:0feb0bdf2e763f289e231147794653fe358a3716384eb3856d608edd7b647a3c

Observation 1e0fa9cd-c0cf-4ae9-b42d-a5049e1a6a6b · outbound

This paper cites Grounding language models to images for multimodal inputs and outputs.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Grounding language models to images for multimodal inputs and outputs

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:43.643872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.565712Z digest=sha256:c173df7106a355bf97025c01f61f05bd54e5849a0b8923672bb0c0fc74d2d93a

Observation 59cbec90-3ee0-49b6-a1e9-860b5e298810 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.570718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.570718Z digest=sha256:2dc1f6d39795d2c786be5855b4dba5b68e8c056f7c40226a2f8a49433a1059ea

Observation a8e6a41d-8ecd-4639-b354-9fae0fccd4a2 · outbound

This paper cites Improving context understanding in multi- modal large language models via multimodal composition learning.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Improving context understanding in multi- modal large language models via multimodal composition learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:43.599678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.575336Z digest=sha256:6db8ae4d7cedd9a078746b47c41905e85a9fa173e3eab8649aef557492609e05

Observation 8be22f4c-8343-4f4e-a7e4-e46acc24ce45 · outbound

This paper cites Microsoft coco: Common objects in context.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Microsoft coco: Common objects in context

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.580182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.580182Z digest=sha256:e32dd95926b16e3284884348949ab07b29e2f2ae8d0d3ad1acf81f2ec5e28e3a

Observation 28076032-59d4-4bfc-88f5-7972ff5da59b · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Improved baselines with visual instruction tuning, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:43.541650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.584561Z digest=sha256:fbb7df959545e9b27bf4f8299cc15ed0c4ecc06f970c2d270637a9a7ebdec80e

Observation d8bb9fe8-6475-4e19-bed9-d79013b6cdf8 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.589649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.589649Z digest=sha256:806931e43c7d873ef4c7f09b5e42fcba46c371dd7171696e1f57fd706c88a162

Observation 50dcfcfa-b9ef-4a01-945e-619e9309aa5a · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.594557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.594557Z digest=sha256:ec56511bcba37b5429a04ec1e90c54caeb5c0e65f5d5a9da367068ff2a39d3c7

Observation d63ff2e8-ed48-4e5e-8d61-3ffefcaee84a · outbound

This paper cites Visual instruction tuning.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:43.417008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.599029Z digest=sha256:b5b168c8999d36b714283f52ce9819f208eb29330d9f15c9901dcccd2e386ad1

Observation d1c5c96f-d2fc-4b84-aab7-f86ee42cd1dd · outbound

This paper cites Image retrieval on real-life images with pre-trained vision-and-language models.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Image retrieval on real-life images with pre-trained vision-and-language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:43.308028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.704715Z digest=sha256:e6d6a6b908ea224f245f51e0192e30d2bb599c0b8f02fbff3e83205f5fbf9cf6

Observation 043148bf-ef48-4b2f-9a20-7f9362be0d4a · outbound

This paper cites Image retrieval on real-life images with pre- trained vision-and-language models.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Image retrieval on real-life images with pre- trained vision-and-language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:43.178420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.792808Z digest=sha256:63d1220cc5201cd1262fb4cf99a8ab658dc1b134d2dc198cfb7883d2f0f2fd8a

Observation 4d6e0c25-73b0-44ac-a678-d5bd30e6ea41 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Representation Learning with Contrastive Predictive Coding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.869088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.869088Z digest=sha256:ececafb9c62776df3ae1d7b77175bebcf8e39f1aa73e052cb6e599ce31edd8a2

Observation 0dd56536-ed00-4831-8aa2-ab922bd2609b · outbound

This paper cites Training language models to follow instructions with human feedback.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Training language models to follow instructions with human feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.875750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.875750Z digest=sha256:59c0f02bbaee97c629193d766d8a81bf2535b7d538f2d4599a62b07371b144f1

Observation 0e928b77-cd43-4554-9d3d-c9fae0e61cbd · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Learning transferable visual models from natural language supervi- sion

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:38.881850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:38.881850Z digest=sha256:d69c0bddf0bcd27cf00bf0a22da91c68e026892760da32343052db21aaa9727a

Observation 78ca80f0-4fcb-43e3-987a-ea740e5bb187 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Direct preference optimization: Your language model is secretly a reward model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:43.042379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.886480Z digest=sha256:668ef5e075d2b4f19e5bacee3d95784522e3a33fc8a2edc9b7e48f1d09584a24

Observation 5fd95c1c-dd2c-4b52-a9be-087933cc9264 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Zero: Memory optimizations toward training trillion parameter models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:42.975459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.890635Z digest=sha256:82b834f3cd06a20f4216a5a93670d9e9ab66d2e0a14787f6866d71b9c2688198

Observation 7949d080-3ec7-47ad-a565-a4f532d09f9a · outbound

This paper cites Pic2word: Mapping pictures to words for zero-shot composed image retrieval.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Pic2word: Mapping pictures to words for zero-shot composed image retrieval

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:42.957568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:38.982829Z digest=sha256:406d015ed34562be716cc7697709696a0088eb4ba742b9a27bf1e3d3caa6996d

Observation 6041b27e-9247-4c4f-907c-c41c2f929955 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:42.819721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:39.070424Z digest=sha256:7368dfb6a2853c7abac2d2c02b80e56d71a20b613b4ea94750ede546e134bada

Observation 7645cde4-abe7-45c5-a75c-4d30078de11b · outbound

This paper cites ”foil it! find one mismatch between image and language caption”.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning ”foil it! find one mismatch between image and language caption”

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:42.801465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:39.224902Z digest=sha256:325396597bf38d2a7f6dc68a547f7e298588eb08f04d060fef60b7a2ba60da42

Observation 5262a08a-2df5-432f-bcd6-e2392b3808a9 · outbound

This paper cites Knowledge-enhanced dual-stream zero-shot composed im- age retrieval.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Knowledge-enhanced dual-stream zero-shot composed im- age retrieval

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:42.734509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:39.231540Z digest=sha256:72775f556418dba533aba3b3b72fef3c2831b24d4423474b2b7c23b23b2fb422

Observation 3a1f294d-5a88-4a43-a772-119e6c1bea2a · outbound

This paper cites Context-i2w: Mapping images to context-dependent words for accurate zero-shot composed image retrieval.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Context-i2w: Mapping images to context-dependent words for accurate zero-shot composed image retrieval

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:42.450794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:39.237121Z digest=sha256:68e151cb03a2618fffd793684afd39a6c0d797f511965b5ea242229331bedb9c

Observation 4edf9d82-ba3d-4c36-b8c5-f78db82cab71 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:39.243296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:39.243296Z digest=sha256:5b45bdf19e25bd569c59e181167a8f9238929a332d2c38d1e1cdaf74edd27729

Observation ad6ba9f9-fe8e-4ca0-87ea-7324613c8ef4 · outbound

This paper cites Genecis: A benchmark for general conditional image similarity.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Genecis: A benchmark for general conditional image similarity

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:41.889403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:39.248874Z digest=sha256:2fbac058682d6ca877d950729232395eb096d8f143c284b6d578122ad7032142

Observation a5ecbf14-ff90-479b-8c1b-5834f36de698 · outbound

This paper cites Covr: Learning composed video retrieval from web video captions.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Covr: Learning composed video retrieval from web video captions

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:41.524875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:39.254687Z digest=sha256:dd550fd2798ceca28dc79cda990c0c98d4110e9d2dd520632fc5e275bab733a8

Observation aefc1b5b-c1c5-40a8-8961-fb062b484464 · outbound

This paper cites Improving Text Embeddings with Large Language Models.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Improving Text Embeddings with Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:39.259288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:39.259288Z digest=sha256:a76e008cb86f549e66d5f29d1de737e96cbe2c206d5f64319b2fd8c3adc37e91

Observation ad0242fd-a5d9-4639-bf61-2642203c8a32 · outbound

This paper cites Uniir: Training and benchmarking universal multimodal information retriev- ers, 2023.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Uniir: Training and benchmarking universal multimodal information retriev- ers, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:41.376774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:39.265070Z digest=sha256:346174096986a154206e5802cacc024b85df772aac8a8e2120d814fd279ddbcb

Observation 47e0b32b-6939-4a80-81db-5d4f3378a4c5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:41.141585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:39.269655Z digest=sha256:347bb8eb84bd18deff6a1f816a0a9ceb3cb891a12911105983188742b11be5bd

Observation 0589c548-8de5-42f5-8fde-62c8d1523fb8 · outbound

This paper cites The fashion iq dataset: Retrieving images by combining side information and relative natural language feedback.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning The fashion iq dataset: Retrieving images by combining side information and relative natural language feedback

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:40.692496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:39.388157Z digest=sha256:c97454226b9b6b715a214c45ec9ec8f3f1fd44f2e9f9dde2e16744f2425e2fff

Observation 41e306a9-bec9-471c-a9fb-9b6e9da5f6dd · outbound

This paper cites A Survey on Multimodal Large Language Models.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning A Survey on Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:39.473351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:39.473351Z digest=sha256:d809bf1d3e582d61f03b4f7501700c98296e7bf8396d364298f04b3bd54235b1

Observation daa16a05-d6c2-4846-92eb-d10f58f21792 · outbound

This paper cites Attention Prompting on Image for Large Vision-Language Models.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Attention Prompting on Image for Large Vision-Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:28:39.870518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:39.553852Z digest=sha256:f0b42b75b961b2c1f2f51f351514a3aef4978001e504ac73fdd49c3f9d4e2ad0

Observation ef230554-b52f-4ca9-869d-2f6d63fd3f24 · outbound

This paper cites MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:39.559321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:39.559321Z digest=sha256:820f9a9b0ba01b4e9e455e643d8a9b1c224495489fefad0540e0c4ce912e2ea2

Observation f8678ae0-270c-4fbf-9603-87bbd416ba53 · outbound

This paper cites Instruction tuning for large language models: A survey.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Instruction tuning for large language models: A survey

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:39.565645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:39.565645Z digest=sha256:1adcf4cccbb965e9b43f1dd397b16f96ad3ae12b08232d0c3680749a4a619883

Observation b1cfea2a-9079-4ae6-9a05-af8255e62337 · outbound

This paper cites A Survey of Large Language Models.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning A Survey of Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:39.570872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:39.570872Z digest=sha256:2099c1a9f54599393eed54604a4dd8cc52672799df4c263d22f25ce131abc0f9

Observation ba48bfc9-ffcf-4282-b8da-3fd151b706f8 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T20:28:39.576272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:28:39.576272Z digest=sha256:4433e06f0b338768f6e8346a401f91ad468f95d22d39e114fa4dbac75b1252f2

Observation 4fbdb120-1a1c-4708-bc82-48dc3dbc8dc6 · outbound

This paper cites a very typical bus station.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning a very typical bus station

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:40.547856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:39.580737Z digest=sha256:da40f123ed4d475ca844072c2519f3e754eff748bffd7d23a15bebcfc545d250

Observation 1d224a39-7b9b-48cd-a60f-ed5d5e171257 · outbound

This paper cites Table 8 shows the sizes of training datasets.

Compositional Image Retrieval via Instruction-Aware Contrastive Learning Table 8 shows the sizes of training datasets

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:28:40.386621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T20:28:39.587387Z digest=sha256:c75c784328baef4a7d6b3c851abaf4b4c096398db455fb278543448a8dbeed52

Pith citing papers

Observation fe1fac11-dc12-4666-acb2-990a031842ad · inbound

MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval cites this paper.

MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval Compositional Image Retrieval via Instruction-Aware Contrastive Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:36.479575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:36.479575Z digest=sha256:272dd11e1ac723c8a1d4018fa19113112172e62a7a980bd50d347cc5038b3a9c

Observation 7159ee2e-b43d-4a3d-a704-81725a921591 · inbound

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval cites this paper.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Compositional Image Retrieval via Instruction-Aware Contrastive Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.059408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.059408Z digest=sha256:a4e8029430b7aa800e8f1e0f5caece2a305e88a78b2e03abbad427e09c6bfc81

Observation f0d5ce91-7dd8-4804-9186-84cdc4b9e03a · inbound

Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering cites this paper.

Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering Compositional Image Retrieval via Instruction-Aware Contrastive Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:58.111382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:28:50.127362Z digest=sha256:83f5d41433f0c98edb5498ea1d16019902e55b6167f2b6177c7feed2d2c89f00

Observation 12930f30-d7d0-482c-ab7b-941ac426ffdd · inbound

Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent cites this paper.

Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent Compositional Image Retrieval via Instruction-Aware Contrastive Learning

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:49:57.656128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T01:19:19.253076Z digest=sha256:9c240119fda825285a4a9208d210cf640969760f77361164dee5b299c816bae0