Pith. sign in

Paper Citation Record · LEDGER

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

As of 9 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 7 inbound Pith citation observations for arXiv:2505.19952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19952 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:07:40.730204Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:21:10.770726Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:05:41.143246Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact4
  • verified fuzzy15
  • unresolved36
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f6c19bf-fa02-4ae7-8ad0-fd89ac999fa8 · outbound

This paper cites Distribution consistency guided hashing for cross-modal retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Distribution consistency guided hashing for cross-modal retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:34.933454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:34.933454Z digest=sha256:78d74a07083b0906328ce95e29af4d13d121e7426628e281db7c795430a5522d

Observation 538ab951-db2c-40b6-b51a-8f47eef29e71 · outbound

This paper cites Similarity transitivity broken-aware multi-modal hashing.IEEE Trans.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Similarity transitivity broken-aware multi-modal hashing.IEEE Trans

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.023106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.023106Z digest=sha256:283326a54bd355aac4614ce3c3a6a10f9c7d20f21a5a05ea127dd1bc6c5d35a5

Observation 9d9f5b31-c484-4436-8dec-1c341b1491cd · outbound

This paper cites Data-aware proxy hashing for cross-modal retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Data-aware proxy hashing for cross-modal retrieval

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.062654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.062654Z digest=sha256:b3885635865f9c96af14a0e0a7d1979b4337281f7f48be9f6b7aff158aeae6d6

Observation 51b2802e-9064-479a-9e56-6f461d220588 · outbound

This paper cites Where does the performance improvement come from?: - A reproducibility concern about image-text retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Where does the performance improvement come from?: - A reproducibility concern about image-text retrieval

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.168550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.168550Z digest=sha256:8a364c8d398732b8569e0283f97c6b17f76a6d2445403eae90babb0598df2aa0

Observation a8855576-087d-471a-9ace-60361e5e8f85 · outbound

This paper cites Com- posing text and image for image retrieval - an empirical odyssey.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Com- posing text and image for image retrieval - an empirical odyssey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.291206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.291206Z digest=sha256:ad403bca3b4d101244afd14561047b64d5a52d7d669733e546ccd64b8a0ec2f8

Observation 8c94eda2-0b0d-4e67-b969-4e28046fa20d · outbound

This paper cites Sim- ple but effective raw-data level multimodal fusion for composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Sim- ple but effective raw-data level multimodal fusion for composed image retrieval

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:47.951301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:35.377153Z digest=sha256:dbb8c73e3524cf5aac36064433d5a2aff8cb6a5457f84fd023ac11ded07783af

Observation 1754bc0b-4231-47f3-bbb5-5a5c2e5a0b7b · outbound

This paper cites Dual-path semantic construction network for composed query-based image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Dual-path semantic construction network for composed query-based image retrieval

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.561828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.561828Z digest=sha256:c6f0f7168c7c9a1becfb6186f99dceef6421b4dd6d834e967f9957ea1b46fcd3

Observation ed7928b0-7826-495d-aec2-c096137479c1 · outbound

This paper cites Multi-modal transformer with global-local alignment for composed query image retrieval.IEEE Trans.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Multi-modal transformer with global-local alignment for composed query image retrieval.IEEE Trans

Reference 8

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:07:43.016620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:35.694647Z digest=sha256:ea4bb53f7f42e0008efead3bc4b85ba19e3088d5df5cc4a4d6aee0feae4b10a0

Observation 301e70e5-0c60-4b16-bd29-1a0d622023ad · outbound

This paper cites Sentence-level Prompts Benefit Composed Image Retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Sentence-level Prompts Benefit Composed Image Retrieval

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.796615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.796615Z digest=sha256:09f147ca6f84e4509fbbcee4a747859c58fd55e754d841c83d43cfe376c4e68b

Observation fbb65504-4c99-47e3-bd34-736153cc65b2 · outbound

This paper cites Image Search with Text Feedback by Additive Attention Compositional Learning.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Image Search with Text Feedback by Additive Attention Compositional Learning

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:07:41.386495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:35.903005Z digest=sha256:43dcf04652a8597e0d6c06392b1c2dc3c38070dd2c8ba85d0d0aebaeb8c950e4

Observation 30f596e3-4083-4d51-b197-01f032f122e8 · outbound

This paper cites Image retrieval on real-life images with pre-trained vision-and-language models.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Image retrieval on real-life images with pre-trained vision-and-language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.986269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.986269Z digest=sha256:8f77538d5ea996218d559c291b843bc38c4bb653a7c5b3ffe82960d6189cbc4c

Observation 2f231df7-737b-47a7-aa08-b5b9b016f965 · outbound

This paper cites Composed image retrieval using contrastive learning and task-oriented clip-based features.ACM Trans.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Composed image retrieval using contrastive learning and task-oriented clip-based features.ACM Trans

Reference 12

Resolution
verified exact
doi, observed 2026-08-07T14:07:41.240420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:36.053819Z digest=sha256:531d7bb916c1de664a5099e929eca581967a2e30cc9f06857f23b18a8e4f0cf7

Observation 7dd376e6-6da1-4e3d-a81c-4eabcdadfd49 · outbound

This paper cites Fine- grained textual inversion network for zero-shot composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Fine- grained textual inversion network for zero-shot composed image retrieval

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:36.160343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:36.160343Z digest=sha256:76d5ba7eb2e5c92d44a22fccaaaa2791c837446ccae8abf45d365f04d246a77f

Observation 9270e1f6-a011-4e8f-9a48-b3e09ed9b047 · outbound

This paper cites Pic2word: Mapping pictures to words for zero-shot composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Pic2word: Mapping pictures to words for zero-shot composed image retrieval

Reference 14

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:07:42.601773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:36.254019Z digest=sha256:c85c8dbef635d0d0bbc0561afa8ba9bda375f2a77e4d178535ebbfa7f70fc305

Observation 94655246-9d83-41c3-8d1f-951209eb33a1 · outbound

This paper cites MLLM-I2W: harnessing multimodal large language model for zero-shot composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval MLLM-I2W: harnessing multimodal large language model for zero-shot composed image retrieval

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:47.724543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:36.416601Z digest=sha256:436c782f365fe675bfaf9c0335eaec2509ad0a26af27f22d7f0d99072a814be1

Observation 8cd08409-8037-41ba-8f94-b7fa225194d5 · outbound

This paper cites Kankanhalli.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Kankanhalli

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:47.508680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:36.510398Z digest=sha256:2b44c3153288d0b539a530ab1700c71e109e891852ba1625ceff7b67abcbed37

Observation 7c16a5f1-68a2-4fd6-9e75-b18ac824a136 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Learning transferable visual models from natural language supervi- sion

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:47.274574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:36.588283Z digest=sha256:8de1c299c8204094088835dd8e5495a61203c110ff312941326f44e144e9116f

Observation eb65907c-f8e9-45b5-bd4e-372cc5f342da · outbound

This paper cites an unresolved cited work.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:07:47.063818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:36.710310Z digest=sha256:14cbdd775cf634a6b0c51c0a8ddc490439563a9d3b55dbaa41af98f05b75de9b

Observation d28659d0-78ad-4ad8-86c1-077abbf24ef5 · outbound

This paper cites Seeing what you miss: Vision-language pre-training with semantic completion learning.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Seeing what you miss: Vision-language pre-training with semantic completion learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:36.788327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:36.788327Z digest=sha256:d1d8a9b43745eac183e5fd3ab55dd17bed8921055601dc9ca999b7f9ba4604b6

Observation a939717c-1883-4fa2-99d7-4d71ecaf91b4 · outbound

This paper cites Global and Local Semantic Completion Learning for Vision-Language Pre-training.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Global and Local Semantic Completion Learning for Vision-Language Pre-training

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:07:41.095133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:36.870380Z digest=sha256:5e5e130f9c42b9df0cc6afa7b210b852eceafc9f4baf1f078063abb61866b893

Observation 22b2247f-6e88-4a19-8f53-2fe7d56eb5dc · outbound

This paper cites Zero-shot composed text- image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Zero-shot composed text- image retrieval

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:46.822771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:36.972948Z digest=sha256:4527a290e7e6368c49db9de7fc8c3b968d015e38439fb5964d7957173d36b0d8

Observation 7159ee2e-b43d-4a3d-a704-81725a921591 · outbound

This paper cites Compositional Image Retrieval via Instruction-Aware Contrastive Learning.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Compositional Image Retrieval via Instruction-Aware Contrastive Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.059408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.059408Z digest=sha256:c681f4fbd82caea883cadcf482d6730f1a67bcf2e511e2caeb6a46b31bb8a8ed

Observation 6cf8b24b-55bf-48d4-a2c2-8e94cd17290c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval LLaMA: Open and Efficient Foundation Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.160966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.160966Z digest=sha256:45d3d7e8c519edcea10625c6d292113957a02083f2872add3656295aee13cd27

Observation 2f83fe13-f59a-4027-9956-1c6ef3d9215c · outbound

This paper cites Large Language Model Agent: A Survey on Methodology, Applications and Challenges.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Large Language Model Agent: A Survey on Methodology, Applications and Challenges

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.240894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.240894Z digest=sha256:ad49e8f5b9ee547226deb49c48b767041dd84d881bc172b8ab0a49bc24b66fa8

Observation c023c4a0-12fb-4a6f-afc6-8b68a27e560a · outbound

This paper cites Visual instruction tuning.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Visual instruction tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:46.516968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:37.359497Z digest=sha256:51f248ed589655cc09de775391e35fac2203ff5a9b037e85807a7e1cc155628e

Observation 4faaafec-6882-45dc-89c9-338a09a643c9 · outbound

This paper cites SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.487699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.487699Z digest=sha256:a5b4be186d5389c03f1de440ac6be1adcb464dfc1755ea0e2c8ae2d9d8e3fea8

Observation b74f79d4-0c60-463c-8822-0f8326ab012f · outbound

This paper cites Grounding language models to images for multimodal inputs and outputs.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Grounding language models to images for multimodal inputs and outputs

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:46.295587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:37.592129Z digest=sha256:10668670b4b987ccec5a6e75b989ef4c821f0935842840ef7efdc3c0286b0377

Observation 7b1ae7af-ba40-4db3-b761-89f4dc576f85 · outbound

This paper cites an unresolved cited work.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:07:46.003833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:37.696524Z digest=sha256:e9f09dd9fd0147841029c70e49ea735cabd9a32484c2655f8d0c3ac03a719fee

Observation 11038ea0-4bce-4a62-bb5f-b711dbd1a21b · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.755205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.755205Z digest=sha256:17ab2a36391d8780e42c694b4fde571a0916cfd728ad82efa50a45d9fa50014e

Observation 61e154bc-4477-4ca2-a4e9-eed858a478dc · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Representation Learning with Contrastive Predictive Coding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.919492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.919492Z digest=sha256:17cdd7b1d0551238b32a7843f7b8fc892e88e6e34685f1c2cce53078b844df7d

Observation 84835a9a-54c1-4b11-9ecc-8ea290504b13 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:37.813610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:37.813610Z digest=sha256:a7341b0c3d49e750a50fac24ffe673ed101ce232ccc9acb873931cfb837d3e66

Observation 8cf5a0ca-7cae-4890-9b59-9ebab07e4ae0 · outbound

This paper cites Weighted gaussian loss based hamming hashing.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Weighted gaussian loss based hamming hashing

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:38.165615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:38.165615Z digest=sha256:2d2c47115a50832efc7dad99832d4b7f0f872dd685909092c5d8394c5240eeaa

Observation 020557ce-a9b2-4711-bb62-d7b1292c9eea · outbound

This paper cites Unsupervised hashing with semantic concept mining.Proc.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Unsupervised hashing with semantic concept mining.Proc

Reference 33

Resolution
verified exact
doi, observed 2026-08-07T14:07:40.884160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:38.080713Z digest=sha256:f4a0bfaae1c559d2eeda40df92f772227be0f7c78eba1bffa578dbf7438e7985

Observation 1b6d306c-3e34-41eb-aed3-f31e3e77bd37 · outbound

This paper cites Unsupervised cross-modal hashing with modality-interaction.IEEE Trans.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Unsupervised cross-modal hashing with modality-interaction.IEEE Trans

Reference 34

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:07:42.189094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:38.379544Z digest=sha256:b8086dbf226fd2d6593a14948f4bab0a20b1610cc1a35d74eabd8bd97a687f89

Observation d416cac1-f94a-4443-bf6e-f460ef08e3ee · outbound

This paper cites Partial-softmax loss based deep hashing.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Partial-softmax loss based deep hashing

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:38.290845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:38.290845Z digest=sha256:69d41e57390ccaafef7480b1b6f46543008e461d5b77f49a89f569c5071b5b92

Observation c7699576-e266-4c04-9ef3-62c61a9746f4 · outbound

This paper cites Unsupervised cross-modal hashing via semantic text mining.IEEE Trans.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Unsupervised cross-modal hashing via semantic text mining.IEEE Trans

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:38.566628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:38.566628Z digest=sha256:49fc71bd96c034833aec42212394b33e8788c62b8c59e1f6b215f9cd85858ff0

Observation af33825f-e7e6-4eb9-abe4-0863552461e6 · outbound

This paper cites Deep cross-modal proxy hashing.IEEE Trans.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Deep cross-modal proxy hashing.IEEE Trans

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:38.471242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:38.471242Z digest=sha256:25d6aa4888a605f929a83f18bc90fccd4608d0be5356aba477ea2f8af9f634b5

Observation a962b845-f363-4556-9c1b-11339166c34e · outbound

This paper cites Knowledge-enhanced dual-stream zero- shot composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Knowledge-enhanced dual-stream zero- shot composed image retrieval

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:45.438771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:38.893398Z digest=sha256:88b3c643f312621f3341184afeed20a152d9b197b60c70b0d19c331cfb3062ba

Observation 4abf15b5-e33a-4f23-a3ba-adcf8076cb5d · outbound

This paper cites Zero-shot composed image retrieval with textual inversion.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Zero-shot composed image retrieval with textual inversion

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:45.689940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:38.676505Z digest=sha256:92baed96a66f73eed10f2d6acfa0ab73f8d51dcd30b9e8e07dc3ff9fb573c785

Observation 9fbf06e4-991f-4bd4-a342-b94d9b5b34cd · outbound

This paper cites Target-guided composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Target-guided composed image retrieval

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:45.186347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:39.258733Z digest=sha256:2c8c6d50d0bdea2d12c97c2c467bc2dc5f0e0ab8c3e04d1a5ddd5d2b71054466

Observation d8838646-f428-4d69-b99a-625bd9726bce · outbound

This paper cites Enhance composed image retrieval via multi-level collaborative localization and semantic activeness perception.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Enhance composed image retrieval via multi-level collaborative localization and semantic activeness perception

Reference 43

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:07:41.699024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:39.075665Z digest=sha256:a935bcc317f9d4f2699875c883e4a635a0764fdda670cf5dbb4b59c4468b20a1

Observation 4d0facfa-3bef-45b5-a616-90a4e93fa75f · outbound

This paper cites Data roaming and quality assessment for composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Data roaming and quality assessment for composed image retrieval

Reference 44

Resolution
malformed identifier
no resolver link, observed 2026-08-07T14:07:39.516160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:39.516160Z digest=sha256:49de2e06383713101ba8680b258b6f578783bc0bcac3571146c31e0ea50e3360

Observation 3cc6934b-9fde-4ca2-b134-d7dec363c68b · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:39.618249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:39.618249Z digest=sha256:29adfedfb56abab4a25d37c92ba8fd9d14a10fd3e1f99bf0d5935b44a3cf26a8

Observation 3d1f6201-d720-47ac-bb11-baf71655fb1f · outbound

This paper cites Self- training boosted multi-factor matching network for composed image retrieval.IEEE Trans.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Self- training boosted multi-factor matching network for composed image retrieval.IEEE Trans

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:44.969022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:39.359562Z digest=sha256:5fb5144b87ba12195121274482e63b8584fd998021fca351e1894ef2e4e9d154

Observation ff304f35-27d3-410d-b920-5ca08d8da2b6 · outbound

This paper cites Context- i2w: Mapping images to context-dependent words for accurate zero-shot composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Context- i2w: Mapping images to context-dependent words for accurate zero-shot composed image retrieval

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:39.449109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:39.449109Z digest=sha256:6ed397e9d4a5c0fb0b7754994e250e70ad9d0218e58561403fffde0a83696837

Observation ca4a488a-0f0f-4da9-bf66-635be1128220 · outbound

This paper cites Fashion IQ: A new dataset towards retrieving im- ages by natural language feedback.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Fashion IQ: A new dataset towards retrieving im- ages by natural language feedback

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:39.932800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:39.932800Z digest=sha256:f04ed2fe2314030c9515908b33f81fd4345e1c2173d0f32197c73525bdb9e023

Observation b2fb9b68-3761-4a0e-b362-2c2ddd5034b5 · outbound

This paper cites A corpus for reasoning about natural language grounded in photographs.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval A corpus for reasoning about natural language grounded in photographs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:40.096640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:40.096640Z digest=sha256:7c6429543d446fd12789d0dbf62f3c7bc7b442a35bede146ec7db875c7ee95c5

Observation 4129a7a0-e2b1-42e7-b0d3-67e399357e06 · outbound

This paper cites Mini-batch optimization of contrastive loss.Transactions on Machine Learning Research, 2024.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Mini-batch optimization of contrastive loss.Transactions on Machine Learning Research, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:44.691636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:39.674668Z digest=sha256:dde610c5a29d1780ec4a743967150f97c2b39aad30f87be1d3c4983ab4446293

Observation aa12343d-fb2d-48de-8a7a-c69161ef55b3 · outbound

This paper cites Bridging mini-batch and asymptotic analysis in contrastive learning: From infoNCE to kernel-based losses.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Bridging mini-batch and asymptotic analysis in contrastive learning: From infoNCE to kernel-based losses

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:44.425859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:39.786720Z digest=sha256:a8c2e07b33f2a862b11923acd354738767f2bdc7a189da428bb14576b93423d4

Observation 60c97a70-de33-4cc0-9da2-9ccaf7a8c04e · outbound

This paper cites an unresolved cited work.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:07:44.117488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:39.862668Z digest=sha256:54098e353125f1f3d72083eae7c98a510c412f8011a25bdfbdaf53afab3cbfad

Observation ef43d3dc-2c58-44b9-951a-117d568cacea · outbound

This paper cites Language- only efficient training of zero-shot composed image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Language- only efficient training of zero-shot composed image retrieval

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:40.428606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:40.428606Z digest=sha256:88931f8a2519460efb94cf899b4c45f97cbfbe0719138aa91be2fb8c9498b40a

Observation 9cc13558-3ba6-494a-8a96-f7ec2024f060 · outbound

This paper cites iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:40.499510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:40.499510Z digest=sha256:211c3adccb780ec2fe9a89ca4dfe6bd4f4492cd3b378f53b9c06a2d549b5ad85

Observation 60ccd460-c987-4c18-b7f0-ae258f6f9267 · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:40.153180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:40.153180Z digest=sha256:ca09a120f904eca338296e0b7cbb1d5091ca778c811f0017eb397eeac274931c

Observation 7a072da1-1de3-4991-afd7-b8a93b41041c · outbound

This paper cites Bernstein, Alexander C.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Bernstein, Alexander C

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:40.228472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:40.228472Z digest=sha256:50ecf55d60a3268f9ea1a29ecedb07f082210f6924c2ca85668b9c50ae74f771

Observation 1891be75-4b69-4f6a-8173-5eb5de5d6d3a · outbound

This paper cites Decoupled weight decay regularization.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Decoupled weight decay regularization

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:40.384950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:40.384950Z digest=sha256:67a834fb934468475f9609deb0353cdb9ef707f8c1001962bfad65c0bbdbd61d

Observation 1f139446-d555-4488-8c72-cb646f2a48bc · outbound

This paper cites Vision-by- language for training-free compositional image retrieval.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Vision-by- language for training-free compositional image retrieval

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:43.833760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:40.575614Z digest=sha256:b1c686c341c83f4d9e754d2a6e562acd2b166793f5fe3483268afb595eb1c4aa

Observation 95c7083a-1aa8-49db-9d8c-7bd100e3bc53 · outbound

This paper cites Effective conditioned and composed image retrieval combining clip-based features.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval Effective conditioned and composed image retrieval combining clip-based features

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:40.687358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:40.687358Z digest=sha256:ba33a71e338d12e4d11512b7056e5e37a73dfc3eb3f0cfa982d6db07b2e48af5

Observation 605c0739-24a8-4707-8b81-9bf4aab364f2 · outbound

This paper cites cap1" and image 2 with the caption.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval cap1" and image 2 with the caption

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:07:43.504874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:07:40.730204Z digest=sha256:8e4b64d1ec3d7f52da63c05218f885bacf09b6d2a15a7f188e2de97948348803

Observation 05e1be0d-d4c5-4bda-a1e3-ce88948d567a · outbound

This paper cites URL https://doi.org/10.1109/ICCV51070.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval URL https://doi.org/10.1109/ICCV51070

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:38.798145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:38.798145Z digest=sha256:1e68f440c79846e5a369489de926347956fb3333284eebe8958eb105f77b556e

Observation bbf5ee11-3fd1-4061-a73d-4b5b733e86cb · outbound

This paper cites URL https://doi.org/10.1145/3626772.3657727.

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval URL https://doi.org/10.1145/3626772.3657727

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:35.458335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:35.458335Z digest=sha256:08fc036a9f1e957dd13ca1b47d55668b6bec288841eb717b3388fb983fa7c7cb

Pith citing papers

Observation e452ef4b-739f-4f84-ad68-a889a2b58f50 · inbound

SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding cites this paper.

SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.770726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.770726Z digest=sha256:22cb11981937c3393c18aaafd0e4c57aa9970e62c3159fd29506ba7ffc683fac

Observation 9b1c5e39-0d5b-436d-98fc-ec0b8f51f774 · inbound

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories cites this paper.

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T01:02:35.070287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:02:35.070287Z digest=sha256:74f20f062d1472e6203d8dcc89bb1343a9472fff33cb1a8f915da82ad895d666

Observation 16a6f781-dec7-417d-ae33-b5f9e59733cb · inbound

CoVR-R:Reason-Aware Composed Video Retrieval cites this paper.

CoVR-R:Reason-Aware Composed Video Retrieval Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T21:37:55.887477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:37:55.887477Z digest=sha256:9550c7fc2dbd3b3018f30a44b7f6f84b09552bbca21bf87e7b3d0fb661d4ad6e

Observation 922148f1-64cc-45c3-8c8f-05fe704e3ba4 · inbound

DeliCIR: Memory-Guided Test-Time Deliberation via Multi-Agent Collaboration for Composed Image Retrieval cites this paper.

DeliCIR: Memory-Guided Test-Time Deliberation via Multi-Agent Collaboration for Composed Image Retrieval Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:41:10.257851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T06:40:58.445504Z digest=sha256:b0c082f5816ae0aba2d538ff53fc68f3cc9abbe73ea6735559dae20e872327bf

Observation 49782fa1-ea49-4103-a5ea-995e68bdb437 · inbound

DeliCIR: Memory-Guided Test-Time Deliberation via Multi-Agent Collaboration for Composed Image Retrieval cites this paper.

DeliCIR: Memory-Guided Test-Time Deliberation via Multi-Agent Collaboration for Composed Image Retrieval Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:14:56.878699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:10:54.790735Z digest=sha256:ed8475ca9bffa495f7b34dc43ce7b9413e7c7aaf349a595293334bbedbe07dbd

Observation 6d51e281-51a6-4bac-bd03-c5e8505fdf78 · inbound

DeliCIR: Memory-Guided Test-Time Deliberation via Multi-Agent Collaboration for Composed Image Retrieval cites this paper.

DeliCIR: Memory-Guided Test-Time Deliberation via Multi-Agent Collaboration for Composed Image Retrieval Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T13:28:19.474650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:28:19.474650Z digest=sha256:f6835e040e91f529197415e0a2eaef2177aed70e9c0738acfeb6a73ebf091399

Observation de0291f8-879f-4383-a4d6-69e97783b87b · inbound

Thinking Before Retrieving: Robust Zero-Shot Composed Image Retrieval via Strategic Planning and Self-Criticism cites this paper.

Thinking Before Retrieving: Robust Zero-Shot Composed Image Retrieval via Strategic Planning and Self-Criticism Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:05:41.145002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T05:52:20.768497Z digest=sha256:f22faea5095b41528460a22e136f234eb89fb4495d3273a6b792e930fd0f6ac6