Pith. sign in

Paper Citation Record · LEDGER

UniCoRN: Unified Commented Retrieval Network with LMMs

As of 10 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 0 inbound Pith citation observations for arXiv:2502.08254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08254 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:54:16.148845Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

99 of 99 outbound references displayed

  • verified exact2
  • verified fuzzy38
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38d6ddc2-c253-4ba6-a33a-97550b330624 · outbound

This paper cites Pixtral 12B.

UniCoRN: Unified Commented Retrieval Network with LMMs Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.808010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.808010Z digest=sha256:ac634fbc0239fb576ae91051db487fd999a9df9020b068abce8dc3eedd38f1eb

Observation c558e5bf-2f8e-45f7-b00a-23bdc8214473 · outbound

This paper cites Learning attribute representations with local- ization for flexible fashion search.

UniCoRN: Unified Commented Retrieval Network with LMMs Learning attribute representations with local- ization for flexible fashion search

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.812393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.812393Z digest=sha256:6eca88f4284d4f832d4c57e6935b6a64322b9108e6408e1305847acd1a438bf3

Observation 2a2b8b81-bc1c-4b50-9095-641d36a52911 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Bottom-up and top-down attention for image captioning and visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.815852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.815852Z digest=sha256:e3f10e2c3edc3ff1a89003b5c38a3b002854fac344501854f34f99deb537e4de

Observation b52e3554-78f2-4aa7-a5e2-19a3cc08e183 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

UniCoRN: Unified Commented Retrieval Network with LMMs The claude 3 model family: Opus, sonnet, haiku

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.819291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.819291Z digest=sha256:356e80e39a497974d5ef1dbbfc0f697a1fe637b10ca5c06b0cfc71ba3eb3788a

Observation 992f8b97-ed79-42cf-be80-18f80ad192f2 · outbound

This paper cites Vqa: Visual question answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Vqa: Visual question answering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.822811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.822811Z digest=sha256:d961c5697644443435097b9994c0c9b7f0cb08249181477b5daf8f5e84bdce23

Observation d9590b7f-e4e7-4517-aa01-fb2b35046b20 · outbound

This paper cites Effective conditioned and composed im- age retrieval combining clip-based features.

UniCoRN: Unified Commented Retrieval Network with LMMs Effective conditioned and composed im- age retrieval combining clip-based features

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.826342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.826342Z digest=sha256:5a4fc8de92b9deb124c213bf256926767675c68c77168c5c45e39bbffffa2e67

Observation 8034d24e-c299-4e48-9ba7-8187cac9e97b · outbound

This paper cites Zero-shot composed image retrieval with textual inversion.

UniCoRN: Unified Commented Retrieval Network with LMMs Zero-shot composed image retrieval with textual inversion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.829572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.829572Z digest=sha256:f76a7919681c3b8bb27cc376f8911d5c1ba750e6e078a94e4e6d914e751d1767

Observation c2c2326d-a5c4-4552-bb9f-fcde8a26abd2 · outbound

This paper cites Vlmo: Unified vision-language pre-training with mixture-of-modality-experts.

UniCoRN: Unified Commented Retrieval Network with LMMs Vlmo: Unified vision-language pre-training with mixture-of-modality-experts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.833054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.833054Z digest=sha256:d571582e9dcb0b199780c69fe7da9b86ecb24f6c5c2fc56910990be8d8d04152

Observation 06e54353-c76a-4c4a-a521-e9f159eade34 · outbound

This paper cites Tomayto, tomahto.

UniCoRN: Unified Commented Retrieval Network with LMMs Tomayto, tomahto

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.836907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.836907Z digest=sha256:102169d6863ed44155ac8aedb507bb769da227f273818329de1b20ee4ecf1f8d

Observation d4a09c30-4cdf-40c6-a317-5863e240dfcf · outbound

This paper cites A Suite of Generative Tasks for Multi-Level Multimodal Webpage Understanding.

UniCoRN: Unified Commented Retrieval Network with LMMs A Suite of Generative Tasks for Multi-Level Multimodal Webpage Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.840236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.840236Z digest=sha256:9bf1c43412e7f1123280028ea63b3363b8bdc904b16219594c331c981d4220f9

Observation bbd4ef3c-0289-4215-9ebd-abf83f7d2828 · outbound

This paper cites Plummer, Kate Saenko, Jianmo Ni, and Mandy Guo.

UniCoRN: Unified Commented Retrieval Network with LMMs Plummer, Kate Saenko, Jianmo Ni, and Mandy Guo

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.843793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.843793Z digest=sha256:0e55d8381c24533ef20d8257f07e75a2254fa232e3f964a704ae23dafb3331d0

Observation a669727d-692c-4509-ba45-bcbbc5839f3d · outbound

This paper cites Wiki-llava: Hierarchical retrieval-augmented gener- ation for multimodal llms.

UniCoRN: Unified Commented Retrieval Network with LMMs Wiki-llava: Hierarchical retrieval-augmented gener- ation for multimodal llms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.847113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.847113Z digest=sha256:9e5201abcb69aed73a621630fb65d6ff0a93d04120784183b1f2e7ce99a2e8fa

Observation 9a0755fe-5633-45ce-a0dc-c5be88c1260e · outbound

This paper cites MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text.

UniCoRN: Unified Commented Retrieval Network with LMMs MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.850481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.850481Z digest=sha256:8ab5685a24e4c85b34255a635cba6218faff104abdc3c0a69afdfd475c998779

Observation 1369fc8e-ca41-41cd-ab5d-3da27cc877e4 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

UniCoRN: Unified Commented Retrieval Network with LMMs PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.853944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.853944Z digest=sha256:521e94ee42f7ed72b10e4c8516ce35fdeb86e1372c3acec05d2803b79856acb3

Observation bed027d4-31ba-4427-b29e-619051466185 · outbound

This paper cites Image search with text feedback by visiolinguistic attention learn- ing.

UniCoRN: Unified Commented Retrieval Network with LMMs Image search with text feedback by visiolinguistic attention learn- ing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.857665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.857665Z digest=sha256:64bd5dbb69941454eb705ee1dd72bb0340a1b1c6330625763300b22a1fbb6229

Observation 3add3c2b-6dd9-44a2-b881-b45c5638c29a · outbound

This paper cites Image search with text feedback by visiolinguistic attention learn- ing.

UniCoRN: Unified Commented Retrieval Network with LMMs Image search with text feedback by visiolinguistic attention learn- ing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.860861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.860861Z digest=sha256:5eef0edb3788b3784611e4e1a55f7ec1ceaa6b1af722eeb5072e0211a63f3821

Observation 3109636d-67a5-415a-baff-0e63cb21a129 · outbound

This paper cites Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?.

UniCoRN: Unified Commented Retrieval Network with LMMs Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.864357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.864357Z digest=sha256:a75c72205e7f99dfb252051054ab5f4b77d5ecb6dab38fb1bdd15ac2f62cead6

Observation 4ef45408-0c61-4f29-abdb-4b6c0de8ec1d · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open- source suites, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs How far are we to gpt-4v? closing the gap to commercial multimodal models with open- source suites, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.868000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.868000Z digest=sha256:3df69e94ab139b04b9c9eaaf52690c559516cebebc6f2abb6df3144ca3036a02

Observation 425d9efd-8881-43fa-b621-d4a2072ea104 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

UniCoRN: Unified Commented Retrieval Network with LMMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.871260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.871260Z digest=sha256:44f017a2a9a1fa82d44c3f6bf5f0dd76cd5df32a9eaa3d95abb90001810c2922

Observation bafc4799-8cb8-4f13-a03e-a855dd5c3956 · outbound

This paper cites Meteor universal: Lan- guage specific translation evaluation for any target language.

UniCoRN: Unified Commented Retrieval Network with LMMs Meteor universal: Lan- guage specific translation evaluation for any target language

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.874549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.874549Z digest=sha256:4585c1b32cb7c4070a4e9bcbc17eda4b6d25842094650ae044a45298e44d240f

Observation fcec966b-795a-4e81-85ad-db0c89a83a69 · outbound

This paper cites Hyper- bolic image-text representations.

UniCoRN: Unified Commented Retrieval Network with LMMs Hyper- bolic image-text representations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.877970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.877970Z digest=sha256:ecb714f4165b4400f69acfc1771ec64456d644413fa7558bc8983b1c2900096a

Observation 5ff08aba-4c39-439f-b204-4a38fe95f237 · outbound

This paper cites Toutanova.

UniCoRN: Unified Commented Retrieval Network with LMMs Toutanova

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.881162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.881162Z digest=sha256:4c5b00c81fffc92f50fb7ec3d17c2d65e96e97d6f3f8996e724473782c2a0c0a

Observation 1073d162-89b2-4794-983c-09792b776666 · outbound

This paper cites The Llama 3 Herd of Models.

UniCoRN: Unified Commented Retrieval Network with LMMs The Llama 3 Herd of Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.884750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.884750Z digest=sha256:7b14d6abaf5c51f13f6621017954d38e7ab27c7941e4fd6dabe72ff66842f67c

Observation 16874e68-f265-4254-b4cf-48e4ff28f391 · outbound

This paper cites Entities as Experts: Sparse Memory Access with Entity Supervision.

UniCoRN: Unified Commented Retrieval Network with LMMs Entities as Experts: Sparse Memory Access with Entity Supervision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.888711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.888711Z digest=sha256:9cada7af7c9abf0a30342703a818c194853700ac21fac66375d55ee0d530c7dd

Observation 033cdd40-4ff8-48c0-8b12-7f28458d3ef7 · outbound

This paper cites Pyramidclip: Hierarchi- cal feature alignment for vision-language model pretrain- ing.

UniCoRN: Unified Commented Retrieval Network with LMMs Pyramidclip: Hierarchi- cal feature alignment for vision-language model pretrain- ing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.892674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.892674Z digest=sha256:040b094e1d181aa736cb3eb85ffe9915344d31cd41ba3276625a1251e459e2da

Observation dc07771b-3808-4fa9-8b0a-7919c7744148 · outbound

This paper cites SoftCLIP: Softer Cross-modal Alignment Makes CLIP Stronger.

UniCoRN: Unified Commented Retrieval Network with LMMs SoftCLIP: Softer Cross-modal Alignment Makes CLIP Stronger

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:54:16.446151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.895975Z digest=sha256:f8d13416c7425e63dfaed73f6d21aae1364a19ad2f50603fa31768d28bd0421d

Observation 4f9696d8-933d-4d91-8b84-0d05be0fc9a4 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

UniCoRN: Unified Commented Retrieval Network with LMMs Making LLaMA SEE and Draw with SEED Tokenizer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.899573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.899573Z digest=sha256:4db09062e295b8b629ce6cd9cfbea8929b71c9588f665cf738d74d9de3527a9a

Observation 67f67b07-c0a4-4475-8e0d-f043e95af363 · outbound

This paper cites Cyclip: Cyclic contrastive language-image pretraining.

UniCoRN: Unified Commented Retrieval Network with LMMs Cyclip: Cyclic contrastive language-image pretraining

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.902988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.902988Z digest=sha256:aee5eaf19cd73dfee1c74d86593e533755ef29444d117163b479c9ab82098b04

Observation f34a8fe8-19ff-46b6-9529-34b77634a9ef · outbound

This paper cites Fashionvlp: Vision language transformer for fashion re- trieval with feedback.

UniCoRN: Unified Commented Retrieval Network with LMMs Fashionvlp: Vision language transformer for fashion re- trieval with feedback

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.946462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.906213Z digest=sha256:dbf8902e7daf1dc19e0af8f1f04a628eefa0203e9c39393121f75741db60f6db

Observation 53408d33-2e4e-41c6-8db8-c566824c86b0 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

UniCoRN: Unified Commented Retrieval Network with LMMs Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.935833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.909463Z digest=sha256:164afbed7147558188efaebfd7058416ecf505e94767b19c591339d8363c1b60

Observation a5810dab-5893-47e6-b4fe-c280901bc02d · outbound

This paper cites Language-only training of zero-shot com- posed image retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Language-only training of zero-shot com- posed image retrieval

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.925444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.912732Z digest=sha256:1e2e1c86e75556fb12689599f14132de57d79bb19a45b1604d8e13d6ef24b2aa

Observation efdad6e1-a744-4786-9a69-dbe3c71d3d7b · outbound

This paper cites Dialog-based interactive image retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Dialog-based interactive image retrieval

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.915546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.915885Z digest=sha256:c6b2d23f7f0f1485b5a65a9556fc77caa6cfde62719082e47c38f8f1f7d19972

Observation 1ffed5d6-3957-444a-8568-7c2949410145 · outbound

This paper cites Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback.

UniCoRN: Unified Commented Retrieval Network with LMMs Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.919336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.919336Z digest=sha256:bd513ff04c23804d3a00a420a16b603a2772d292d48d94d85a9d579e95fcecdd

Observation c70ead6c-7350-4f67-9286-00a82bb5d4b7 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

UniCoRN: Unified Commented Retrieval Network with LMMs Vizwiz grand challenge: Answering visual questions from blind people

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.922958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.922958Z digest=sha256:7d3ceee04d62d4ed5ba16b159b1ce5c8d8305f82f35a897e92cf3de9f75ce9e7

Observation 4eca5215-0015-4a15-99d2-9b72b539afee · outbound

This paper cites Retrieval augmented language model pre- training.

UniCoRN: Unified Commented Retrieval Network with LMMs Retrieval augmented language model pre- training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.926563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.926563Z digest=sha256:b1dc36a70e449511b1d3134b10e1b46c9adbdcb6f9afc63845b742965dd63bb4

Observation a3130d19-48be-4574-bc75-17cbf7b07ff9 · outbound

This paper cites Au- tomatic spatially-aware fashion concept discovery.

UniCoRN: Unified Commented Retrieval Network with LMMs Au- tomatic spatially-aware fashion concept discovery

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.893871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.929768Z digest=sha256:acc5b5ca49c9020f90f5766c7ca1d184945220349a81f98ed044dfacd93cb9a1

Observation 8472be4a-c96f-495f-b465-d90702b91cfa · outbound

This paper cites Learning attribute-driven disentangled represen- tations for interactive fashion retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Learning attribute-driven disentangled represen- tations for interactive fashion retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.884577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.933138Z digest=sha256:515b2c6d01ec39ef828766247c9ac9496a3fe47de1c07961e36b2f21896a3d55

Observation 6e15b028-1da2-4866-996b-1d52a2d3d25a · outbound

This paper cites Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities.

UniCoRN: Unified Commented Retrieval Network with LMMs Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.874761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.936531Z digest=sha256:e0a96348cf82e364c84d6889e674cb8907727cd40f1b3c4d94a8d1d971a75523

Observation 97fb0400-2a4e-4e51-8ccb-fad1c82afafc · outbound

This paper cites Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge mem- ory.

UniCoRN: Unified Commented Retrieval Network with LMMs Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge mem- ory

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.865028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.939817Z digest=sha256:f9a1bf6b456e2e13f4ed6b7c946b91143f487eba13f6faef2ff7625edd1a3aa6

Observation 88d8d982-2371-488e-9f1d-7496f0ef343d · outbound

This paper cites Openclip, 2021.

UniCoRN: Unified Commented Retrieval Network with LMMs Openclip, 2021

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.855413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.943126Z digest=sha256:e85020d49c51386d5ea113844bca1473f2bce4021ae99fecd3f226381578e5fc

Observation aeb94a3a-7f2b-4238-a29d-50dc08a65f5f · outbound

This paper cites Mantis: Interleaved multi-image instruction tuning, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs Mantis: Interleaved multi-image instruction tuning, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.845876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.946618Z digest=sha256:a759e62b8ceb9d285397a1a57f65caddf132acebc1fc8196e51a95820e67f880

Observation 11b4e977-6a6e-4ec8-8e66-6549b62d0d4e · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

UniCoRN: Unified Commented Retrieval Network with LMMs VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.949814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.949814Z digest=sha256:9a85851176b00b12c46870333de0fb51b30b793e4090747163f4342e8af9bf82

Observation 40be1038-6517-4deb-a45e-396b5470b8e5 · outbound

This paper cites Dense Passage Retrieval for Open-Domain Question Answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Dense Passage Retrieval for Open-Domain Question Answering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.953157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.953157Z digest=sha256:8ea1a28cf1a897185d994c831967573ae05018458360843cd3ce5c690bbb7884

Observation a2540e4b-15b0-4ffc-b70e-47acf8f5fe44 · outbound

This paper cites Vision-by-Language for Training-Free Compositional Image Retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Vision-by-Language for Training-Free Compositional Image Retrieval

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.956672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.956672Z digest=sha256:dc8280e5c866e3f86a68dc0d667ba8d36e215aef0be04e5b83151a6c0192fc2d

Observation 2fefdf49-ab73-4198-b080-6f82ff96ad3f · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

UniCoRN: Unified Commented Retrieval Network with LMMs Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.835520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.960045Z digest=sha256:241ae9521a11a7fa8f7840d67c08f6cd483200813a66ab7e1cae6daf7ae61d1d

Observation 449e2b1a-41d0-4cf8-bf8d-2f5a379d0626 · outbound

This paper cites Grounding language models to images for multimodal in- puts and outputs.

UniCoRN: Unified Commented Retrieval Network with LMMs Grounding language models to images for multimodal in- puts and outputs

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.824825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.963359Z digest=sha256:feeb9d3b9aaf4574fad16be8ce3cf9f0c383daeed6725004c32d0e09b235cf38

Observation a4e72049-9e5a-47b4-9670-8f0756296a55 · outbound

This paper cites UniCLIP: Unified Framework for Contrastive Language-Image Pre-training.

UniCoRN: Unified Commented Retrieval Network with LMMs UniCLIP: Unified Framework for Contrastive Language-Image Pre-training

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.966675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.966675Z digest=sha256:dfb73ea43d5555e1bc63796476497c7a069b1c378e737f823b7b069a19353f50

Observation 24c7f62a-e6d4-4eb0-85a4-49e7e456ee08 · outbound

This paper cites Latent Retrieval for Weakly Supervised Open Domain Question Answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Latent Retrieval for Weakly Supervised Open Domain Question Answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.970255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.970255Z digest=sha256:678df239d6889b93793f548a74c3086cfeb915fefe49400f2d430dc9874eb69c

Observation a26127f8-216b-4c52-99f4-09b99c72868e · outbound

This paper cites Chatting makes perfect: Chat-based image retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Chatting makes perfect: Chat-based image retrieval

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.814878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.974015Z digest=sha256:5ac0777eeffe615503bb79c1ae6055d1b84b76bafbd877dac24e1d6cc74bd304

Observation 707197bd-11e0-4063-856c-5150dd2e5bbf · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

UniCoRN: Unified Commented Retrieval Network with LMMs Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.805378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.977277Z digest=sha256:3af4fb55eee1ff87bc5e33a4534b3f5e4f53d427e04b7840cb0e42401a7979f9

Observation c1f5bef1-4759-4615-879e-d60368870870 · outbound

This paper cites Retrieval-augmented genera- tion for knowledge-intensive nlp tasks, 2021.

UniCoRN: Unified Commented Retrieval Network with LMMs Retrieval-augmented genera- tion for knowledge-intensive nlp tasks, 2021

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.795786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.980488Z digest=sha256:67d3207cc044e8d61aaefc32cea7ae2f3f15c7fa3fcdff43463fc6b21cd5877e

Observation 96d48774-0e2e-4239-a0e7-7e03dee3e2a1 · outbound

This paper cites Textbind: Multi-turn interleaved multimodal instruction- following in the wild.

UniCoRN: Unified Commented Retrieval Network with LMMs Textbind: Multi-turn interleaved multimodal instruction- following in the wild

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.785517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.983696Z digest=sha256:cee6ed54c06a5b047cb506803bf8f5bdb4c77272d39cac155d3983f486a7141d

Observation fd3196f1-a8bc-467e-954c-0fcacac5b9c5 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

UniCoRN: Unified Commented Retrieval Network with LMMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.987047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.987047Z digest=sha256:3e66c8eac38b71d415a6c5f87c774d42cb79a8ee4334e2359a19b3e87b77c32a

Observation 2e4636fd-fcbc-470c-b437-31dd485575b2 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

UniCoRN: Unified Commented Retrieval Network with LMMs Rouge: A package for automatic evaluation of summaries

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.990318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.990318Z digest=sha256:da6a5eecab85ec4e314f7156abece5807033413fc1154049bddcb88bf35f4372

Observation 591ce93e-7e28-434c-9f98-8987dc53e7b8 · outbound

This paper cites MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs.

UniCoRN: Unified Commented Retrieval Network with LMMs MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.993546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.993546Z digest=sha256:3e27b2d9747046fa7376fb6fc634c9abc172a577373f45c7190800cff5751fcb

Observation 7ee84982-63be-4cba-a493-145152d42283 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

UniCoRN: Unified Commented Retrieval Network with LMMs Improved baselines with visual instruction tuning, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.762848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:15.997037Z digest=sha256:66c30ccc0ca6cb4f0602c216dc731dc1c91c198d79e288dd1b59f4ae86c54a2d

Observation 88b233ce-38d6-4a2a-9fd8-41b3558cbc82 · outbound

This paper cites Visual instruction tuning, 2023.

UniCoRN: Unified Commented Retrieval Network with LMMs Visual instruction tuning, 2023

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.751844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.000680Z digest=sha256:d6a37e50f6290e576de7637201f3c7001f3d827260a0a7408125c4d5482926ca

Observation 5d36c078-a63d-47a0-9cb2-52bd215e595f · outbound

This paper cites Image retrieval on real-life images with pre-trained vision-and-language models.

UniCoRN: Unified Commented Retrieval Network with LMMs Image retrieval on real-life images with pre-trained vision-and-language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.739618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.003945Z digest=sha256:2b4b7320aa7d3ca3a77c68c4602ca06409d33c6956463d8b99b6cdb9ac8b73ce

Observation cc3c7b49-a857-4c7a-bcc2-088717bdaeae · outbound

This paper cites Image retrieval on real-life images with pre- trained vision-and-language models.

UniCoRN: Unified Commented Retrieval Network with LMMs Image retrieval on real-life images with pre- trained vision-and-language models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.007282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.007282Z digest=sha256:96928fd150a22119cc0a463485b23f3d82980d6506b2c5c5c2b85a90b40e9a4d

Observation 4bb2271a-8d20-46fd-ae2b-601b1da7987c · outbound

This paper cites Bi-directional training for composed im- age retrieval via text prompt learning.

UniCoRN: Unified Commented Retrieval Network with LMMs Bi-directional training for composed im- age retrieval via text prompt learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.723592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.010587Z digest=sha256:630ea23610d48a243c69c924eca64919e81adc8efa21949aebdda9d0172d3e77

Observation 6d68932e-aee0-4a36-9613-52f3a90496fe · outbound

This paper cites Three facets of visual and verbal learners: Cognitive ability, cognitive style, and learning preference.

UniCoRN: Unified Commented Retrieval Network with LMMs Three facets of visual and verbal learners: Cognitive ability, cognitive style, and learning preference

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.713434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.013721Z digest=sha256:bce9fc255fecc85c2efadc0ce2e03155818bed0c83211b2d15999efea8f5aea8

Observation e1354ed6-e14f-4de9-8f5e-bbb089595898 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

UniCoRN: Unified Commented Retrieval Network with LMMs MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.017042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.017042Z digest=sha256:8ce08477448393c4336a7884a4d5af4ad2f9808595dba048cbec929d678ee842

Observation f95c9202-b167-4158-9d88-f3e1a0bdc10d · outbound

This paper cites Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories.

UniCoRN: Unified Commented Retrieval Network with LMMs Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.703498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.020707Z digest=sha256:72204dabaf2309e84ef4c611d2ad385a761639d3da437885ad0efef8531b40a1

Observation 4b4b17d4-a55e-4555-89c1-548ddd5e7829 · outbound

This paper cites Gpt-4 technical report, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs Gpt-4 technical report, 2024

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.024210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.024210Z digest=sha256:5c6dcf44841f43106d535c510cdd57883aaf2c4a07fcbcdeacf9bbdced2e643f

Observation 342eb1c4-f58d-44fa-931b-f0ff86bb93fd · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

UniCoRN: Unified Commented Retrieval Network with LMMs Bleu: a method for automatic evaluation of machine translation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.687287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.027601Z digest=sha256:5375007695319faab9e3e77bd961b420563f683cdccb228830ea1bb2ca8eb172

Observation a94d9f92-3caa-43da-9e5f-d4e191272feb · outbound

This paper cites Rora-vlm: Robust retrieval-augmented vision language models, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs Rora-vlm: Robust retrieval-augmented vision language models, 2024

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.677710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.031009Z digest=sha256:3c8be2300e277737ad8c6ff7282ff12f5e57eb69cdf82268e83a62fbee06529b

Observation 19c69085-513f-47e0-9035-33bc8a69694c · outbound

This paper cites Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation.

UniCoRN: Unified Commented Retrieval Network with LMMs Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.035119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.035119Z digest=sha256:1bfd6179b1928ec89b781fc4772cae3c8c39f3edbd76ed202ff21c40576e38e6

Observation 7b0afb44-c106-4e9a-8e51-81e36db8fc49 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

UniCoRN: Unified Commented Retrieval Network with LMMs Learning Transferable Visual Models From Natural Language Supervision

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.038739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.038739Z digest=sha256:ecd67dd717bd9f4bab5a123f40c0aa4b600e4ca1c5640e6b30de3848679a1ce5

Observation 645a6847-16b2-456f-b901-afd2f19c4813 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

UniCoRN: Unified Commented Retrieval Network with LMMs LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.042435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.042435Z digest=sha256:71bcff92126d7e856b0a44cf07b6c687fce3f7aa5964815c5e3edc83e251f2bd

Observation 94e1b93c-2782-482f-8579-efdd03bdd4a6 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

UniCoRN: Unified Commented Retrieval Network with LMMs A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.667678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.045951Z digest=sha256:5c16671d4292bdf44618ecae25616f87e424c286b5a58e45ee9f5eac893da33a

Observation 084a3a0d-94b5-4e4c-b964-2dc08e73e7bd · outbound

This paper cites Kvqa: Knowledge-aware visual question answering.

UniCoRN: Unified Commented Retrieval Network with LMMs Kvqa: Knowledge-aware visual question answering

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.657548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.049269Z digest=sha256:bf70f5aed09b311b7831408be2b6c8233ef6cbb8a153eea61757c2b5dfc21573

Observation 5932d71f-7388-4f8e-885e-54ca21648e15 · outbound

This paper cites Towards vqa models that can read.

UniCoRN: Unified Commented Retrieval Network with LMMs Towards vqa models that can read

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.646486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.052474Z digest=sha256:aeb861f95f5a2feedb0bd335ae49221f617bd917e44ed1aa54bf910c23b11aff

Observation b0f5bc17-34e7-4571-bb41-9805ece0c23a · outbound

This paper cites Knowledge-enhanced dual-stream zero-shot composed im- age retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Knowledge-enhanced dual-stream zero-shot composed im- age retrieval

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.635984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.056213Z digest=sha256:2749d8b0677f6b2102377abd899c0a5f347a7f358944e2ef65ca5e4b24c94579

Observation e5afdb99-5d1e-4271-a2e3-02333f732556 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

UniCoRN: Unified Commented Retrieval Network with LMMs Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.059617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.059617Z digest=sha256:9c5b5917632a33fec9802312f48fd62ccd031f63219867bbc4a4ec0300ad173e

Observation 17392f1e-64e6-4937-80c0-e3260e4c6c9f · outbound

This paper cites Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024.

UniCoRN: Unified Commented Retrieval Network with LMMs Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.625967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.063185Z digest=sha256:677da09737bd2bfaa5b5282cbf979101800119dc746fef3245d1024d72475605

Observation 918b30d8-a3b9-4390-9cec-7a5a6d442a61 · outbound

This paper cites MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer.

UniCoRN: Unified Commented Retrieval Network with LMMs MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.066474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.066474Z digest=sha256:563289affe3ff76e4951c7df0d7cc30f18897974d5be7cc030c0543796fa5797

Observation faeba532-7ee5-473f-a4ec-791afc12438a · outbound

This paper cites Genecis: A benchmark for general conditional image similarity.

UniCoRN: Unified Commented Retrieval Network with LMMs Genecis: A benchmark for general conditional image similarity

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.615682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.069876Z digest=sha256:6ef143581a2292d0231d51f6831042ac5bb4326277eca0e204e2d3097b33f5e9

Observation c4307a2c-2981-4e25-ac03-72277d72af1e · outbound

This paper cites Composing text and image for image retrieval - an empirical odyssey.

UniCoRN: Unified Commented Retrieval Network with LMMs Composing text and image for image retrieval - an empirical odyssey

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.605302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.073150Z digest=sha256:63fff1e83f5a9b6e7302f03939abef1532c771abdd2ee0d35d6f1766e263e2e3

Observation 8a47854d-ddaf-4635-ba28-df7b6e540b35 · outbound

This paper cites Composing text and image for image retrieval-an empirical odyssey.

UniCoRN: Unified Commented Retrieval Network with LMMs Composing text and image for image retrieval-an empirical odyssey

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.076600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.076600Z digest=sha256:6da135eb545649099932077953e842083bbb54ef701ab6bb5e384c5e31692153

Observation d573bb81-416d-4a88-87cc-ffb22d910204 · outbound

This paper cites Cross-modal feature alignment and fusion for com- posed image retrieval.

UniCoRN: Unified Commented Retrieval Network with LMMs Cross-modal feature alignment and fusion for com- posed image retrieval

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.589090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.079980Z digest=sha256:073288b8e282b3b7f8a856ae4af918ff1a8daca27dd02ca735b53ba52c920eae

Observation 891a069f-594a-423f-a1ca-f90029ac0587 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

UniCoRN: Unified Commented Retrieval Network with LMMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.083564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.083564Z digest=sha256:b69a48acb47dad5fb14480323f44b9deb582e8230a63ce8896f80f75b80faa3e

Observation 182dbbdc-3271-4672-97ca-4db2bebad9f2 · outbound

This paper cites Image as a foreign language: Beit pretraining for vision and vision- language tasks.

UniCoRN: Unified Commented Retrieval Network with LMMs Image as a foreign language: Beit pretraining for vision and vision- language tasks

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.578986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.087345Z digest=sha256:17c099a71a6d9a37386ec53a06954234eb1de6ba535c94e86bc7a33f794f0023

Observation 8462c595-82f9-4c73-9e8a-f598044ab3da · outbound

This paper cites UniIR: Training and Benchmarking Universal Multimodal Information Retrievers.

UniCoRN: Unified Commented Retrieval Network with LMMs UniIR: Training and Benchmarking Universal Multimodal Information Retrievers

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.090827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.090827Z digest=sha256:d25768c74d002e002a5e6543a976221b9fe1ddfa3d4ff5e4358922eaf2e9962c

Observation 7faf818e-0027-4616-b6a6-dba6c7623a02 · outbound

This paper cites Fashion iq: A new dataset towards retrieving images by natural language feedback.

UniCoRN: Unified Commented Retrieval Network with LMMs Fashion iq: A new dataset towards retrieving images by natural language feedback

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.569127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.094338Z digest=sha256:944a5ade834c32b7278d5be4ead946efca78d07de83a06d22ac69920c5245989

Observation f480dfe3-c184-478b-a138-44100814778a · outbound

This paper cites Fashion iq: A new dataset towards retrieving images by natural language feedback.

UniCoRN: Unified Commented Retrieval Network with LMMs Fashion iq: A new dataset towards retrieving images by natural language feedback

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.558907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.097683Z digest=sha256:f3a40f4ec8ab270f9bddfa28c2373860cb9b760f7b2417f67e2c5a175b147354

Observation f981b359-073d-4059-9f45-42a1fba8df2e · outbound

This paper cites Visual question answer- ing: A survey of methods and datasets.

UniCoRN: Unified Commented Retrieval Network with LMMs Visual question answer- ing: A survey of methods and datasets

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.101159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.101159Z digest=sha256:624f53bd401ee50358ee6821ef06880561a48d0ff315d67a7eeb63ee72c14c47

Observation cb9d029c-c0d8-463e-80cb-d7253de4152e · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

UniCoRN: Unified Commented Retrieval Network with LMMs NExT-GPT: Any-to-Any Multimodal LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.104506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.104506Z digest=sha256:6c39b72a2f28f0fe5e66711e32447e483945d95fe8db40d7d0840015c8449f8a

Observation 22ec78be-1e28-48a4-af82-7030b566ab26 · outbound

This paper cites EchoSight: Advancing visual- language models with Wiki knowledge.

UniCoRN: Unified Commented Retrieval Network with LMMs EchoSight: Advancing visual- language models with Wiki knowledge

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.542820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.108201Z digest=sha256:136afbae57ddb290474fc6c853a7574bc7e2da95a3a68d27e71bf66679725fa4

Observation 44e175a0-ec5e-4e91-9d78-8d43d48b7cc9 · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

UniCoRN: Unified Commented Retrieval Network with LMMs FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.111778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.111778Z digest=sha256:e703c1f7b6f14175b38c1422356fdb754234716d048701e5a8a9c1af8671623a

Observation f3622b38-edbd-4faf-8681-7c8a7e89af82 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

UniCoRN: Unified Commented Retrieval Network with LMMs mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.115407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.115407Z digest=sha256:c44f170e268dff31616e475171446bb7c44a4c0103c856f828c16f3a2bb10d70

Observation 962fba5b-9b5b-4d40-b138-8a495af5dd8b · outbound

This paper cites A Survey on Multimodal Large Language Models.

UniCoRN: Unified Commented Retrieval Network with LMMs A Survey on Multimodal Large Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.119075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.119075Z digest=sha256:39cf2cbe2170e8b2d6a623ea605c4db73412ba429ffef56445cbbd1034af2aae

Observation 43394a8b-a74f-4e12-ad5a-39f6911df660 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

UniCoRN: Unified Commented Retrieval Network with LMMs CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.122413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.122413Z digest=sha256:b41093b0e6ac373bd39edc2b02ddf2474fc0b60c5dbce95ceec43cd1bd5659d6

Observation 00b79990-a65d-40ee-9b7d-cc78b3d733dc · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

UniCoRN: Unified Commented Retrieval Network with LMMs Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.126629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.126629Z digest=sha256:dad2283b4b594b36fc7c3f3cda3c5b7b4a03f673b017435556b1d32d86f757d2

Observation 3382d6d5-1c24-40b4-a9db-3d32e9e0b111 · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

UniCoRN: Unified Commented Retrieval Network with LMMs VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.130352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.130352Z digest=sha256:2235110ca5b3f710de6e1db59d7dcc00e7384cb2bb1301e339e9b53d5d30cf56

Observation 5ebec1ff-b09f-4a7e-ba09-4104836ad76b · outbound

This paper cites A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark.

UniCoRN: Unified Commented Retrieval Network with LMMs A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.134206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.134206Z digest=sha256:6ac45bcec84fa0c46ee5807cb4765cd959bb06d5d313bb9e99dcb21c8ab20e6f

Observation a892aaff-44e9-478b-91be-e937e72ad97f · outbound

This paper cites Sigmoid loss for language image pre-training.

UniCoRN: Unified Commented Retrieval Network with LMMs Sigmoid loss for language image pre-training

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.532426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.137913Z digest=sha256:e7a603007f883f4df3aa00bdbfdbfbc4837acc249ca4df88ceaf24143fbb8975

Observation 151d9cde-c624-4470-b210-bd4e414baa0f · outbound

This paper cites MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions.

UniCoRN: Unified Commented Retrieval Network with LMMs MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.141327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.141327Z digest=sha256:7502816a21a7043a30cedfb75b91ffeb2043c37b1daf3f0139972f31bebc6de2

Observation a4528704-16cb-4ea1-b256-d898499e5c46 · outbound

This paper cites Non-Contrastive Learning Meets Language-Image Pre-Training.

UniCoRN: Unified Commented Retrieval Network with LMMs Non-Contrastive Learning Meets Language-Image Pre-Training

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:54:16.184255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.145165Z digest=sha256:a8eebd1bd40c99d4e860e6ab68cd362a89068c660c00a54c6c73e1049f85b716

Observation e6957db2-f559-4b6d-bc2e-ba3a7e2a2f08 · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language models.

UniCoRN: Unified Commented Retrieval Network with LMMs MiniGPT-4: Enhancing vision-language understanding with advanced large language models

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:54:16.522347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T05:54:16.148845Z digest=sha256:299be5c8281d3805cf2a04b51d37c260d1cbeb8a3f95be4941dba99abf0b09a9

Pith citing papers

No inbound Pith citation observations are available.