Pith. sign in

Paper Citation Record · LEDGER

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture

As of 10 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2605.14448.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.14448 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T02:51:39.142437Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:43:23.733753Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T11:43:28.407453Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact11
  • verified fuzzy31
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 94adbe60-0729-4d73-8a99-83e7efbbafc4 · outbound

This paper cites Qwen3-VL Technical Report.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:53:33.620043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:e6d4205d0694b8d9e0f46d97ba7184cf42d98a5a43b986eab9d68d834bc9927d

Observation d3eb9b61-dd0a-413b-b66d-f29fe0365f3b · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:53:33.616969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:2fd9dbb26fa7231957d9cd589f02c30c954dfc32a5179e48e6340e5da494ee30

Observation 42fd8a27-34c3-4abc-9e9a-53c10427595e · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.298030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:5ed000abb12fe96facc6f34e02b9a9ba6567697d7d07949ae4011b6b1de59ed0

Observation 25ce55d7-42cb-4c1f-848d-445effd6ea1d · outbound

This paper cites arXiv preprint arXiv:2510.05014 , year=.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture arXiv preprint arXiv:2510.05014 , year=

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:53:33.574647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:bcbbac7e6f6a1e3bb925307ef561ea53f76f52ce0f9efa57a57df2cf1a5fe1d8

Observation 40922353-9b79-4575-8926-64edb62500c5 · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.274547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:ce59111fb1f33a23f0b608be8b69f93770fd893b2f3bc9e62ea7017d9f3374b0

Observation f90352f8-3ccd-499d-a2cd-d0e67a4075f8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:53:33.607157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:d778e60bfdd58b9d7b35bd2e8b36d68f5b64dd0ff9852d0b26aa8b0c6103fc65

Observation 7b41eda3-caf2-4586-bef9-e41f5f07707f · outbound

This paper cites Colpali: Efficient document retrieval with vision language models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Colpali: Efficient document retrieval with vision language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.278740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:9429b57e2a3efb215a6714aaac301656307d90b0b0d2eff0c720081870ede154

Observation 3af8c85c-284f-4954-9b19-3eddcf13b494 · outbound

This paper cites Scaling deep contrastive learning batch size under memory limited setup.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Scaling deep contrastive learning batch size under memory limited setup

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.288208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:0d0dd0e2c2bd86d2d9c34ed7c29ae24577d59651c59d5a17f2218387353567e9

Observation 2206bcc1-fde0-43e8-aab8-39fefd06315b · outbound

This paper cites Breaking the modality barrier: Universal embedding learning with multimodal llms.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Breaking the modality barrier: Universal embedding learning with multimodal llms

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.644603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:7e39b55968e8267b8ff05185c66b9af6c3aa9b0fce1390c4b912973417f0ed85

Observation 77ac9636-68f6-4152-aa96-d27a2844eb3f · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Lora: Low-rank adaptation of large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.625790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:d80dcc406c33af280c7ab9ad599180658a8d78dd2ca542b091db4317f66a3a69

Observation 726c1a68-4710-4e0f-8e58-62d996a9e19c · outbound

This paper cites Cumulated gain-based evaluation of IR techniques.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Cumulated gain-based evaluation of IR techniques

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.632221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:56760adba872a5fc380241c09613e412c8790961f901c924f7c7c6f9e67964ce

Observation eea7682d-c7ae-40c2-afb1-4ad1dc97be29 · outbound

This paper cites Le, Yun- Hsuan Sung, Zhen Li, and Tom Duerig.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Le, Yun- Hsuan Sung, Zhen Li, and Tom Duerig

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.638328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:f0f373664c3ac2f5e5f113a37af0316535de1c0f3f4ad966b7dac0604b949b8f

Observation f298b2ba-57b1-4556-9b70-3f0eb4cef87a · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:52:21.013841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:d05a3ad19327500cd9eaef16c7e14715d4a48423cc2a8f7a471d30c9f7baa626

Observation 8b067369-f4a1-41b0-8240-1d9f39b73375 · outbound

This paper cites Vlm2vec: Training vision-language models for massive multimodal embedding tasks.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Vlm2vec: Training vision-language models for massive multimodal embedding tasks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.650503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:f5cb617f8c01ff9d1d8ef3ea68c5c3ba84ddfcc6f885471bcb3d2cea2a6e0c38

Observation 53ea582e-0153-4b8c-ad25-f86a72722762 · outbound

This paper cites Large language models are zero-shot reasoners.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Large language models are zero-shot reasoners

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.598543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:e4b947be6e64d7154ee82d032b691cf63f501fe507eb0d51e39a0fb623ac319f

Observation 172de98b-cec0-4fb6-a8ab-e5474d3aaed5 · outbound

This paper cites Llave: Large language and vision embedding models with hardness-weighted contrastive learning.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Llave: Large language and vision embedding models with hardness-weighted contrastive learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.564439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:a01a49695bfa1b50ce76c793010645bd1689b40e34a20f6e32297faecab1a9b8

Observation 2d20370a-6f3c-4ca1-b10f-0a4df1dc1923 · outbound

This paper cites arXiv preprint arXiv:2511.00405 , year=.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture arXiv preprint arXiv:2511.00405 , year=

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:53:33.610275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:c6dd0a1079c024c57c2959f9113f7bb53073b6d49c7e45582712e2d8bb066d72

Observation 27772d70-64f0-4b92-b333-4f37453b81e0 · outbound

This paper cites Nv-embed: Improved techniques for training llms as generalist embedding models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Nv-embed: Improved techniques for training llms as generalist embedding models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.570183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:38d5bb32030dc4c53bc7dd404bfea39a183393a736a8252265b09fa10450e70d

Observation 7c55fd41-6d7d-4067-b70b-f541a8ba9599 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.663216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:598f332402eba1317e496f6b35886c285e8103f8b4ae40c9fd410611847b0ca7

Observation b8ec17a7-4ed5-45de-8d2d-951bcfe8d1ef · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.656404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:6bbb2384bf84e05b63e37d469f000335c969db7131924fddc5afe92bfe086f88

Observation 7820e316-f812-4cb0-af06-1b8e88d78800 · outbound

This paper cites Mm-embed: Universal multimodal retrieval with multimodal llms.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Mm-embed: Universal multimodal retrieval with multimodal llms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.293336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:0e35e337c3f590016afa749bea83f5816c79a403a56ca3345a1a03d569a4834d

Observation 90618f00-1002-4306-8f97-d57f97f28423 · outbound

This paper cites Reasoning guided embeddings: Leveraging mllm reasoning for improved multimodal retrieval.arXiv preprint arXiv:2511.16150, 2025a.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Reasoning guided embeddings: Leveraging mllm reasoning for improved multimodal retrieval.arXiv preprint arXiv:2511.16150, 2025a

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:53:33.613684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:4a20f3c8db41775407090d9e583d7b5ed82215474e974c205ae144780aeb118b

Observation 22cc78b9-ade5-430f-8089-07b9cd8217b0 · outbound

This paper cites Visual instruction tuning.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Visual instruction tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.521297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:9030a6919d14a8d7c14e7cd6d444cc83212d55af04f0d022245aad5ab6d3eca3

Observation f2b3302f-c83e-4bb9-9bb9-e1470ed0f58d · outbound

This paper cites Lamra: Large multimodal model as your advanced retrieval assistant.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Lamra: Large multimodal model as your advanced retrieval assistant

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.546804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:e0fa4ecb09916fe273c8750dbee48220623de76fea7a49ef12470e461c07403a

Observation bf140399-66ac-4dd8-b1e3-dc555b67bee1 · outbound

This paper cites Vlm2vec-v2: Advancing multimodal embedding for videos, images, and visual documents.Transactions on Machine Learning Research.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Vlm2vec-v2: Advancing multimodal embedding for videos, images, and visual documents.Transactions on Machine Learning Research

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.533510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:12e5e342082cca8b2db5dee5caa584eed71f58e1848531b88acb4d5d74db590f

Observation 7dec594e-9711-436f-9a7a-fc61c71f2de6 · outbound

This paper cites Mteb: Massive text embedding benchmark.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Mteb: Massive text embedding benchmark

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.488914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:961c6ce196ff6bc314d7b16e0c396abadb979577dfc721ce033dcf3839972888

Observation f775eba8-d4f0-4432-a150-44943d056710 · outbound

This paper cites Vladva: Discriminative fine-tuning of lvlms.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Vladva: Discriminative fine-tuning of lvlms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.494114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:d3d9a253cd622f05c9f9a2e96ae9f8102a39801f28b72e48877552c220c31adf

Observation bb206138-b18b-4a66-8360-34b9b6206aea · outbound

This paper cites Learning transferable visual models from natural language supervision.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Learning transferable visual models from natural language supervision

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.529126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:7048a30c5437ff06a964e0036c31476ffca14a0be171a438e37dec30a9b399f5

Observation 3d8133bb-2658-49a1-b1a8-686d3b852be5 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.479475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:674ec89c837b21ab5e6299144377d810aff2174e5bfc78ff43e6d84efc37f3ad

Observation 757f8968-967f-4c29-a2a3-252cf2bd40c1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:53:33.603874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:6ce43a45cbc0bbb74a18721df7c2a7d8a74afa8c2212d561e9185479bdc0aafa

Observation 85f17632-76b2-4724-93ea-09451092bffc · outbound

This paper cites Smith, Luke Zettlemoyer, and Tao Yu.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Smith, Luke Zettlemoyer, and Tao Yu

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.282994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:de6d3826933610c03599337de6e0829cdc0e26a7a5ba635a0252b7adba2a1aab

Observation 8bc26bd1-8c86-4ce7-b1a9-0b05849ae7f8 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Representation Learning with Contrastive Predictive Coding

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:53:33.580487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:82a5abbaa933a58fc17c6d453eeee0c15b8d168a85620109231c0536bf1fec67

Observation a06c2e40-b8e6-41f8-934d-a382f565b79a · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:53:33.600217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:e43c59e04ae84eb50ddb6908cb820e15b451c4eefa91a4cfbf95da8d98f8de83

Observation 3e391284-0470-486c-bc79-a579880df57c · outbound

This paper cites Improving text embeddings with large language models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Improving text embeddings with large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.266541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:3e159b9fd62c847078062415464bc72a4ba5db1ab64d8aa8c24d62f63a9a43a5

Observation ebe631c6-5aa0-40e7-afea-907fc4ea5f7d · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:53:33.577662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:95d430c074f9f7c17e1e573aa4bf795a8340f2c1c7092e0aa0d53e2c07111746

Observation 34e63aaa-94ee-4414-a698-afd43ccfc5ed · outbound

This paper cites Chi, Fei Xia, Quoc Le, and Denny Zhou.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Chi, Fei Xia, Quoc Le, and Denny Zhou

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:48.270504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:75d17483ee1bedd3d4894104feb9e4d4d6c4656b0f93b3c6659209e22103f1ec

Observation e7129446-75d7-4c23-9a26-1a6bb699f922 · outbound

This paper cites Cafe: Unifying representation and generation with contrastive-autoregressive finetuning.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Cafe: Unifying representation and generation with contrastive-autoregressive finetuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.461203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:913965957b9e76b60abbe7ad9e44954c32401e2a97e8c2dd4739051cf9da8613

Observation 4f3b7b09-ed5a-4b8d-af1f-4f3417d2b966 · outbound

This paper cites Visrag: Vision-based retrieval-augmented generation on multi-modality documents.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Visrag: Vision-based retrieval-augmented generation on multi-modality documents

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.466795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:e819f2d2b3b1a8d9bed48f503666d2fd6d2ead24dcd63aff4cb47e42e570e97b

Observation 637e55fb-c38b-4171-9d01-648dc1e5cd25 · outbound

This paper cites Gradient surgery for multi-task learning.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Gradient surgery for multi-task learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.473299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:06a03838eef4d1d8c00367e991eecd061c6f26f23a71747bce9a2cae6879358a

Observation 33c1b78a-bf01-4b73-9cc9-8f5f9d76b865 · outbound

This paper cites Sigmoid loss for language image pre-training.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Sigmoid loss for language image pre-training

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.809767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:784356fe67179aeaad1f826e3230d47b4f76dc6df1f1616b8a8749173c419088

Observation 1c230e21-1155-4bc8-99b6-5a18b2dfed93 · outbound

This paper cites Hauptmann, Yonatan Bisk, and Yiming Yang.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Hauptmann, Yonatan Bisk, and Yiming Yang

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.483846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:8540b8e3a9d6b1fe5da612f5cbe6e127bfe0a0685f7f6677f14c3ced2e531a75

Observation 4ed72dd8-0c61-4dbe-9ffd-e1f44cc25ef9 · outbound

This paper cites Bridging modalities: Improving universal multimodal retrieval by multimodal large language models.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture Bridging modalities: Improving universal multimodal retrieval by multimodal large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T03:59:47.551142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:a2905d86a4abad28d6b910a0bb9a9cb2420bbee5144a941ae320314e16f25d73

Observation 56230514-a7eb-4d62-9193-f769326d580a · outbound

This paper cites think" field completely empty (.

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture think" field completely empty (

Reference 43

Resolution
malformed identifier
raw_fallback, observed 2026-05-15T03:59:47.455293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:51:39.142437Z digest=sha256:6be6325ee5b5bdd4c740cb0b210f5712b8ac172b61c5c6989d1d0cdc2b0d977a

Pith citing papers

Observation c1fef859-a397-4ace-b531-59308763c5da · inbound

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding cites this paper.

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-05T11:43:28.454660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T11:43:23.733753Z digest=sha256:5cfe2fd3c783c69ef933dabe8a26f546385863371f55e54aec27b45d9f8f3fb5