Pith. sign in

REVIEW 2 cited by

PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.08327 v2 pith:4PLA2XMA submitted 2024-02-13 cs.CL

classification cs.CL
keywords multi-modaltaskskb-vqam2krpreflmrretrieversfine-grainedgeneral-purpose
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Multimodal Models (LMMs) excel in natural language and visual understanding but are challenged by exacting tasks such as Knowledge-based Visual Question Answering (KB-VQA) which involve the retrieval of relevant information from document collections to use in shaping answers to questions. We present an extensive training and evaluation framework, M2KR, for KB-VQA. M2KR contains a collection of vision and language tasks which we have incorporated into a single suite of benchmark tasks for training and evaluating general-purpose multi-modal retrievers. We use M2KR to develop PreFLMR, a pre-trained version of the recently developed Fine-grained Late-interaction Multi-modal Retriever (FLMR) approach to KB-VQA, and we report new state-of-the-art results across a range of tasks. We also present investigations into the scaling behaviors of PreFLMR intended to be useful in future developments in general-purpose multi-modal retrievers.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AR-RAG: Autoregressive Retrieval Augmentation for Image Generation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Autoregressive patch-level retrieval augmentation improves text-to-image generation on GenEval, DPG-Bench, and Midjourney-30K, with a training-free decoding variant and a fine-tuned variant.

  2. UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering

    cs.IR 2026-08 conditional novelty 6.0 of 10

    UniHEAR combines image-to-image and image-to-text candidate retrieval with source-aware attention reranking, improving Recall@1 over prior reranking methods on E-VQA and InfoSeek.

Pith tools