REVIEW 12 cited by
ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Neural information retrieval (IR) has greatly advanced search and other knowledge-intensive language tasks. While many neural IR methods encode queries and documents into single-vector representations, late interaction models produce multi-vector representations at the granularity of each token and decompose relevance modeling into scalable token-level computations. This decomposition has been shown to make late interaction more effective, but it inflates the space footprint of these models by an order of magnitude. In this work, we introduce ColBERTv2, a retriever that couples an aggressive residual compression mechanism with a denoised supervision strategy to simultaneously improve the quality and space footprint of late interaction. We evaluate ColBERTv2 across a wide range of benchmarks, establishing state-of-the-art quality within and outside the training domain while reducing the space footprint of late interaction models by 6--10$\times$.
Forward citations
Cited by 12 Pith papers
-
Semantic Homogenization in Italian Popular Music: A Diachronic Analysis
Sanremo lyrics exhibit rising semantic homogeneity over decades, consistently recovered by full-text, portion, topic and word-level embedding analyses.
-
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning
A tool-augmented reward model trained with GRPO on 27K synthetic pairs beats existing reward models on long-form QA judgment and improves downstream alignment.
-
Zero-Shot Contextual Embeddings via Offline Synthetic Corpus Generation
ZEST shows that a small synthetic corpus generated by GPT-4o from five examples can stand in for the real target corpus in a frozen context-aware embedding model, losing under 0.5% retrieval accuracy on MTEB.
-
Identifying Origins of Place Names via Retrieval Augmented Generation
A RAG pipeline using ColBERTv2 and Llama2 retrieves Melbourne street-name origins from DBpedia, but language models under-use spatial context, limiting top-1 accuracy.
-
FlexRAG: A Flexible and Comprehensive Framework for Retrieval-Augmented Generation
FlexRAG is a modular, open-source RAG framework with text, multimodal, and web retrieval, plus evaluation tools and efficient memory-mapped indexing.
-
ASARL: Autonomous Social-Aware Relevance Learning for QQ Search
An agent-loop data-curation pipeline with social-aware chain-of-thought, preference, and distillation training improves QQ group/channel search relevance in offline and online evaluation.
-
JobMatchAI-An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search and Explainable AI
On the new JobSearch-XS benchmark, the hybrid JobMatchAI pipeline reaches NDCG@10 of 0.81 (about 7% over BM25) with a white-box, factor-level reranker and LLM explanations.
-
HyST: LLM-Powered Hybrid Retrieval over Semi-Structured Tabular Data
A hybrid retrieval system that combines LLM-generated attribute filters with embedding search outperforms several baselines on a small, curated semi-structured product benchmark.
-
Enhanced Arabic Text Retrieval with Attentive Relevance Scoring
An Arabic dense retriever using a trainable attentive scoring module instead of dot-product similarity reports improved top-k passage retrieval on ArabicaQA.
-
GOLFer: Smaller LM-Generated Documents Hallucination Filter & Combiner for Query Expansion in Information Retrieval
GOLFer filters hallucinated sentences from small-LM-generated hypothetical documents and reweights the rest into the query, improving retrieval at lower cost than large LLM expansion.
-
Ask, Retrieve, Summarize: A Modular Pipeline for Scientific Literature Summarization
XSum, a question-generation plus editor RAG pipeline, produces survey-style summaries from multiple scientific papers and reports improved scores on the SurveySum benchmark.
-
Semantic Certainty Assessment in Vector Retrieval Systems: A Novel Framework for Embedding Quality Evaluation
A query-level score combining quantization stability and neighborhood density predicts retrieval performance and is claimed to improve Recall@10 by only 2 to 3 percent per dataset.
Discussion (0). Sign in to comment.