Pith. sign in

REVIEW 9 cited by

Document Ranking with a Pretrained Sequence-to-Sequence Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.06713 v1 pith:XZ7HYNYH submitted 2020-03-14 cs.IR cs.LG

classification cs.IRcs.LG
keywords modelrankingapproachmodelspretrainedsequence-to-sequencetargetwords
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work proposes a novel adaptation of a pretrained sequence-to-sequence model to the task of document ranking. Our approach is fundamentally different from a commonly-adopted classification-based formulation of ranking, based on encoder-only pretrained transformer architectures such as BERT. We show how a sequence-to-sequence model can be trained to generate relevance labels as "target words", and how the underlying logits of these target words can be interpreted as relevance probabilities for ranking. On the popular MS MARCO passage ranking task, experimental results show that our approach is at least on par with previous classification-based models and can surpass them with larger, more-recent models. On the test collection from the TREC 2004 Robust Track, we demonstrate a zero-shot transfer-based approach that outperforms previous state-of-the-art models requiring in-dataset cross-validation. Furthermore, we find that our approach significantly outperforms an encoder-only model in a data-poor regime (i.e., with few training examples). We investigate this observation further by varying target words to probe the model's use of latent knowledge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Models

    cs.CL 2025-08 conditional novelty 7.0 of 10

    On a new benchmark of post-April 2025 queries, LLM rerankers show a 5-15% performance drop compared with familiar benchmarks, and lightweight models match them on efficiency and sometimes accuracy.

  2. DocReRank: Single-Page Hard Negative Query Generation for Training Multi-Modal RAG Rerankers

    cs.IR 2025-05 conditional novelty 7.0 of 10

    Training a reranker on VLM-verified hard negative queries, generated per page from LLM rephrasings of positive queries, outperforms training on document-level hard negatives in multimodal RAG retrieval.

  3. Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain

    cs.IR 2025-09 conditional novelty 6.0 of 10

    Misleading health documents in RAG context sharply lower LLM accuracy, and heavily helpful-biased retrieval pools restore it.

  4. Universal Biological Sequence Reranking for Improved De Novo Peptide Sequencing

    cs.LG 2025-05 conditional novelty 6.0 of 10

    RankNovo, a list-wise reranker with mass-deviation supervision, improves de novo peptide sequencing accuracy by selecting among candidates from multiple base models.

  5. Seeing the Forest Through the Trees: Knowledge Retrieval for Streamlining Particle Physics Analysis

    hep-ex 2025-09 conditional novelty 5.0 of 10

    Retrieving answers from the LHCb literature using hierarchical paper trees and a knowledge graph of uncertainties modestly outperforms standard chunk-based RAG in a proof-of-concept.

  6. FlexRAG: A Flexible and Comprehensive Framework for Retrieval-Augmented Generation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    FlexRAG is a modular, open-source RAG framework with text, multimodal, and web retrieval, plus evaluation tools and efficient memory-mapped indexing.

  7. Are Optimal Algorithms Still Optimal? Rethinking Sorting in LLM-Based Pairwise Ranking with Batching and Caching

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Under an LLM-inference cost model, Quicksort with batching uses roughly 44% fewer inference calls than Heapsort for pairwise document ranking, reversing the classical comparison-count ordering.

  8. Automating AI Failure Tracking: Semantic Association of Reports in AI Incident Database

    cs.CY 2025-07 conditional novelty 4.0 of 10

    Sentence-embedding retrieval ranks the correct AI Incident in the top three for about 98% of test reports when titles and descriptions are combined, but possible train/test leakage likely inflates that number.

  9. Exp4Fuse: A Rank Fusion Framework for Enhanced Sparse Retrieval using Large Language Model-based Query Expansion

    cs.IR 2025-06 conditional novelty 4.0 of 10

    Exp4Fuse improves sparse retrieval by fusing the ranked lists from the original query and an LLM-expanded query using a modified reciprocal rank fusion.

Pith tools