Pith. sign in

REVIEW 3 major objections 4 minor 2 cited by

State Space Models are Strong Text Rerankers

T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Mamba-based rerankers match transformer ranking quality at similar parameter counts, but trail transformers with Flash Attention in training and inference speed.

desk verdict Solid empirical benchmark of Mamba rerankers, but the 'competitive' claim is weaker than the title suggests once pretraining token budgets are accounted for. read the letter →

arxiv 2412.14354 v3 pith:T77R2CAV submitted 2024-12-18 cs.CL cs.IR

classification cs.CLcs.IR
keywords statespacemodelsMambaMamba-2textrerankinginformationretrievalefficiencybenchmarkFlashAttentioncross-encoderreranker
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether state space models, specifically Mamba-1 and Mamba-2, can perform the fine-grained query-document interaction required by text reranking, a task transformers have dominated. It benchmarks rerankers built on Mamba and on transformer language models of similar size on MS MARCO passage and document ranking, plus BEIR out-of-domain sets. The central finding is threefold: Mamba rerankers match transformer rerankers of comparable parameter count on ranking accuracy; they are slower in both training and single-pass inference than transformers using Flash Attention; and Mamba-2 improves on Mamba-1 on both axes. A sympathetic reader takes this as evidence that SSMs are a viable architectural alternative for reranking, with efficiency still the open gap.

What carries the argument

The central objects are the Mamba selective state space models. Mamba-1 makes the SSM parameters input-dependent and uses a hardware-aware selective scan, compressing context into a hidden state of size N; Mamba-2 restricts the A matrix to a scalar times identity, introduces an SSM head dimension analogous to transformer heads, and uses the structured state space duality algorithm so computation runs through matrix multiplications. The comparison's operational machinery is the standard cross-encoder reranker: query and document are concatenated into one input, a linear layer scores the final token, and training uses a softmax loss over one positive and hard negatives sampled from a first-stage retriever. That setup lets the authors attribute differences in ranking quality and speed primarily to the backbone architecture.

What would settle it

Train a Mamba and a transformer reranker from checkpoints pre-trained on the same corpus with matched token budgets and the same fine-tuning setup, then measure ranking metrics and per-query latency; if the ranking gap widens or the efficiency gap reverses under matched pre-training, the paper's architecture-level conclusions would not transfer, while if both hold, its claims are confirmed.

Watch

Extended reading notes

Core claim

The paper's central claim is that Mamba-based language models, despite compressing context into a fixed-size recurrent state, can learn the query-document interactions needed for text reranking and reach ranking quality comparable to transformer-based rerankers of similar scale. In passage reranking, Mamba-2-370M scores close to BERT-large on in-domain sets, and Mamba-2-1.3B averages 53.6 NDCG@10 across 13 BEIR datasets, slightly ahead of OPT-1.3B's 52.7. In document reranking, the best sub-1-billion-parameter model is the 780M Mamba-2 in the long-context setting. The paper also claims that the theoretical O(1) inference advantage of SSMs does not materialize on this task: Mamba models are slower in measured queries per second and substantially slower to train than transformers with Flash Attention, because reranking needs a single forward pass and Mamba's operators are less I/O-efficient. Within the SSM family, Mamba-2 outperforms Mamba-1 in both ranking performance and efficiency.

Load-bearing premise

The load-bearing premise is that publicly released checkpoints with very different pre-training token budgets and objectives, such as Mamba at roughly 300B tokens versus BERT at 3.3B-33B and Llama-3.2 at 15T tokens, can be compared as representative instantiations of their architectures, so observed ranking and efficiency differences are read as architectural rather than as artifacts of pre-training compute.

Editorial extensions

If this is right

  • Mamba-2 or similar SSMs could serve as the backbone of production rerankers when matching transformer ranking quality at a given parameter budget is the goal.
  • The advertised inference benefit of SSMs does not apply to reranking's single-pass workload, so IR deployments should expect SSM latency to be worse than Flash-Attention transformers until I/O optimizations improve.
  • Mamba-2's better training memory footprint lets larger SSM rerankers fit on the same GPU, as seen when Mamba-1 runs out of memory at 1.3B parameters while Mamba-2 trains.
  • Hybrid transformer-SSM models are a natural next step, since pure SSMs close the quality gap but not the efficiency gap in this workload.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the comparison mixes pre-training budgets, the ranking parity result is most conservatively read as 'SSMs can be fine-tuned into competitive rerankers,' not as proof that architecture alone determines ranking quality; an apples-to-apples pre-training study could change the ranking comparison.
  • The profiling result suggests a concrete optimization target: if Mamba-1's scalar-extraction operations and Mamba-2's MambaSplitConv1D-dominated load are replaced by fused tensor kernels, the measured inference gap may close independently of architecture choice.
  • The same benchmarking template could be applied to dense retrieval with SSM encoders and to hybrid Mamba-transformer models, where the efficiency result may differ because retrieval and generation workloads have different forward-pass profiles.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper presents a benchmarking study of Mamba-1 and Mamba-2 state-space models as text rerankers, comparing them with transformer-based models (BERT, RoBERTa, ELECTRA, BART, OPT, Llama-3.2) across passage and document reranking. The authors fine-tune public checkpoints with a standard cross-encoder setup, evaluate on MS MARCO Dev, TREC DL19/DL20, and 13 BEIR datasets, and additionally measure training throughput, inference speed, and operator-level execution time. The paper claims that (1) Mamba models achieve competitive ranking performance compared with transformers of similar size, (2) they are less efficient than flash-attention transformers in training and inference, and (3) Mamba-2 outperforms Mamba-1 in both performance and efficiency.

Significance. The study is a useful and extensive empirical benchmark: it covers passage and document ranking, in-domain and out-of-domain evaluation, a wide range of parameter scales, training and inference throughput, and operator-level profiling. The code release and the use of public datasets and checkpoints are strengths, and all central claims are direct measurements rather than derived predictions, so there is no circularity. If the conclusions were fully supported, the paper would give the IR community a practical answer about whether SSM rerankers can substitute for transformers in the single-pass scoring setting. However, the headline architecture-level claims are currently stronger than the controlled comparisons and tables justify, and the Mamba-2-versus-Mamba-1 performance claim is not uniformly supported by the data.

major comments (3)
  1. [§4.1, Table 1, Limitations] The paper's central architecture-level claim that Mamba models are 'competitive, comparable to transformer-based models of similar size' is confounded by the pre-training token budget. Table 1 shows the Mamba checkpoints use 300B tokens, whereas the transformer counterparts use 3.3B–33B (BERT, ELECTRA, BART), 180B (OPT), or 15T (Llama-3.2) tokens. Since reranking fine-tuning starts from these checkpoints, the observed ranking differences can be attributed to pre-training data and compute as much as to architecture. The Limitations section explicitly concedes this is not apples-to-apples. In addition, Tables 9 and 10 show that global batch size and epochs differ across the compared models (for example, Mamba-1-130M uses batch size 8 while Mamba-2-130M uses batch size 4, and Mamba-1-790M trains for 1 epoch while smaller models train for 2 epochs), and the authors themselves note that batch size affects reranking quality. Please add same-token decoder-only transformer baselines such as Pythia, or equivalently reframe the claim as one about publicly available checkpoints rather than about the architecture per se.
  2. [§4.2, §4.4, Tables 11 and 4] The claim that 'Mamba-2 outperforms Mamba-1 in both performance and efficiency' is not uniformly supported by the reported results. On the BEIR average in Table 11, Mamba-1-790M achieves 54.4 NDCG@10 versus Mamba-2-780M's 53.9, and Mamba-1-370M achieves 53.6 versus Mamba-2-370M's 53.0. In document reranking with the FirstP setting in Table 4, Mamba-1-370M beats Mamba-2-370M on Dev MRR@100 (42.5 vs 41.0) and DL19 NDCG@10 (67.8 vs 67.2). The performance superiority of Mamba-2 over Mamba-1 is therefore not a general result; it holds in the in-domain passage experiments but not in out-of-domain or long-document settings. Please either restrict the conclusion to the settings where it holds or provide significance tests or multiple-seed results that establish a consistent trend.
  3. [§4.4, §4.5, Tables 4 and 5] The efficiency claims are internally inconsistent with parts of the reported data. Section 4.4 states that Mamba-2 models 'in general require less GPU memory' during training, based on a FirstP run where Mamba-1-1.4B OOMs but Mamba-2-1.3B does not. However, the LongP block of Table 4 shows that both Mamba-1-1.4B and Mamba-2-1.3B OOM, and the Table 10 note confirms that Mamba-2-1.3B OOMs in the LongP setting despite all optimization techniques. The evidence supports only 'Mamba-2 fits in some configurations where Mamba-1 does not,' not a general memory-efficiency advantage. Similarly, the inference speed numbers in Table 5 do not uniformly support the abstract claim that Mamba models are less efficient than transformers with flash attention: Mamba-2-1.3B achieves 0.30 queries/second versus OPT-1.3B's 0.29 at length 512, and 0.29 versus 0.28 at length 1536. The efficiency conclusion is solid for Mamba-1 and for the 370M scale, but should be tempered for Mamba-2 at the 1.3B scale, and memory claims should report actual measured memory or be scoped to the configurations tested.
minor comments (4)
  1. [§4.2] The text says Llama-3.2-1B was 'trained on more tokens (15B)' but Table 1 lists the pre-training token count as 15T; please fix the unit.
  2. [Table 4] There are model-name inconsistencies in the document reranking table: 'Mamba1-790MD' lacks a hyphen, and 'Mamba-1-1.3BD' appears in the FirstP block while the model is called Mamba-1-1.4B in Table 1 and elsewhere. Please unify the names and sizes.
  3. [§4.5, Figure 1] The figure caption states that at batch size 8 all models except OPT-FlashAttn and Mamba-2 OOM with 48 GB VRAM, but the text says 'Mamba-1-370M does not train with batch size 8.' Please clarify whether the OOM statement applies to all Mamba-1 sizes and specify which model and batch size each throughput point corresponds to.
  4. [Table 2 caption] The caption contains grammatical errors ('other models reranks') and mixes the reranking thresholds for RankLlama and the other models; please revise for clarity.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: all central claims are direct empirical measurements on public benchmarks with no fitted quantity renamed as a prediction.

full rationale

This paper is a benchmarking study, and every central claim is a direct empirical measurement rather than a derivation from an input premise. The reranker is defined as a linear layer on top of a language model and is fine-tuned on MS MARCO and evaluated on held-out TREC DL19/DL20 and BEIR test sets; no parameter is fitted to the evaluation data and then reported as a prediction. The Mamba-vs-transformer performance comparisons are measured, not derived from the SSM equations in Section 2, and the efficiency comparisons in Section 4.5 and Table 5 are direct throughput and latency measurements. Self-citations such as RankMamba (Xu, 2024) appear only as contextual related work or as one of several sources for the standard hard-negative training setup, and the paper's conclusions do not rest on those citations. The acknowledged pretraining-token mismatch in the Limitations section is a genuine threat to the architecture-level comparison, but it is a validity and confounding concern, not a circularity: the reported scores are independent of the paper's own definitions and would stand or fall on external benchmark measurements. No equation reduces to its own input, and no fitted parameter is renamed as a prediction, so the circularity score is 0.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The central comparison rests on available pretrained checkpoints with widely different token budgets; several experimental configuration choices (batch size, learning rate) vary across models. No new model parameters or entities are introduced, but these uncontrolled variables make the architecture comparison approximate.

free parameters (3)
  • Pre-training token budget (confound) = Mamba 300B; OPT 180B; BERT/ELECTRA 3.3B-33B; Llama-3.2 15T; BART 33B
    Chosen by what is publicly released, not controlled across architectures. Acknowledged in Limitations. The paper's central 'comparable performance' claim bundles architecture with pre-training compute.
  • Global batch size = 4 or 8 depending on model (Tables 9-10)
    Selected by memory constraints; reranking performance is sensitive to batch size (per Zhuang et al. 2023), so unequal batch sizes are a confound in the architecture comparison.
  • Learning rate = 2e-5 for base models; 1e-5 for large/1B+ models
    Chosen per model size following common practice; not controlled, could affect relative results.
assumptions (3)
  • domain assumption Publicly released checkpoints on Huggingface Hub fairly represent each architecture family
    The comparison uses off-the-shelf checkpoints; no controlled pre-training is performed (acknowledged in Limitations).
  • domain assumption Concatenating query and document and using a linear head on [EOS]/[CLS] is an equally fair protocol for SSMs and transformers
    Mamba and OPT are unidirectional decoder-only models while BERT/ELECTRA/BART use bidirectional encoding; Section 4.2 hypothesizes bidirectionality helps reranking, so the protocol may favor the encoder/encoder-decoder family.
  • domain assumption Huggingface implementations of Mamba kernels are representative of Mamba's practical throughput
    Efficiency measurements in Section 4.5 and Figure 2 rely on the available software stack; official or future kernels could change the speed comparison.

how reviews work

0 comments
Cite this review

Pith. "Pith review of State Space Models are Strong Text Rerankers." pith.science (2026). https://pith.science/paper/T77R2CAV

@misc{pith2026241214354,
  author       = {Pith},
  title        = {Pith review of: State Space Models are Strong Text Rerankers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/T77R2CAV}},
  note         = {Machine review of arXiv:2412.14354}
}
abstract

Transformers dominate NLP and IR; but their inference inefficiencies and challenges in extrapolating to longer contexts have sparked interest in alternative model architectures. Among these, state space models (SSMs) like Mamba offer promising advantages, particularly $O(1)$ time complexity in inference. Despite their potential, SSMs' effectiveness at text reranking -- a task requiring fine-grained query-document interaction and long-context understanding -- remains underexplored. This study benchmarks SSM-based architectures (specifically, Mamba-1 and Mamba-2) against transformer-based models across various scales, architectures, and pre-training objectives, focusing on performance and efficiency in text reranking tasks. We find that (1) Mamba architectures achieve competitive text ranking performance, comparable to transformer-based models of similar size; (2) they are less efficient in training and inference compared to transformers with flash attention; and (3) Mamba-2 outperforms Mamba-1 in both performance and efficiency. These results underscore the potential of state space models as a transformer alternative and highlight areas for improvement in future IR applications.

Figures

Figures reproduced from arXiv: 2412.14354 by the authors.

Figure 1
Figure 1. Training throughput comparison between models≈330M. For batch_size=8, all models except OPT-FlashAttn and Mamba-2 run out of memory with a 48 GB VRAM GPU. in terms of the task performance, Mamba-based rerankers are comparable to their Transformer￾based counterparts for every parameter budget. No￾tably, among the sub-1 billion parameter models, the best model is the 780 million Mamba-2 model trained with 1536 context… view at source ↗
Figure 2
Figure 2. Inference profiling results for Mamba models versus OPT models of similar size. [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Understanding What State Space Models Learn About Code

    cs.AI 2026-02 conditional novelty 6.0 of 10

    SSM code models capture code syntax and semantics better than Transformers before fine-tuning, forget short-range structure when fine-tuned on type inference, and an added high-frequency path or more kernels recovers ...

  2. MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation

    cs.CL 2024-12 conditional novelty 4.0 of 10

    MAIN-RAG filters noisy retrieved documents with multiple LLM agents and an adaptive score threshold, improving QA accuracy by 2 to 11 percent over standard RAG.

Reference graph

Works this paper leans on

86 extracted references · 26 canonical work pages · cited by 2 Pith papers

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, Joseph Chee Chang, Kyle Lo, Luca Soldaini, Sergey Feldman, Mike D'arcy, et al. 2024. Openscholar: Synthesizing scientific literature with retrieval-augmented lms. arXiv preprint arXiv:2411.14199

  4. [4]

    Anas Awadalla, Mitchell Wortsman, Gabriel Ilharco, Sewon Min, Ian Magnusson, Hannaneh Hajishirzi, and Ludwig Schmidt. 2022. Exploring the landscape of distributional robustness for question answering models. arXiv preprint arXiv:2210.12517

  5. [5]

    Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al. 2016. Ms marco: A human generated machine reading comprehension dataset. arXiv preprint arXiv:1611.09268

  6. [6]

    Alexander Bondarenko, Maik Fr \"o be, Meriem Beloucif, Lukas Gienapp, Yamen Ajjour, Alexander Panchenko, Chris Biemann, Benno Stein, Henning Wachsmuth, Martin Potthast, et al. 2020. Overview of touch \'e 2020: argument retrieval. In Experimental IR Meets Multilinguality, Multimodality, and Interaction: 11th International Conference of the CLEF Association...

  7. [7]

    Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. 2016. A full-text learning to rank dataset for medical information retrieval. In Advances in Information Retrieval: 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20--23, 2016. Proceedings 38, pages 716--722. Springer

  8. [8]

    Leonid Boytsov, Tianyi Lin, Fangwei Gao, Yutian Zhao, Jeffrey Huang, and Eric Nyberg. 2022. Understanding performance of long-document ranking models through comprehensive evaluation and leaderboarding. arXiv preprint arXiv:2207.01262

Show all 86 references
  1. [9]

    Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020. Electra: Pre-training text encoders as discriminators rather than generators. In International Conference on Learning Representations

  2. [10]

    Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S Weld. 2020. Specter: Document-level representation learning using citation-informed transformers. arXiv preprint arXiv:2004.07180

  3. [11]

    Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2021. Overview of the trec 2020 deep learning track. corr abs/2102.07662 (2021). arXiv preprint arXiv:2102.07662

  4. [12]

    Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M Voorhees. 2020. Overview of the trec 2019 deep learning track. arXiv preprint arXiv:2003.07820

  5. [13]

    Zhuyun Dai and Jamie Callan. 2019. Deeper text understanding for ir with contextual neural language modeling. In Proceedings of the 42nd international ACM SIGIR conference on research and development in information retrieval, pages 985--988

  6. [14]

    Tri Dao. 2024. https://openreview.net/forum?id=mZn2Xyh9Ec Flashattention-2: Faster attention with better parallelism and work partitioning . In The Twelfth International Conference on Learning Representations

  7. [15]

    Tri Dao and Albert Gu. 2024. https://proceedings.mlr.press/v235/dao24a.html Transformers are SSM s: Generalized models and efficient algorithms through structured state space duality . In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proce...

  8. [16]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...

  9. [17]

    Thomas Diggelmann, Jordan Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold. 2020. Climate-fever: A dataset for verification of real-world climate claims. arXiv preprint arXiv:2012.00614

  10. [18]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783

  11. [19]

    Luyu Gao and Jamie Callan. 2022. Unsupervised corpus aware language model pre-training for dense passage retrieval. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2843--2853

  12. [20]

    Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. Rethink training of bert rerankers in multi-stage retrieval pipeline. In Advances in Information Retrieval: 43rd European Conference on IR Research, ECIR 2021, Virtual Event, March 28--April 1, 2021, Proceedings, Part II 43, pages ...

  13. [21]

    Paolo Glorioso, Quentin Anthony, Yury Tokpanov, James Whittington, Jonathan Pilault, Adam Ibrahim, and Beren Millidge. 2024. Zamba: A compact 7b ssm hybrid model. arXiv preprint arXiv:2405.16712

  14. [22]

    Albert Gu and Tri Dao. 2023. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752

  15. [23]

    Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher R \'e . 2020. Hippo: Recurrent memory with optimal polynomial projections. Advances in neural information processing systems, 33:1474--1487

  16. [24]

    Albert Gu, Karan Goel, and Christopher Re. 2021 a . Efficiently modeling long sequences with structured state spaces. In International Conference on Learning Representations

  17. [25]

    Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher R \'e . 2021 b . Combining recurrent, convolutional, and continuous-time models with linear state space layers. Advances in neural information processing systems, 34:572--585

  18. [26]

    Ankit Gupta, Albert Gu, and Jonathan Berant. 2022. Diagonal state spaces are as effective as structured state spaces. Advances in Neural Information Processing Systems, 35:22982--22994

  19. [27]

    Ashim Gupta, Carter Blum, Temma Choji, Yingjie Fei, Shalin Shah, Alakananda Vempala, and Vivek Srikumar. 2023. Don’t retrain, just rewrite: Countering adversarial perturbations by rewriting text. In Proceedings of the 61st Annual Meeting of the Association for Computational Li...

  20. [28]

    Ashim Gupta, Rishanth Rajendhran, Nathan Stringham, Vivek Srikumar, and Ana Marasovi \'c . 2024 a . Whispers of doubt amidst echoes of triumph in nlp robustness. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistic...

  21. [29]

    Ashim Gupta, Sina Mahdipour Saravani, P Sadayappan, and Vivek Srikumar. 2024 b . An empirical investigation of matrix factorization methods for pre-trained transformers. arXiv preprint arXiv:2406.11307

  22. [30]

    Faegheh Hasibi, Fedor Nikolaev, Chenyan Xiong, Krisztian Balog, Svein Erik Bratsberg, Alexander Kotov, and Jamie Callan. 2017. Dbpedia-entity v2: a test collection for entity search. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in I...

  23. [31]

    Sebastian Hofst \"a tter, Bhaskar Mitra, Hamed Zamani, Nick Craswell, and Allan Hanbury. 2021. Intra-document cascading: learning to select passages for neural document ranking. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Inform...

  24. [32]

    Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021. Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations

  25. [33]

    Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.550 Dense passage retrieval for open-domain question answering . In Proceedings of the 2020 Conference on Empiric...

  26. [34]

    Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav...

  27. [35]

    Barak Lenz, Alan Arazi, Amir Bergman, Avshalom Manevich, Barak Peleg, Ben Aviram, Chen Almagor, Clara Fridman, Dan Padnos, et al. 2024. Jamba-1.5: Hybrid transformer-mamba models at scale. arXiv preprint arXiv:2408.12570

  28. [36]

    Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. https://doi.org/10.18653/v1/2020.acl-main.703 BART : Denoising sequence-to-sequence pre-training for natural language generation, translatio...

  29. [37]

    Canjia Li, Andrew Yates, Sean MacAvaney, Ben He, and Yingfei Sun. 2023. Parade: Passage representation aggregation fordocument reranking. ACM Transactions on Information Systems, 42(2):1--26

  30. [38]

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2023. Holistic evaluation of language models. Trans. Mach. Learn. Res

  31. [39]

    Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, et al. 2024. Jamba: A hybrid transformer-mamba language model. arXiv preprint arXiv:2403.19887

  32. [40]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692

  33. [41]

    Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2023. Fine-tuning llama for multi-stage text retrieval. arXiv preprint arXiv:2310.08319

  34. [42]

    Macedo Maia, Siegfried Handschuh, Andr \'e Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018. Www'18 open challenge: financial opinion mining and question answering. In Companion proceedings of the the web conference 2018, pages 1941--1942

  35. [43]

    Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al. 2022. Text and code embeddings by contrastive pre-training. arXiv preprint arXiv:2201.10005

  36. [44]

    Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith Hall, Daniel Cer, and Yinfei Yang. 2022. https://doi.org/10.18653/v1/2022.findings-acl.146 Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models . In Findings of the Association for Computa...

  37. [45]

    Rodrigo Nogueira and Kyunghyun Cho. 2019. Passage re-ranking with bert. arXiv preprint arXiv:1901.04085

  38. [46]

    Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020. Document ranking with a pretrained sequence-to-sequence model. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 708--718

  39. [47]

    Rodrigo Nogueira, Wei Yang, Jimmy Lin, and Kyunghyun Cho. 2019. Document expansion by query prediction. arXiv preprint arXiv:1904.08375

  40. [48]

    Nvidia. 2025. Nemotron-h: A family of accurate, efficient hybrid mamba-transformer models

  41. [49]

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32

  42. [50]

    Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Leon Derczynski, Xingjian Du, Matteo Grella, Kranthi Gv, Xuzheng He, Haowen Hou, Przemyslaw Kazienko, Jan Kocon, Jiaming Kong, Bart omiej Koptyra, Hay...

  43. [51]

    Ronak Pradeep, Rodrigo Nogueira, and Jimmy Lin. 2021. The expando-mono-duo design pattern for text ranking with pretrained sequence-to-sequence models. arXiv preprint arXiv:2101.05667

  44. [52]

    Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun, and Yiran Zhong. 2024. Hgrn2: Gated linear rnns with state expansion. arXiv preprint arXiv:2404.07904

  45. [53]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1--67

  46. [54]

    Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al. 1995. Okapi at trec-3. Nist Special Publication Sp, 109:109

  47. [55]

    Hinrich Sch \"u tze, Christopher D Manning, and Prabhakar Raghavan. 2008. Introduction to information retrieval, volume 39. Cambridge University Press Cambridge

  48. [56]

    Jimmy TH Smith, Andrew Warrington, and Scott Linderman. 2022. Simplified state space layers for sequence modeling. In The Eleventh International Conference on Learning Representations

  49. [57]

    Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063

  50. [58]

    Nandan Thakur, Nils Reimers, Andreas R \"u ckl \'e , Abhishek Srivastava, and Iryna Gurevych. 2021. Beir: A heterogeneous benchmark for zero-shot evaluation of information retrieval models

  51. [59]

    James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. Fever: a large-scale dataset for fact extraction and verification. arXiv preprint arXiv:1803.05355

  52. [60]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288

  53. [61]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30

  54. [62]

    Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2021. Trec-covid: constructing a pandemic information retrieval test collection. In ACM SIGIR Forum, volume 54, pages 1--12. ACM New York, NY, USA

  55. [63]

    Henning Wachsmuth, Shahbaz Syed, and Benno Stein. 2018. Retrieval of the best counterargument without prior topic knowledge. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 241--251

  56. [64]

    David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or fiction: Verifying scientific claims. arXiv preprint arXiv:2004.14974

  57. [65]

    Roger Waleffe, Wonmin Byeon, Duncan Riach, Brandon Norick, Vijay Korthikanti, Tri Dao, Albert Gu, Ali Hatamizadeh, Sudhakar Singh, Deepak Narayanan, et al. 2024. An empirical study of mamba-based language models. arXiv preprint arXiv:2406.07887

  58. [66]

    Junxiong Wang, Daniele Paliotta, Avner May, Alexander Rush, and Tri Dao. 2024. The mamba in the llama: Distilling and accelerating hybrid models. Advances in Neural Information Processing Systems, 37:62432--62457

  59. [67]

    Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022. Simlm: Pre-training with representation bottleneck for dense passage retrieval. arXiv preprint arXiv:2207.02578

  60. [68]

    T Wolf. 2019. Huggingface's transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771

  61. [69]

    Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighof. 2023. C-pack: Packaged resources to advance general chinese embedding. arXiv preprint arXiv:2309.07597

  62. [70]

    Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, et al. 2024. Effective long-context scaling of foundation models. In Proceedings of the 2024 Conference of the North American ...

  63. [71]

    Zhichao Xu. 2023. Context-aware decoding reduces hallucination in query-focused summarization. arXiv preprint arXiv:2312.14335

  64. [72]

    Zhichao Xu. 2024. Rankmamba, benchmarking mamba's document ranking performance in the era of transformers. arXiv preprint arXiv:2403.18276

  65. [73]

    Zhichao Xu and Daniel Cohen. 2023. A lightweight constrained generation alternative for query-focused summarization. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 1745--1749

  66. [74]

    Zhichao Xu, Daniel Cohen, Bei Wang, and Vivek Srikumar. 2024 a . https://doi.org/10.18653/v1/2024.findings-naacl.167 In-context example ordering guided by label distributions . In Findings of the Association for Computational Linguistics: NAACL 2024, pages 2623--2640, Mexico C...

  67. [75]

    Zhichao Xu, Ashim Gupta, Tao Li, Oliver Bentham, and Vivek Srikumar. 2024 b . https://doi.org/10.18653/v1/2024.findings-emnlp.901 Beyond perplexity: Multi-dimensional safety evaluation of LLM compression . In Findings of the Association for Computational Linguistics: EMNLP 202...

  68. [76]

    Zhichao Xu, Hemank Lamba, Qingyao Ai, Joel Tetreault, and Alex Jaimes. 2024 c . https://doi.org/10.1145/3664190.3672508 Cfe2: Counterfactual editing for search result explanation . In Proceedings of the 2024 ACM SIGIR International Conference on Theory of Information Retrieval...

  69. [77]

    Zhichao Xu, Fengran Mo, Zhiqi Huang, Crystina Zhang, Puxuan Yu, Bei Wang, Jimmy Lin, and Vivek Srikumar. 2025. A survey of model architectures in information retrieval. arXiv preprint arXiv:2502.14822

  70. [78]

    Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. 2023. Gated linear attention transformers with hardware-efficient training. arXiv preprint arXiv:2312.06635

  71. [79]

    Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim. 2024. https://openreview.net/forum?id=y8Rm4VNRPH Parallelizing linear transformers with the delta rule over sequence length . In The Thirty-eighth Annual Conference on Neural Information Processing Systems

  72. [80]

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600

  73. [81]

    Andrew Yates, Rodrigo Nogueira, and Jimmy Lin. 2021. Pretrained transformers for text ranking: Bert and beyond. In Proceedings of the 14th ACM International Conference on web search and data mining, pages 1154--1156

  74. [82]

    Hanqi Zhang, Chong Chen, Lang Mei, Qi Liu, and Jiaxin Mao. 2024. Mamba retriever: Utilizing mamba for effective and efficient dense retrieval. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, pages 4268--4272

  75. [83]

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068

  76. [84]

    Yue Zhang, ChengCheng Hu, Yuqi Liu, Hui Fang, and Jimmy Lin. 2021. https://doi.org/10.18653/v1/2021.sustainlp-1.8 Learning to rank in the age of M uppets: Effectiveness -- efficiency tradeoffs in multi-stage ranking . In Proceedings of the Second Workshop on Simple and Efficie...

  77. [85]

    Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. 2024. Vision mamba: Efficient visual representation learning with bidirectional state space model. In Forty-first International Conference on Machine Learning

  78. [86]

    Honglei Zhuang, Zhen Qin, Rolf Jagerman, Kai Hui, Ji Ma, Jing Lu, Jianmo Ni, Xuanhui Wang, and Michael Bendersky. 2023. Rankt5: Fine-tuning t5 for text ranking with ranking losses. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Inf...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.