REVIEW 3 major objections 4 minor 2 cited by
State Space Models are Strong Text Rerankers
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Mamba-based rerankers match transformer ranking quality at similar parameter counts, but trail transformers with Flash Attention in training and inference speed.
desk verdict Solid empirical benchmark of Mamba rerankers, but the 'competitive' claim is weaker than the title suggests once pretraining token budgets are accounted for. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central objects are the Mamba selective state space models. Mamba-1 makes the SSM parameters input-dependent and uses a hardware-aware selective scan, compressing context into a hidden state of size N; Mamba-2 restricts the A matrix to a scalar times identity, introduces an SSM head dimension analogous to transformer heads, and uses the structured state space duality algorithm so computation runs through matrix multiplications. The comparison's operational machinery is the standard cross-encoder reranker: query and document are concatenated into one input, a linear layer scores the final token, and training uses a softmax loss over one positive and hard negatives sampled from a first-stage retriever. That setup lets the authors attribute differences in ranking quality and speed primarily to the backbone architecture.
What would settle it
Train a Mamba and a transformer reranker from checkpoints pre-trained on the same corpus with matched token budgets and the same fine-tuning setup, then measure ranking metrics and per-query latency; if the ranking gap widens or the efficiency gap reverses under matched pre-training, the paper's architecture-level conclusions would not transfer, while if both hold, its claims are confirmed.
Extended reading notes
Core claim
The paper's central claim is that Mamba-based language models, despite compressing context into a fixed-size recurrent state, can learn the query-document interactions needed for text reranking and reach ranking quality comparable to transformer-based rerankers of similar scale. In passage reranking, Mamba-2-370M scores close to BERT-large on in-domain sets, and Mamba-2-1.3B averages 53.6 NDCG@10 across 13 BEIR datasets, slightly ahead of OPT-1.3B's 52.7. In document reranking, the best sub-1-billion-parameter model is the 780M Mamba-2 in the long-context setting. The paper also claims that the theoretical O(1) inference advantage of SSMs does not materialize on this task: Mamba models are slower in measured queries per second and substantially slower to train than transformers with Flash Attention, because reranking needs a single forward pass and Mamba's operators are less I/O-efficient. Within the SSM family, Mamba-2 outperforms Mamba-1 in both ranking performance and efficiency.
Load-bearing premise
The load-bearing premise is that publicly released checkpoints with very different pre-training token budgets and objectives, such as Mamba at roughly 300B tokens versus BERT at 3.3B-33B and Llama-3.2 at 15T tokens, can be compared as representative instantiations of their architectures, so observed ranking and efficiency differences are read as architectural rather than as artifacts of pre-training compute.
Editorial extensions
If this is right
- Mamba-2 or similar SSMs could serve as the backbone of production rerankers when matching transformer ranking quality at a given parameter budget is the goal.
- The advertised inference benefit of SSMs does not apply to reranking's single-pass workload, so IR deployments should expect SSM latency to be worse than Flash-Attention transformers until I/O optimizations improve.
- Mamba-2's better training memory footprint lets larger SSM rerankers fit on the same GPU, as seen when Mamba-1 runs out of memory at 1.3B parameters while Mamba-2 trains.
- Hybrid transformer-SSM models are a natural next step, since pure SSMs close the quality gap but not the efficiency gap in this workload.
Reading between the lines
- Because the comparison mixes pre-training budgets, the ranking parity result is most conservatively read as 'SSMs can be fine-tuned into competitive rerankers,' not as proof that architecture alone determines ranking quality; an apples-to-apples pre-training study could change the ranking comparison.
- The profiling result suggests a concrete optimization target: if Mamba-1's scalar-extraction operations and Mamba-2's MambaSplitConv1D-dominated load are replaced by fused tensor kernels, the measured inference gap may close independently of architecture choice.
- The same benchmarking template could be applied to dense retrieval with SSM encoders and to hybrid Mamba-transformer models, where the efficiency result may differ because retrieval and generation workloads have different forward-pass profiles.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a benchmarking study of Mamba-1 and Mamba-2 state-space models as text rerankers, comparing them with transformer-based models (BERT, RoBERTa, ELECTRA, BART, OPT, Llama-3.2) across passage and document reranking. The authors fine-tune public checkpoints with a standard cross-encoder setup, evaluate on MS MARCO Dev, TREC DL19/DL20, and 13 BEIR datasets, and additionally measure training throughput, inference speed, and operator-level execution time. The paper claims that (1) Mamba models achieve competitive ranking performance compared with transformers of similar size, (2) they are less efficient than flash-attention transformers in training and inference, and (3) Mamba-2 outperforms Mamba-1 in both performance and efficiency.
Significance. The study is a useful and extensive empirical benchmark: it covers passage and document ranking, in-domain and out-of-domain evaluation, a wide range of parameter scales, training and inference throughput, and operator-level profiling. The code release and the use of public datasets and checkpoints are strengths, and all central claims are direct measurements rather than derived predictions, so there is no circularity. If the conclusions were fully supported, the paper would give the IR community a practical answer about whether SSM rerankers can substitute for transformers in the single-pass scoring setting. However, the headline architecture-level claims are currently stronger than the controlled comparisons and tables justify, and the Mamba-2-versus-Mamba-1 performance claim is not uniformly supported by the data.
major comments (3)
- [§4.1, Table 1, Limitations] The paper's central architecture-level claim that Mamba models are 'competitive, comparable to transformer-based models of similar size' is confounded by the pre-training token budget. Table 1 shows the Mamba checkpoints use 300B tokens, whereas the transformer counterparts use 3.3B–33B (BERT, ELECTRA, BART), 180B (OPT), or 15T (Llama-3.2) tokens. Since reranking fine-tuning starts from these checkpoints, the observed ranking differences can be attributed to pre-training data and compute as much as to architecture. The Limitations section explicitly concedes this is not apples-to-apples. In addition, Tables 9 and 10 show that global batch size and epochs differ across the compared models (for example, Mamba-1-130M uses batch size 8 while Mamba-2-130M uses batch size 4, and Mamba-1-790M trains for 1 epoch while smaller models train for 2 epochs), and the authors themselves note that batch size affects reranking quality. Please add same-token decoder-only transformer baselines such as Pythia, or equivalently reframe the claim as one about publicly available checkpoints rather than about the architecture per se.
- [§4.2, §4.4, Tables 11 and 4] The claim that 'Mamba-2 outperforms Mamba-1 in both performance and efficiency' is not uniformly supported by the reported results. On the BEIR average in Table 11, Mamba-1-790M achieves 54.4 NDCG@10 versus Mamba-2-780M's 53.9, and Mamba-1-370M achieves 53.6 versus Mamba-2-370M's 53.0. In document reranking with the FirstP setting in Table 4, Mamba-1-370M beats Mamba-2-370M on Dev MRR@100 (42.5 vs 41.0) and DL19 NDCG@10 (67.8 vs 67.2). The performance superiority of Mamba-2 over Mamba-1 is therefore not a general result; it holds in the in-domain passage experiments but not in out-of-domain or long-document settings. Please either restrict the conclusion to the settings where it holds or provide significance tests or multiple-seed results that establish a consistent trend.
- [§4.4, §4.5, Tables 4 and 5] The efficiency claims are internally inconsistent with parts of the reported data. Section 4.4 states that Mamba-2 models 'in general require less GPU memory' during training, based on a FirstP run where Mamba-1-1.4B OOMs but Mamba-2-1.3B does not. However, the LongP block of Table 4 shows that both Mamba-1-1.4B and Mamba-2-1.3B OOM, and the Table 10 note confirms that Mamba-2-1.3B OOMs in the LongP setting despite all optimization techniques. The evidence supports only 'Mamba-2 fits in some configurations where Mamba-1 does not,' not a general memory-efficiency advantage. Similarly, the inference speed numbers in Table 5 do not uniformly support the abstract claim that Mamba models are less efficient than transformers with flash attention: Mamba-2-1.3B achieves 0.30 queries/second versus OPT-1.3B's 0.29 at length 512, and 0.29 versus 0.28 at length 1536. The efficiency conclusion is solid for Mamba-1 and for the 370M scale, but should be tempered for Mamba-2 at the 1.3B scale, and memory claims should report actual measured memory or be scoped to the configurations tested.
minor comments (4)
- [§4.2] The text says Llama-3.2-1B was 'trained on more tokens (15B)' but Table 1 lists the pre-training token count as 15T; please fix the unit.
- [Table 4] There are model-name inconsistencies in the document reranking table: 'Mamba1-790MD' lacks a hyphen, and 'Mamba-1-1.3BD' appears in the FirstP block while the model is called Mamba-1-1.4B in Table 1 and elsewhere. Please unify the names and sizes.
- [§4.5, Figure 1] The figure caption states that at batch size 8 all models except OPT-FlashAttn and Mamba-2 OOM with 48 GB VRAM, but the text says 'Mamba-1-370M does not train with batch size 8.' Please clarify whether the OOM statement applies to all Mamba-1 sizes and specify which model and batch size each throughput point corresponds to.
- [Table 2 caption] The caption contains grammatical errors ('other models reranks') and mixes the reranking thresholds for RankLlama and the other models; please revise for clarity.
Circularity Check
No circularity: all central claims are direct empirical measurements on public benchmarks with no fitted quantity renamed as a prediction.
full rationale
This paper is a benchmarking study, and every central claim is a direct empirical measurement rather than a derivation from an input premise. The reranker is defined as a linear layer on top of a language model and is fine-tuned on MS MARCO and evaluated on held-out TREC DL19/DL20 and BEIR test sets; no parameter is fitted to the evaluation data and then reported as a prediction. The Mamba-vs-transformer performance comparisons are measured, not derived from the SSM equations in Section 2, and the efficiency comparisons in Section 4.5 and Table 5 are direct throughput and latency measurements. Self-citations such as RankMamba (Xu, 2024) appear only as contextual related work or as one of several sources for the standard hard-negative training setup, and the paper's conclusions do not rest on those citations. The acknowledged pretraining-token mismatch in the Limitations section is a genuine threat to the architecture-level comparison, but it is a validity and confounding concern, not a circularity: the reported scores are independent of the paper's own definitions and would stand or fall on external benchmark measurements. No equation reduces to its own input, and no fitted parameter is renamed as a prediction, so the circularity score is 0.
Assumptions & free parameters
free parameters (3)
- Pre-training token budget (confound) =
Mamba 300B; OPT 180B; BERT/ELECTRA 3.3B-33B; Llama-3.2 15T; BART 33B
- Global batch size =
4 or 8 depending on model (Tables 9-10)
- Learning rate =
2e-5 for base models; 1e-5 for large/1B+ models
assumptions (3)
- domain assumption Publicly released checkpoints on Huggingface Hub fairly represent each architecture family
- domain assumption Concatenating query and document and using a linear head on [EOS]/[CLS] is an equally fair protocol for SSMs and transformers
- domain assumption Huggingface implementations of Mamba kernels are representative of Mamba's practical throughput
Cite this review
Pith. "Pith review of State Space Models are Strong Text Rerankers." pith.science (2026). https://pith.science/paper/T77R2CAV
@misc{pith2026241214354,
author = {Pith},
title = {Pith review of: State Space Models are Strong Text Rerankers},
year = {2026},
howpublished = {\url{https://pith.science/paper/T77R2CAV}},
note = {Machine review of arXiv:2412.14354}
}
abstract
Transformers dominate NLP and IR; but their inference inefficiencies and challenges in extrapolating to longer contexts have sparked interest in alternative model architectures. Among these, state space models (SSMs) like Mamba offer promising advantages, particularly $O(1)$ time complexity in inference. Despite their potential, SSMs' effectiveness at text reranking -- a task requiring fine-grained query-document interaction and long-context understanding -- remains underexplored. This study benchmarks SSM-based architectures (specifically, Mamba-1 and Mamba-2) against transformer-based models across various scales, architectures, and pre-training objectives, focusing on performance and efficiency in text reranking tasks. We find that (1) Mamba architectures achieve competitive text ranking performance, comparable to transformer-based models of similar size; (2) they are less efficient in training and inference compared to transformers with flash attention; and (3) Mamba-2 outperforms Mamba-1 in both performance and efficiency. These results underscore the potential of state space models as a transformer alternative and highlight areas for improvement in future IR applications.
Figures
Forward citations
Cited by 2 Pith papers
-
Towards Understanding What State Space Models Learn About Code
SSM code models capture code syntax and semantics better than Transformers before fine-tuning, forget short-range structure when fine-tuned on type inference, and an added high-frequency path or more kernels recovers ...
-
MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation
MAIN-RAG filters noisy retrieved documents with multiple LLM agents and an adaptive score threshold, improving QA accuracy by 2 to 11 percent over standard RAG.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, Joseph Chee Chang, Kyle Lo, Luca Soldaini, Sergey Feldman, Mike D'arcy, et al. 2024. Openscholar: Synthesizing scientific literature with retrieval-augmented lms. arXiv preprint arXiv:2411.14199
arXiv 2024
-
[4]
Anas Awadalla, Mitchell Wortsman, Gabriel Ilharco, Sewon Min, Ian Magnusson, Hannaneh Hajishirzi, and Ludwig Schmidt. 2022. Exploring the landscape of distributional robustness for question answering models. arXiv preprint arXiv:2210.12517
arXiv 2022
-
[5]
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al. 2016. Ms marco: A human generated machine reading comprehension dataset. arXiv preprint arXiv:1611.09268
arXiv 2016
-
[6]
Alexander Bondarenko, Maik Fr \"o be, Meriem Beloucif, Lukas Gienapp, Yamen Ajjour, Alexander Panchenko, Chris Biemann, Benno Stein, Henning Wachsmuth, Martin Potthast, et al. 2020. Overview of touch \'e 2020: argument retrieval. In Experimental IR Meets Multilinguality, Multimodality, and Interaction: 11th International Conference of the CLEF Association...
2020
-
[7]
Vera Boteva, Demian Gholipour, Artem Sokolov, and Stefan Riezler. 2016. A full-text learning to rank dataset for medical information retrieval. In Advances in Information Retrieval: 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20--23, 2016. Proceedings 38, pages 716--722. Springer
2016
- [8]
Show all 86 references
-
[9]
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020. Electra: Pre-training text encoders as discriminators rather than generators. In International Conference on Learning Representations
2020
-
[10]
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S Weld. 2020. Specter: Document-level representation learning using citation-informed transformers. arXiv preprint arXiv:2004.07180
2020 arXiv
-
[11]
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2021. Overview of the trec 2020 deep learning track. corr abs/2102.07662 (2021). arXiv preprint arXiv:2102.07662
2021 arXiv
-
[12]
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Ellen M Voorhees. 2020. Overview of the trec 2019 deep learning track. arXiv preprint arXiv:2003.07820
2020 arXiv
-
[13]
Zhuyun Dai and Jamie Callan. 2019. Deeper text understanding for ir with contextual neural language modeling. In Proceedings of the 42nd international ACM SIGIR conference on research and development in information retrieval, pages 985--988
2019
-
[14]
Tri Dao. 2024. https://openreview.net/forum?id=mZn2Xyh9Ec Flashattention-2: Faster attention with better parallelism and work partitioning . In The Twelfth International Conference on Learning Representations
2024
-
[15]
Tri Dao and Albert Gu. 2024. https://proceedings.mlr.press/v235/dao24a.html Transformers are SSM s: Generalized models and efficient algorithms through structured state space duality . In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proce...
2024
-
[16]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[17]
Thomas Diggelmann, Jordan Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold. 2020. Climate-fever: A dataset for verification of real-world climate claims. arXiv preprint arXiv:2012.00614
2020 arXiv
-
[18]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783
2024 arXiv
-
[19]
Luyu Gao and Jamie Callan. 2022. Unsupervised corpus aware language model pre-training for dense passage retrieval. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2843--2853
2022
-
[20]
Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. Rethink training of bert rerankers in multi-stage retrieval pipeline. In Advances in Information Retrieval: 43rd European Conference on IR Research, ECIR 2021, Virtual Event, March 28--April 1, 2021, Proceedings, Part II 43, pages ...
2021
-
[21]
Paolo Glorioso, Quentin Anthony, Yury Tokpanov, James Whittington, Jonathan Pilault, Adam Ibrahim, and Beren Millidge. 2024. Zamba: A compact 7b ssm hybrid model. arXiv preprint arXiv:2405.16712
2024 arXiv
-
[22]
Albert Gu and Tri Dao. 2023. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752
2023 arXiv
-
[23]
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher R \'e . 2020. Hippo: Recurrent memory with optimal polynomial projections. Advances in neural information processing systems, 33:1474--1487
2020
-
[24]
Albert Gu, Karan Goel, and Christopher Re. 2021 a . Efficiently modeling long sequences with structured state spaces. In International Conference on Learning Representations
2021
-
[25]
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher R \'e . 2021 b . Combining recurrent, convolutional, and continuous-time models with linear state space layers. Advances in neural information processing systems, 34:572--585
2021
-
[26]
Ankit Gupta, Albert Gu, and Jonathan Berant. 2022. Diagonal state spaces are as effective as structured state spaces. Advances in Neural Information Processing Systems, 35:22982--22994
2022
-
[27]
Ashim Gupta, Carter Blum, Temma Choji, Yingjie Fei, Shalin Shah, Alakananda Vempala, and Vivek Srikumar. 2023. Don’t retrain, just rewrite: Countering adversarial perturbations by rewriting text. In Proceedings of the 61st Annual Meeting of the Association for Computational Li...
2023
-
[28]
Ashim Gupta, Rishanth Rajendhran, Nathan Stringham, Vivek Srikumar, and Ana Marasovi \'c . 2024 a . Whispers of doubt amidst echoes of triumph in nlp robustness. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistic...
2024
-
[29]
Ashim Gupta, Sina Mahdipour Saravani, P Sadayappan, and Vivek Srikumar. 2024 b . An empirical investigation of matrix factorization methods for pre-trained transformers. arXiv preprint arXiv:2406.11307
2024 arXiv
-
[30]
Faegheh Hasibi, Fedor Nikolaev, Chenyan Xiong, Krisztian Balog, Svein Erik Bratsberg, Alexander Kotov, and Jamie Callan. 2017. Dbpedia-entity v2: a test collection for entity search. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in I...
2017
-
[31]
Sebastian Hofst \"a tter, Bhaskar Mitra, Hamed Zamani, Nick Craswell, and Allan Hanbury. 2021. Intra-document cascading: learning to select passages for neural document ranking. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Inform...
2021
-
[32]
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021. Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations
2021
-
[33]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.550 Dense passage retrieval for open-domain question answering . In Proceedings of the 2020 Conference on Empiric...
2020 doi
-
[34]
Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav...
2019 doi
-
[35]
Barak Lenz, Alan Arazi, Amir Bergman, Avshalom Manevich, Barak Peleg, Ben Aviram, Chen Almagor, Clara Fridman, Dan Padnos, et al. 2024. Jamba-1.5: Hybrid transformer-mamba models at scale. arXiv preprint arXiv:2408.12570
2024 arXiv
-
[36]
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. https://doi.org/10.18653/v1/2020.acl-main.703 BART : Denoising sequence-to-sequence pre-training for natural language generation, translatio...
2020 doi
-
[37]
Canjia Li, Andrew Yates, Sean MacAvaney, Ben He, and Yingfei Sun. 2023. Parade: Passage representation aggregation fordocument reranking. ACM Transactions on Information Systems, 42(2):1--26
2023
-
[38]
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2023. Holistic evaluation of language models. Trans. Mach. Learn. Res
2023
-
[39]
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, et al. 2024. Jamba: A hybrid transformer-mamba language model. arXiv preprint arXiv:2403.19887
2024 arXiv
-
[40]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692
2019 arXiv
-
[41]
Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2023. Fine-tuning llama for multi-stage text retrieval. arXiv preprint arXiv:2310.08319
2023 arXiv
-
[42]
Macedo Maia, Siegfried Handschuh, Andr \'e Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018. Www'18 open challenge: financial opinion mining and question answering. In Companion proceedings of the the web conference 2018, pages 1941--1942
2018
-
[43]
Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al. 2022. Text and code embeddings by contrastive pre-training. arXiv preprint arXiv:2201.10005
2022 arXiv
-
[44]
Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith Hall, Daniel Cer, and Yinfei Yang. 2022. https://doi.org/10.18653/v1/2022.findings-acl.146 Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models . In Findings of the Association for Computa...
2022 doi
-
[45]
Rodrigo Nogueira and Kyunghyun Cho. 2019. Passage re-ranking with bert. arXiv preprint arXiv:1901.04085
2019 arXiv
-
[46]
Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020. Document ranking with a pretrained sequence-to-sequence model. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 708--718
2020
-
[47]
Rodrigo Nogueira, Wei Yang, Jimmy Lin, and Kyunghyun Cho. 2019. Document expansion by query prediction. arXiv preprint arXiv:1904.08375
2019 arXiv
-
[48]
Nvidia. 2025. Nemotron-h: A family of accurate, efficient hybrid mamba-transformer models
2025
-
[49]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32
2019
-
[50]
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Leon Derczynski, Xingjian Du, Matteo Grella, Kranthi Gv, Xuzheng He, Haowen Hou, Przemyslaw Kazienko, Jan Kocon, Jiaming Kong, Bart omiej Koptyra, Hay...
2023
-
[51]
Ronak Pradeep, Rodrigo Nogueira, and Jimmy Lin. 2021. The expando-mono-duo design pattern for text ranking with pretrained sequence-to-sequence models. arXiv preprint arXiv:2101.05667
2021 arXiv
-
[52]
Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun, and Yiran Zhong. 2024. Hgrn2: Gated linear rnns with state expansion. arXiv preprint arXiv:2404.07904
2024 arXiv
-
[53]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1--67
2020
-
[54]
Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al. 1995. Okapi at trec-3. Nist Special Publication Sp, 109:109
1995
-
[55]
Hinrich Sch \"u tze, Christopher D Manning, and Prabhakar Raghavan. 2008. Introduction to information retrieval, volume 39. Cambridge University Press Cambridge
2008
-
[56]
Jimmy TH Smith, Andrew Warrington, and Scott Linderman. 2022. Simplified state space layers for sequence modeling. In The Eleventh International Conference on Learning Representations
2022
-
[57]
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063
2024
-
[58]
Nandan Thakur, Nils Reimers, Andreas R \"u ckl \'e , Abhishek Srivastava, and Iryna Gurevych. 2021. Beir: A heterogeneous benchmark for zero-shot evaluation of information retrieval models
2021
-
[59]
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. Fever: a large-scale dataset for fact extraction and verification. arXiv preprint arXiv:1803.05355
2018 arXiv
-
[60]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288
2023 arXiv
-
[61]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30
2017
-
[62]
Ellen Voorhees, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, William R Hersh, Kyle Lo, Kirk Roberts, Ian Soboroff, and Lucy Lu Wang. 2021. Trec-covid: constructing a pandemic information retrieval test collection. In ACM SIGIR Forum, volume 54, pages 1--12. ACM New York, NY, USA
2021
-
[63]
Henning Wachsmuth, Shahbaz Syed, and Benno Stein. 2018. Retrieval of the best counterargument without prior topic knowledge. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 241--251
2018
-
[64]
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or fiction: Verifying scientific claims. arXiv preprint arXiv:2004.14974
2020 arXiv
-
[65]
Roger Waleffe, Wonmin Byeon, Duncan Riach, Brandon Norick, Vijay Korthikanti, Tri Dao, Albert Gu, Ali Hatamizadeh, Sudhakar Singh, Deepak Narayanan, et al. 2024. An empirical study of mamba-based language models. arXiv preprint arXiv:2406.07887
2024 arXiv
-
[66]
Junxiong Wang, Daniele Paliotta, Avner May, Alexander Rush, and Tri Dao. 2024. The mamba in the llama: Distilling and accelerating hybrid models. Advances in Neural Information Processing Systems, 37:62432--62457
2024
-
[67]
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022. Simlm: Pre-training with representation bottleneck for dense passage retrieval. arXiv preprint arXiv:2207.02578
2022 arXiv
-
[68]
T Wolf. 2019. Huggingface's transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771
2019 arXiv
-
[69]
Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighof. 2023. C-pack: Packaged resources to advance general chinese embedding. arXiv preprint arXiv:2309.07597
2023 arXiv
-
[70]
Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, et al. 2024. Effective long-context scaling of foundation models. In Proceedings of the 2024 Conference of the North American ...
2024
-
[71]
Zhichao Xu. 2023. Context-aware decoding reduces hallucination in query-focused summarization. arXiv preprint arXiv:2312.14335
2023
-
[72]
Zhichao Xu. 2024. Rankmamba, benchmarking mamba's document ranking performance in the era of transformers. arXiv preprint arXiv:2403.18276
2024
-
[73]
Zhichao Xu and Daniel Cohen. 2023. A lightweight constrained generation alternative for query-focused summarization. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 1745--1749
2023
-
[74]
Zhichao Xu, Daniel Cohen, Bei Wang, and Vivek Srikumar. 2024 a . https://doi.org/10.18653/v1/2024.findings-naacl.167 In-context example ordering guided by label distributions . In Findings of the Association for Computational Linguistics: NAACL 2024, pages 2623--2640, Mexico C...
2024 doi
-
[75]
Zhichao Xu, Ashim Gupta, Tao Li, Oliver Bentham, and Vivek Srikumar. 2024 b . https://doi.org/10.18653/v1/2024.findings-emnlp.901 Beyond perplexity: Multi-dimensional safety evaluation of LLM compression . In Findings of the Association for Computational Linguistics: EMNLP 202...
2024 doi
-
[76]
Zhichao Xu, Hemank Lamba, Qingyao Ai, Joel Tetreault, and Alex Jaimes. 2024 c . https://doi.org/10.1145/3664190.3672508 Cfe2: Counterfactual editing for search result explanation . In Proceedings of the 2024 ACM SIGIR International Conference on Theory of Information Retrieval...
2024
-
[77]
Zhichao Xu, Fengran Mo, Zhiqi Huang, Crystina Zhang, Puxuan Yu, Bei Wang, Jimmy Lin, and Vivek Srikumar. 2025. A survey of model architectures in information retrieval. arXiv preprint arXiv:2502.14822
2025
-
[78]
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. 2023. Gated linear attention transformers with hardware-efficient training. arXiv preprint arXiv:2312.06635
2023 arXiv
-
[79]
Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, and Yoon Kim. 2024. https://openreview.net/forum?id=y8Rm4VNRPH Parallelizing linear transformers with the delta rule over sequence length . In The Thirty-eighth Annual Conference on Neural Information Processing Systems
2024
-
[80]
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600
2018 arXiv
-
[81]
Andrew Yates, Rodrigo Nogueira, and Jimmy Lin. 2021. Pretrained transformers for text ranking: Bert and beyond. In Proceedings of the 14th ACM International Conference on web search and data mining, pages 1154--1156
2021
-
[82]
Hanqi Zhang, Chong Chen, Lang Mei, Qi Liu, and Jiaxin Mao. 2024. Mamba retriever: Utilizing mamba for effective and efficient dense retrieval. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, pages 4268--4272
2024
-
[83]
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068
2022 arXiv
-
[84]
Yue Zhang, ChengCheng Hu, Yuqi Liu, Hui Fang, and Jimmy Lin. 2021. https://doi.org/10.18653/v1/2021.sustainlp-1.8 Learning to rank in the age of M uppets: Effectiveness -- efficiency tradeoffs in multi-stage ranking . In Proceedings of the Second Workshop on Simple and Efficie...
2021 doi
-
[85]
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. 2024. Vision mamba: Efficient visual representation learning with bidirectional state space model. In Forty-first International Conference on Machine Learning
2024
-
[86]
Honglei Zhuang, Zhen Qin, Rolf Jagerman, Kai Hui, Ji Ma, Jing Lu, Jianmo Ni, Xuanhui Wang, and Michael Bendersky. 2023. Rankt5: Fine-tuning t5 for text ranking with ranking losses. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Inf...
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.