REVIEW 3 major objections 3 minor 185 references
Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers
T0 review · 3 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This paper argues that a single frozen Granularity object plus a GranularitySchedule can express all three compression axes—depth, token count, and embedding width—so one trained checkpoint serves every operating point, with prior elastic…
desk verdict A genuinely useful unified abstraction for elastic retrievers, but the 'single forward pass' training-cost claim does not hold for the token axis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the Granularity dataclass paired with a GranularitySchedule. Granularity is a frozen object with independent optional fields—layer for depth, dim for width, keep_ratio and pool_layer for token compression—and prints as a collision-free key such as L28xR0.6xP20. The readout function routes through the hidden-states tuple, the layer module list, and the pooling step that standard transformer interfaces expose, so depth and width need no model-specific code; the token axis adds one per-family upper-runner shim that rebuilds attention masks and rotary position embeddings for the shortened sequence. During training, one forward pass per tower is shared and the loss is the sum over schedule points, while at deployment a prune step removes upper layers for depth and applies width and token choices at readout time.
What would settle it
Use the framework's own launch-flag interface to train a token-compression schedule on a sliding-window attention backbone such as ModernBERT; the paper states this requires a pooling rule compatible with a fixed local window, so if the interface cannot express it without new model code, the no-new-code claim is false for that family. Separately, a batch-128 GPU timing run that deviates from the reported ~2% agreement between measured and FLOP-predicted speedup would falsify the cost model.
Extended reading notes
Core claim
The central claim is that depth, token, and width compression are not separate methods but readout choices on the same residual stream: choose a layer, optionally shorten the sequence at a pooling layer, optionally truncate the embedding dimension. All three are captured by a small frozen dataclass with optional fields, and training sums per-granularity losses over a declarative schedule from one shared forward pass. The paper reports that across three backbones (BERT, ModernBERT, Qwen3) and two tasks, the resulting quality curves are smooth and monotonic, the elasticity tax over a single-size model is small, and measured wallclock speedups match the analytic FLOP-based cost model within about two percent at large batch size. It also states explicitly that the framework does not improve on prior methods in absolute quality and leaves the full depth×width×token composition expressible but untested.
Load-bearing premise
The load-bearing premise is that the readout abstraction—layer selection, token pooling, dimension truncation—reproduces each prior method's quality and speed, and the paper verifies this only on global-attention backbones, explicitly leaving sliding-window attention (ModernBERT) unsupported for the token axis.
Editorial extensions
If this is right
- A single trained checkpoint serves every operating point listed in the schedule, so deployment can switch between latency, throughput, and index-size targets without retraining.
- Prior elastic methods—MRL, early exit, 2D Matryoshka/Starbucks, and LTC—can be reproduced by configuration strings, and new combinations such as width plus token, or MLTC, cost no new modeling code.
- Rerankers inherit depth and token compression but not width, since a scalar score has no embedding to truncate; retrievers admit all three axes.
- For large batches, inference speedup can be predicted from simple FLOP counts; the measured throughput matched within about two percent for reranking at batch 128 and for document encoding on the depth axis.
- Token compression barely helps short-query latency because eight-token queries have little to pool and lower layers still run; it is a lever for long documents.
Reading between the lines
- A natural extension the paper leaves untested is a cost-budget planner that picks the operating point per input: depth for short queries, token compression for long documents, and width for index budgets; the reported smooth quality curves make such selection low-risk.
- The framework's schedule grammar suggests the full depth×width×token composition could be trained directly, and the paper's cost model would predict its speedups; this is the obvious next experiment.
- If token compression is extended to sliding-window attention via a window-compatible pooling rule, the 'launch flag' claim would cover ModernBERT too; until then, that backbone is limited to depth and width.
- The near-zero elasticity tax for rerankers and the small positive tax for the decoder retriever hint that joint multi-size training may act as a mild regularizer on some backbones, a testable hypothesis.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Tevatron-Elastic proposes a unified abstraction for training elastic retrievers and rerankers. A frozen Granularity dataclass with optional fields layer, dim, keep_ratio, and pool_layer names an operating point, and a GranularitySchedule lists the points to train; the readout function routes through hidden states, layer modules, and pooling interfaces exposed by Hugging Face transformers. The paper shows that MRL, early exit, 2D Matryoshka/Starbucks, and LTC can be expressed as configurations, introduces MLTC for jointly training multiple token-compression ratios, trains 20 checkpoints over BERT, ModernBERT, and Qwen3 (with Llama3/Mistral path verification), reports smooth quality curves, an elasticity tax at the full point, and inference speedups that match a FLOP-based cost model to about two percent at large batch sizes. The authors release code and checkpoints and are explicit that the framework does not improve absolute quality and that the token axis is currently limited to global-attention backbones.
Significance. The paper's main value is engineering: it unifies three compression axes in one clean interface, makes prior methods special cases of a single schedule grammar, and ships reproducible code and checkpoints. The cost model is simple and parameter-free, and the measured wall-clock speedups at large batch sizes confirm it. The claims about training efficiency and backbone coverage are, however, stronger than the evidence: the 'single shared forward pass' account does not hold for token-compression readouts, no training-time cost is reported for token schedules, and the MLTC validation omits a separately trained LTC retriever baseline. If these gaps are closed or the claims scoped accordingly, the abstraction would be a solid systems contribution to the IR literature.
major comments (3)
- [Section 3 (Joint training) and Section 4.2/Table 3] Section 3, 'Joint training' paragraph: the statement 'We run the backbone model for a single forward pass per batch; every operating point in the schedule is then a cheap readout off that shared forward pass' is inaccurate for token granularities. In the _readout implementation, when g.pool_layer is set, the code calls pool_sequence(H[g.pool_layer], ...) and then _range_runner over layers g.pool_layer..g.layer, which is a separate partial forward pass for every token granularity, not a readout. For the MLTC schedule '28@1.0/20,...,28@0.4/20', layers 20 through 28 are re-executed four times, including at r=1.0 where pooling is the identity and H[28] from the shared forward pass could be reused. Section 4.2 and Table 3 measure the elasticity tax only as nDCG@10 differences, and Section 4.3 measures only inference speedups; no training wall-clock time or training FLOPs are reported for a token schedule. The introduction's 'little extra cost' claim therefore remains unsupported for the token axis and is inconsistent with the stated mechanism. Please either report training time/FLOPs for a token-compression schedule against a single-point baseline, or restrict the cost claim to depth and width.
- [Section 4.2 (Token, MLTC Retriever)] The MLTC experiment compares only against the plain Qwen3 retriever (0.511 vs 0.513) and explicitly states that no comparison with separately trained LTC retrievers was made 'due to limited bandwidth'. Without a single-ratio LTC retriever trained under the same recipe, the MLTC curve cannot separate the effect of jointly training multiple ratios from the effect of the token-pooling path itself, so the value of MLTC as a training method is not established. Add at least one LTC retriever baseline at a matched ratio (e.g., L28xR0.8xP20), or explicitly frame MLTC as only an implementation sanity check rather than a validated contribution.
- [Table 2 and Section 5 (Future Work)] Table 2 claims 'no new model code' for five backbone families, but the token axis requires a per-family upper runner and is not available for ModernBERT, whose sliding-window attention is excluded in Section 5. The statement 'a new backbone is a launch flag rather than new code' is therefore true only for the depth and width axes and for token compression on global-attention backbones. Please scope the backbone-agnostic claim accordingly, or provide the ModernBERT token runner; otherwise the central abstraction claim overstates its coverage.
minor comments (3)
- [Throughout] Several places have missing spaces between numbers and words, e.g., '0.431at layer 6' and '0.518at layer 28' in Section 4.2; a copyedit pass would improve readability.
- [Section 2.2, Eq. (3)] The notation fp:L for the upper-layer stack is introduced only in prose; define it directly at the equation or in a following sentence to avoid ambiguity.
- [Table 1] The 'starbucks_bert/diagonal(...)' spec is not explained; a one-sentence example of how the diagonal depth-width composition is written would help readers use the grammar.
Circularity Check
No significant circularity: the abstraction is a design artifact, prior methods are reproduced as configurations with fresh measurements, and the analytic cost model is validated against wall-clock speed rather than fitted from it.
full rationale
Tevatron-Elastic makes no prediction that reduces to its own inputs. Section 3 defines Granularity with fields layer, dim, keep_ratio, and pool_layer, so the claim that MRL, early exit, 2D Matryoshka/Starbucks, and LTC are special cases is an encoding property of the dataclass, explicitly presented as an engineering abstraction ('We do not introduce a new pooling operation; instead, we provide one abstraction'), not a derived scientific result. The analytic cost model in Section 4.3 is a closed-form FLOP count, and the paper checks measured wall-clock speedup against it ('layer 4 measures 6.86x against a predicted 7.00x'), which is an independent empirical test rather than a fitted constant. The quality curves and elasticity tax (Table 3) are new BEIR-15 measurements, not reductions of the target result. Citations to co-authored prior work (LTC, Starbucks, Tevatron) are used as external methods to reproduce or as background, and none carries the argument alone; the evaluation is self-contained against an external benchmark. The one significant missupport is the Section 3 claim that 'every operating point in the schedule is then a cheap readout off that shared forward pass': for token-axis granularities, _readout calls pool_sequence and _range_runner, re-running layers pool_layer..layer once per keep ratio, and no training wall-clock data are reported for MLTC, so the 'little extra cost' promise is unverified for the token axis. This is an evidentiary gap and a mechanistic overclaim, not a self-definitional or self-citation-based reduction; it does not constitute circularity.
Assumptions & free parameters
free parameters (2)
- pool_layer p for token axis =
20 (fixed by experiment, not fitted)
- per-granularity training temperature =
not reported
assumptions (3)
- domain assumption Hugging Face transformer hidden-states tuple and layer module list are sufficient to implement all three compression axes without modifying model internals.
- domain assumption A single forward pass per batch, with multiple readouts, trains all operating points adequately without per-size optimization.
- standard math In-batch InfoNCE with shared query-passage readouts is a valid training objective for every granularity.
invented entities (2)
-
Granularity dataclass
independent evidence
-
MLTC (Matryoshka Layerwise Token Compression)
independent evidence
Cite this review
Pith. "Pith review of Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers." pith.science (2026). https://pith.science/paper/T6TR3JUV
@misc{pith2026260808809,
author = {Pith},
title = {Pith review of: Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers},
year = {2026},
howpublished = {\url{https://pith.science/paper/T6TR3JUV}},
note = {Machine review of arXiv:2608.08809}
}
read the original abstract
A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the right trade-off changes with the workload. In the context of information retrieval (IR), a transformer-based model can be made smaller in three ways---using fewer layers, passing fewer tokens through the upper layers, or producing a shorter embedding---and each way saves a different compute resource. These options have been studied one at a time, each as its own method with its own code and training setup, which makes them hard to combine or adapt to a new model. We present~\ours to bring all three under one simple abstraction: a single object names any size the model can run at, and a short schedule lists the sizes to train. Training then produces one checkpoint that serves all of those sizes, and at deployment the user picks any of them. The same abstraction covers both retrievers and rerankers and both encoder and decoder models, as it works through interfaces that Hugging Face transformers already expose; a new backbone is a configuration change, not new modeling code. Prior methods---Matryoshka embeddings, early exit, 2D~Matryoshka (e.g., Starbucks), and layerwise token compression---become special cases of our unified abstraction. The same interface also enables Matryoshka~LTC (MLTC), which jointly trains several token-compression ratios in one retriever checkpoint. To validate our framework, we train 20 checkpoints across three backbones and two tasks: the quality curves are smooth, one checkpoint costs little over a model trained for a single size, and a controlled study confirms the wallclock speedups. We release the framework and all checkpoints as a resource for building elastic retrieval systems.
Reference graph
Works this paper leans on
-
[28]
arXiv preprint arXiv:2001.08361 , year=
Scaling laws for neural language models , author=. arXiv preprint arXiv:2001.08361 , year=
arXiv 2001
-
[1]
Scalable training of
Andrew, Galen and Gao, Jianfeng , booktitle=. Scalable training of. 2007 , url=
2007
-
[2]
Dan Gusfield , title =. 1997
1997
-
[3]
Tetreault , title =
Mohammad Sadegh Rasooli and Joel R. Tetreault , title =. Computing Research Repository , volume =. 2015 , url =
2015
-
[4]
A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =
Ando, Rie Kubota and Zhang, Tong , Issn =. A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =. Journal of Machine Learning Research , Month = dec, Numpages =. 2005 , url=
2005
-
[5]
and Tukey, John W
Cooley, James W. and Tukey, John W. , journal=. An algorithm for the machine calculation of complex. 1965 , url=
1965
-
[6]
SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking , year =
Formal, Thibault and Piwowarski, Benjamin and Clinchant, St\'. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking , year =. Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =
-
[7]
SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval , publisher =
Formal, Thibault and Lassance, Carlos and Piwowarski, Benjamin and Clinchant, Stéphane , keywords =. SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval , publisher =. 2021 , copyright =. doi:10.48550/ARXIV.2109.10086 , url =
Show all 185 references
-
[9]
An Efficiency Study for SPLADE Models , year =
Lassance, Carlos and Clinchant, St\'. An Efficiency Study for SPLADE Models , year =. Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =. doi:10.1145/3477495.3531833 , abstract =
-
[10]
Transactions on Machine Learning Research , issn=
A Survey of Model Architectures in Information Retrieval , author=. Transactions on Machine Learning Research , issn=. 2026 , url=
2026
-
[11]
2022 , publisher=
Pretrained transformers for text ranking: Bert and beyond , author=. 2022 , publisher=
2022
-
[12]
Dense Passage Retrieval for Open-Domain Question Answering
Karpukhin, Vladimir and Oguz, Barlas and Min, Sewon and Lewis, Patrick and Wu, Ledell and Edunov, Sergey and Chen, Danqi and Yih, Wen-tau. Dense Passage Retrieval for Open-Domain Question Answering. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Pr...
2020 doi
-
[13]
arXiv preprint arXiv:2308.07107 , year=
Large language models for information retrieval: A survey , author=. arXiv preprint arXiv:2308.07107 , year=
-
[14]
BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina. BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Hum...
2019 doi
-
[15]
arXiv preprint arXiv:2307.09288 , year=
Llama 2: Open foundation and fine-tuned chat models , author=. arXiv preprint arXiv:2307.09288 , year=
-
[16]
Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
Fine-tuning llama for multi-stage text retrieval , author=. Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
-
[17]
ArXiv , year=
Mistral 7B , author=. ArXiv , year=
-
[18]
arXiv preprint arXiv:1611.09268 , year=
Ms marco: A human generated machine reading comprehension dataset , author=. arXiv preprint arXiv:1611.09268 , year=
-
[19]
Robertson, Stephen and Zaragoza, Hugo , title =. Found. Trends Inf. Retr. , month = apr, pages =. 2009 , issue_date =. doi:10.1561/1500000019 , abstract =
2009 doi
-
[20]
Nist Special Publication Sp , volume=
Okapi at TREC-3 , author=. Nist Special Publication Sp , volume=. 1995 , publisher=
1995
-
[21]
Journal of documentation , volume=
A statistical interpretation of term specificity and its application in retrieval , author=. Journal of documentation , volume=. 1972 , publisher=
1972
-
[22]
Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
Learning passage impacts for inverted indexes , author=. Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
-
[24]
arXiv preprint arXiv:2010.02666 , year=
Improving efficient neural ranking models with cross-architecture knowledge distillation , author=. arXiv preprint arXiv:2010.02666 , year=
2010 arXiv
-
[25]
International Conference on Learning Representations , year=
Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval , author=. International Conference on Learning Representations , year=
-
[29]
arXiv preprint arXiv:2203.15556 , year=
Training compute-optimal large language models , author=. arXiv preprint arXiv:2203.15556 , year=
-
[30]
Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =
Fang, Yan and Zhan, Jingtao and Ai, Qingyao and Mao, Jiaxin and Su, Weihang and Chen, Jia and Liu, Yiqun , title =. Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =. 2024 , isbn =. doi:10.1145/3626772.365...
2024
-
[31]
Proceedings of the 27th annual international ACM SIGIR conference on Research and development in information retrieval , pages=
A formal study of information retrieval heuristics , author=. Proceedings of the 27th annual international ACM SIGIR conference on Research and development in information retrieval , pages=
-
[32]
arXiv preprint arXiv:1903.06733 , year=
Dying relu and initialization: Theory and numerical examples , author=. arXiv preprint arXiv:1903.06733 , year=
1903 arXiv
-
[33]
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration , url =
Lin, Ji and Tang, Jiaming and Tang, Haotian and Yang, Shang and Chen, Wei-Ming and Wang, Wei-Chen and Xiao, Guangxuan and Dang, Xingyu and Gan, Chuang and Han, Song , booktitle =. AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration , url =
-
[34]
2023 , url=
Elias Frantar and Saleh Ashkboos and Torsten Hoefler and Dan Alistarh , booktitle=. 2023 , url=
2023
-
[35]
int8 (): 8-bit matrix multiplication for transformers at scale , author=
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale , author=. Advances in neural information processing systems , volume=
-
[36]
arXiv preprint arXiv:2402.15449 , year=
Repetition improves language model embeddings , author=. arXiv preprint arXiv:2402.15449 , year=
-
[37]
torchao: PyTorch native quantization and sparsity for training and inference , author =
-
[38]
Communications of the ACM , volume=
A vector space model for automatic indexing , author=. Communications of the ACM , volume=. 1975 , publisher=
1975
-
[39]
Proceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region , pages=
Neural Lexical Search with Learned Sparse Retrieval , author=. Proceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region , pages=
2024
-
[40]
European Conference on Information Retrieval , pages=
A unified framework for learned sparse retrieval , author=. European Conference on Information Retrieval , pages=. 2023 , organization=
2023
-
[41]
Proceedings of the 42nd international ACM SIGIR conference on research and development in information retrieval , pages=
Deeper text understanding for IR with contextual neural language modeling , author=. Proceedings of the 42nd international ACM SIGIR conference on research and development in information retrieval , pages=
-
[42]
arXiv preprint arXiv:2010.00768 , year=
SparTerm: Learning term-based sparse representation for fast text retrieval , author=. arXiv preprint arXiv:2010.00768 , year=
2010 arXiv
-
[43]
Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
SLIM: Sparsified late interaction for multi-vector retrieval with inverted indexes , author=. Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
-
[44]
International Conference on Learning Representations , year=
Minimizing FLOPs to Learn Efficient Sparse Representations , author=. International Conference on Learning Representations , year=
-
[45]
arXiv preprint arXiv:2205.01068 , year=
Opt: Open pre-trained transformer language models , author=. arXiv preprint arXiv:2205.01068 , year=
-
[48]
arXiv preprint arXiv:1907.11692 , year=
Roberta: A robustly optimized bert pretraining approach , author=. arXiv preprint arXiv:1907.11692 , year=
1907 arXiv
-
[49]
Don`t Stop Pretraining: Adapt Language Models to Domains and Tasks
Gururangan, Suchin and Marasovi \'c , Ana and Swayamdipta, Swabha and Lo, Kyle and Beltagy, Iz and Downey, Doug and Smith, Noah A. Don`t Stop Pretraining: Adapt Language Models to Domains and Tasks. Proceedings of the 58th Annual Meeting of the Association for Computational Li...
2020 doi
-
[50]
arXiv preprint arXiv:2405.17428 , year=
Nv-embed: Improved techniques for training llms as generalist embedding models , author=. arXiv preprint arXiv:2405.17428 , year=
-
[51]
Advances in Neural Information Processing Systems , volume=
MosaicBERT: A bidirectional encoder optimized for fast pretraining , author=. Advances in Neural Information Processing Systems , volume=
-
[52]
R ocket QA : An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering
Qu, Yingqi and Ding, Yuchen and Liu, Jing and Liu, Kai and Ren, Ruiyang and Zhao, Wayne Xin and Dong, Daxiang and Wu, Hua and Wang, Haifeng. R ocket QA : An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering. Proceedings of the 2021 Confe...
2021 doi
-
[53]
2024 , url=
Parishad BehnamGhader and Vaibhav Adlakha and Marius Mosbach and Dzmitry Bahdanau and Nicolas Chapados and Siva Reddy , booktitle=. 2024 , url=
2024
-
[54]
Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =
Gao, Luyu and Ma, Xueguang and Lin, Jimmy and Callan, Jamie , title =. Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =. 2023 , isbn =. doi:10.1145/3539618.3591805 , abstract =
2023
-
[55]
arXiv preprint arXiv:2102.10073 , year=
Pyserini: An easy-to-use python toolkit to support replicable ir research with sparse and dense representations , author=. arXiv preprint arXiv:2102.10073 , year=
-
[57]
Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2) , year=
Nandan Thakur and Nils Reimers and Andreas R. Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2) , year=
-
[58]
P rompt R eps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval
Zhuang, Shengyao and Ma, Xueguang and Koopman, Bevan and Lin, Jimmy and Zuccon, Guido. P rompt R eps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval. Proceedings of the 2024 Conference on Empirical Methods in Natur...
2024 doi
-
[59]
S im LM : Pre-training with Representation Bottleneck for Dense Passage Retrieval
Wang, Liang and Yang, Nan and Huang, Xiaolong and Jiao, Binxing and Yang, Linjun and Jiang, Daxin and Majumder, Rangan and Wei, Furu. S im LM : Pre-training with Representation Bottleneck for Dense Passage Retrieval. Proceedings of the 61st Annual Meeting of the Association fo...
2023 doi
-
[60]
Large Dual Encoders Are Generalizable Retrievers
Ni, Jianmo and Qu, Chen and Lu, Jing and Dai, Zhuyun and Hernandez Abrego, Gustavo and Ma, Ji and Zhao, Vincent and Luan, Yi and Hall, Keith and Chang, Ming-Wei and Yang, Yinfei. Large Dual Encoders Are Generalizable Retrievers. Proceedings of the 2022 Conference on Empirical ...
2022 doi
-
[61]
arXiv 2021 , author=
Lora: Low-rank adaptation of large language models. arXiv 2021 , author=. arXiv preprint arXiv:2106.09685 , year=
2021 arXiv
-
[62]
Transactions on Machine Learning Research , year=
Lora learns less and forgets less , author=. Transactions on Machine Learning Research , year=
-
[63]
arXiv preprint arXiv:2307.08691 , year=
Flashattention-2: Faster attention with better parallelism and work partitioning , author=. arXiv preprint arXiv:2307.08691 , year=
-
[64]
arXiv preprint arXiv:2304.11277 , year=
Pytorch FSDP: experiences on scaling fully sharded data parallel , author=. arXiv preprint arXiv:2304.11277 , year=
-
[65]
Journal of machine learning research , volume=
Exploring the limits of transfer learning with a unified text-to-text transformer , author=. Journal of machine learning research , volume=
-
[66]
arXiv preprint arXiv:1910.03771 , year=
Huggingface's transformers: State-of-the-art natural language processing , author=. arXiv preprint arXiv:1910.03771 , year=
1910 arXiv
-
[67]
arXiv preprint arXiv:1606.08415 , year=
Gaussian error linear units (gelus) , author=. arXiv preprint arXiv:1606.08415 , year=
-
[68]
The Twelfth International Conference on Learning Representations , year=
Non-negative Contrastive Learning , author=. The Twelfth International Conference on Learning Representations , year=
-
[69]
arXiv preprint arXiv:1611.01144 , year=
Categorical reparameterization with gumbel-softmax , author=. arXiv preprint arXiv:1611.01144 , year=
-
[70]
Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval , pages=
Expansion via prediction of importance with contextualization , author=. Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval , pages=
-
[71]
Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
Splate: Sparse late interaction retrieval , author=. Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
-
[72]
Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression
Xu, Zhichao and Gupta, Ashim and Li, Tao and Bentham, Oliver and Srikumar, Vivek. Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression. Findings of the Association for Computational Linguistics: EMNLP 2024. 2024. doi:10.18653/v1/2024.findings-emnlp.901
2024 doi
-
[73]
International Conference on Machine Learning , pages=
The case for 4-bit precision: k-bit inference scaling laws , author=. International Conference on Machine Learning , pages=. 2023 , organization=
2023
-
[74]
International Conference on Machine Learning , pages=
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression , author=. International Conference on Machine Learning , pages=. 2024 , organization=
2024
-
[75]
arXiv preprint arXiv:2012.00614 , year=
Climate-fever: A dataset for verification of real-world climate claims , author=. arXiv preprint arXiv:2012.00614 , year=
2012 arXiv
-
[76]
ACM SIGIR Forum , volume=
TREC-COVID: constructing a pandemic information retrieval test collection , author=. ACM SIGIR Forum , volume=. 2021 , organization=
2021
-
[77]
arXiv preprint arXiv:1809.09600 , year=
HotpotQA: A dataset for diverse, explainable multi-hop question answering , author=. arXiv preprint arXiv:1809.09600 , year=
-
[78]
and Uszkoreit, Jakob and Le, Quoc and Petrov, Slav
Kwiatkowski, Tom and Palomaki, Jennimaria and Redfield, Olivia and Collins, Michael and Parikh, Ankur and Alberti, Chris and Epstein, Danielle and Polosukhin, Illia and Devlin, Jacob and Lee, Kenton and Toutanova, Kristina and Jones, Llion and Kelcey, Matthew and Chang, Ming-W...
2019 doi
-
[79]
Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Retrieval of the best counterargument without prior topic knowledge , author=. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[80]
Overview of Touch
Bondarenko, Alexander and Fr. Overview of Touch. Experimental IR Meets Multilinguality, Multimodality, and Interaction: 11th International Conference of the CLEF Association, CLEF 2020, Thessaloniki, Greece, September 22--25, 2020, Proceedings 11 , pages=. 2020 , organization=
2020
-
[81]
Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
DBpedia-entity v2: a test collection for entity search , author=. Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
-
[82]
arXiv preprint arXiv:2004.07180 , year=
Specter: Document-level representation learning using citation-informed transformers , author=. arXiv preprint arXiv:2004.07180 , year=
2004 arXiv
-
[83]
arXiv preprint arXiv:1803.05355 , year=
FEVER: a large-scale dataset for fact extraction and VERification , author=. arXiv preprint arXiv:1803.05355 , year=
-
[84]
arXiv preprint arXiv:2004.14974 , year=
Fact or fiction: Verifying scientific claims , author=. arXiv preprint arXiv:2004.14974 , year=
2004 arXiv
-
[85]
Advances in Information Retrieval: 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20--23, 2016
A full-text learning to rank dataset for medical information retrieval , author=. Advances in Information Retrieval: 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20--23, 2016. Proceedings 38 , pages=. 2016 , organization=
2016
-
[86]
Companion proceedings of the the web conference 2018 , pages=
Www'18 open challenge: financial opinion mining and question answering , author=. Companion proceedings of the the web conference 2018 , pages=
2018
-
[87]
arXiv preprint arXiv:2202.08904 , year=
Sgpt: Gpt sentence embeddings for semantic search , author=. arXiv preprint arXiv:2202.08904 , year=
-
[88]
arXiv preprint arXiv:2212.03533 , year=
Text embeddings by weakly-supervised contrastive pre-training , author=. arXiv preprint arXiv:2212.03533 , year=
-
[89]
Improving Text Embeddings with Large Language Models
Wang, Liang and Yang, Nan and Huang, Xiaolong and Yang, Linjun and Majumder, Rangan and Wei, Furu. Improving Text Embeddings with Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. doi:1...
2024 doi
-
[90]
Proceedings of the 1984 ACM SIGMOD International Conference on Management of Data , pages =
Salton, Gerard , title =. Proceedings of the 1984 ACM SIGMOD International Conference on Management of Data , pages =. 1984 , isbn =. doi:10.1145/602259.602295 , abstract =
1984
-
[91]
Proceedings of the 2024 ACM SIGIR International Conference on Theory of Information Retrieval , pages=
CFE2: Counterfactual Editing for Search Result Explanation , author=. Proceedings of the 2024 ACM SIGIR International Conference on Theory of Information Retrieval , pages=
2024
-
[92]
Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
A lightweight constrained generation alternative for query-focused summarization , author=. Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
-
[93]
Journal of Machine Learning Research , volume=
Scaling instruction-finetuned language models , author=. Journal of Machine Learning Research , volume=
-
[94]
arXiv preprint arXiv:2407.10759 , year=
Qwen2-audio technical report , author=. arXiv preprint arXiv:2407.10759 , year=
-
[95]
arXiv preprint arXiv:2506.05176 , year=
Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models , author=. arXiv preprint arXiv:2506.05176 , year=
-
[96]
Distillation versus Contrastive Learning: How to Train Your Rerankers
Xu, Zhichao and Huang, Zhiqi and Zhuang, Shengyao and Srikumar, Vivek. Distillation versus Contrastive Learning: How to Train Your Rerankers. Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapte...
2025 doi
-
[97]
2015 , eprint=
Distilling the Knowledge in a Neural Network , author=. 2015 , eprint=
2015
-
[98]
2025 , eprint=
Towards Competitive Search Relevance For Inference-Free Learned Sparse Retrievers , author=. 2025 , eprint=
2025
-
[99]
2023 , eprint=
Context-aware Decoding Reduces Hallucination in Query-focused Summarization , author=. 2023 , eprint=
2023
-
[100]
2024 , eprint=
RankMamba: Benchmarking Mamba's Document Ranking Performance in the Era of Transformers , author=. 2024 , eprint=
2024
-
[101]
Latent Retrieval for Weakly Supervised Open Domain Question Answering
Lee, Kenton and Chang, Ming-Wei and Toutanova, Kristina. Latent Retrieval for Weakly Supervised Open Domain Question Answering. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. doi:10.18653/v1/P19-1612
2019 doi
-
[102]
CSPLADE : Learned Sparse Retrieval with Causal Language Models
Xu, Zhichao and Feng, Aosong and Tian, Yijun and Ding, Haibo and Cheong, Lin Lee. CSPLADE : Learned Sparse Retrieval with Causal Language Models. Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Ch...
2025 doi
-
[103]
2024 , eprint=
Mistral-SPLADE: LLMs for better Learned Sparse Retrieval , author=. 2024 , eprint=
2024
-
[104]
Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
Scaling sparse and dense retrieval in decoder-only llms , author=. Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
-
[105]
2025 , eprint=
Nomic Embed: Training a Reproducible Long Context Text Embedder , author=. 2025 , eprint=
2025
-
[106]
2024 , eprint=
Arctic-Embed 2.0: Multilingual Retrieval Without Compromise , author=. 2024 , eprint=
2024
-
[107]
2024 , eprint=
Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents , author=. 2024 , eprint=
2024
-
[108]
2024 , eprint=
Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models , author=. 2024 , eprint=
2024
-
[109]
Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup
Gao, Luyu and Zhang, Yunyi and Han, Jiawei and Callan, Jamie. Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup. Proceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2021). 2021. doi:10.18653/v1/2021.repl4nlp-1.31
2021 doi
-
[111]
The Thirteenth International Conference on Learning Representations , year=
Making Text Embedders Few-Shot Learners , author=. The Thirteenth International Conference on Learning Representations , year=
-
[112]
Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
Efficient inverted indexes for approximate retrieval over learned sparse representations , author=. Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
-
[113]
2025 , eprint=
Efficient Sketching and Nearest Neighbor Search Algorithms for Sparse Vector Sets , author=. 2025 , eprint=
2025
-
[114]
2017 , eprint=
Billion-scale similarity search with GPUs , author=. 2017 , eprint=
2017
-
[115]
MTEB : Massive Text Embedding Benchmark
Muennighoff, Niklas and Tazi, Nouamane and Magne, Loic and Reimers, Nils. MTEB : Massive Text Embedding Benchmark. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. doi:10.18653/v1/2023.eacl-main.148
2023 doi
-
[116]
Transactions on Machine Learning Research , issn=
Unsupervised Dense Information Retrieval with Contrastive Learning , author=. Transactions on Machine Learning Research , issn=. 2022 , url=
2022
-
[117]
2024 , eprint=
Text Embeddings by Weakly-Supervised Contrastive Pre-training , author=. 2024 , eprint=
2024
-
[118]
2020 , eprint=
SparTerm: Learning Term-based Sparse Representation for Fast Text Retrieval , author=. 2020 , eprint=
2020
-
[119]
SPARTA : Efficient Open-Domain Question Answering via Sparse Transformer Matching Retrieval
Zhao, Tiancheng and Lu, Xiaopeng and Lee, Kyusong. SPARTA : Efficient Open-Domain Question Answering via Sparse Transformer Matching Retrieval. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tec...
2021 doi
-
[120]
Proceedings of the 27th ACM international conference on information and knowledge management , pages=
From neural re-ranking to neural ranking: Learning a sparse representation for inverted indexing , author=. Proceedings of the 27th ACM international conference on information and knowledge management , pages=
-
[121]
2024 , eprint=
Multimodal Learned Sparse Retrieval for Image Suggestion , author=. 2024 , eprint=
2024
-
[122]
European Conference on Information Retrieval , pages=
Multimodal learned sparse retrieval with probabilistic expansion control , author=. European Conference on Information Retrieval , pages=. 2024 , organization=
2024
-
[123]
International Workshop on Knowledge-Enhanced Information Retrieval , pages=
Leveraging decoder architectures for learned sparse retrieval , author=. International Workshop on Knowledge-Enhanced Information Retrieval , pages=. 2025 , organization=
2025
-
[124]
Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
Effective inference-free retrieval for learned sparse representations , author=. Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=
-
[125]
2025 , eprint=
Milco: Learned Sparse Retrieval Across Languages via a Multilingual Connector , author=. 2025 , eprint=
2025
-
[126]
2025 , eprint=
Luxical: High-Speed Lexical-Dense Text Embeddings , author=. 2025 , eprint=
2025
-
[127]
Sentence- BERT : Sentence Embeddings using S iamese BERT -Networks
Reimers, Nils and Gurevych, Iryna. Sentence- BERT : Sentence Embeddings using S iamese BERT -Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP...
2019 doi
-
[128]
International Conference on Learning Representations , year=
Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring , author=. International Conference on Learning Representations , year=
-
[130]
Proceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval , pages =
Mackenzie, Joel and Zhuang, Shengyao and Zuccon, Guido , title =. Proceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval , pages =. 2023 , isbn =. doi:10.1145/3578337.3605129 , abstract =
2023
-
[132]
2025 , eprint=
Training Sparse Mixture Of Experts Text Embedding Models , author=. 2025 , eprint=
2025
-
[133]
International Conference on Learning Representations , volume=
Generative representational instruction tuning , author=. International Conference on Learning Representations , volume=
-
[134]
2026 , eprint=
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning , author=. 2026 , eprint=
2026
-
[135]
Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , articleno =
Rajbhandari, Samyam and Rasley, Jeff and Ruwase, Olatunji and He, Yuxiong , title =. Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , articleno =. 2020 , isbn =
2020
-
[136]
2023 , eprint=
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel , author=. 2023 , eprint=
2023
-
[137]
2025 , eprint=
Qwen3 Technical Report , author=. 2025 , eprint=
2025
-
[138]
Proceedings of the 2019 ACM SIGIR International Conference on Theory of Information Retrieval , pages =
Bruch, Sebastian and Wang, Xuanhui and Bendersky, Michael and Najork, Marc , title =. Proceedings of the 2019 ACM SIGIR International Conference on Theory of Information Retrieval , pages =. 2019 , isbn =. doi:10.1145/3341981.3344221 , abstract =
2019
-
[139]
Proceedings of the 29th Symposium on Operating Systems Principles , pages =
Kwon, Woosuk and Li, Zhuohan and Zhuang, Siyuan and Sheng, Ying and Zheng, Lianmin and Yu, Cody Hao and Gonzalez, Joseph and Zhang, Hao and Stoica, Ion , title =. Proceedings of the 29th Symposium on Operating Systems Principles , pages =. 2023 , isbn =. doi:10.1145/3600006.36...
2023
-
[140]
2020 , eprint=
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism , author=. 2020 , eprint=
2020
-
[141]
2021 , eprint=
Improving Efficient Neural Ranking Models with Cross-Architecture Knowledge Distillation , author=. 2021 , eprint=
2021
-
[142]
Proceedings of the 22nd international conference on Machine learning , pages=
Learning to rank using gradient descent , author=. Proceedings of the 22nd international conference on Machine learning , pages=
-
[143]
Proceedings of the 24th international conference on Machine learning , pages=
Learning to rank: from pairwise approach to listwise approach , author=. Proceedings of the 24th international conference on Machine learning , pages=
-
[144]
2025 , url=
Hongjin SU and Howard Yen and Mengzhou Xia and Weijia Shi and Niklas Muennighoff and Han-yu Wang and Liu Haisu and Quan Shi and Zachary S Siegel and Michael Tang and Ruoxi Sun and Jinsung Yoon and Sercan O Arik and Danqi Chen and Tao Yu , booktitle=. 2025 , url=
2025
-
[145]
Gonzalez and Clark Barrett and Ying Sheng , booktitle=
Lianmin Zheng and Liangsheng Yin and Zhiqiang Xie and Chuyue Sun and Jeff Huang and Cody Hao Yu and Shiyi Cao and Christos Kozyrakis and Ion Stoica and Joseph E. Gonzalez and Clark Barrett and Ying Sheng , booktitle=. 2024 , url=
2024
-
[146]
Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =
Zhuang, Honglei and Qin, Zhen and Jagerman, Rolf and Hui, Kai and Ma, Ji and Lu, Jing and Ni, Jianmo and Wang, Xuanhui and Bendersky, Michael , title =. Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages =. 2...
2023
-
[148]
Advances in Neural Information Processing Systems , volume=
Matryoshka representation learning , author=. Advances in Neural Information Processing Systems , volume=
-
[151]
2020 , editor =
Goyal, Saurabh and Choudhury, Anamitra Roy and Raje, Saurabh and Chakaravarthy, Venkatesan and Sabharwal, Yogish and Verma, Ashish , booktitle =. 2020 , editor =
2020
-
[152]
2025 , eprint=
Jasper-Token-Compression-600M Technical Report , author=. 2025 , eprint=
2025
-
[153]
Proceedings of the 34th International Conference on Neural Information Processing Systems , articleno =
Zhou, Wangchunshu and Xu, Canwen and Ge, Tao and McAuley, Julian and Xu, Ke and Wei, Furu , title =. Proceedings of the 34th International Conference on Neural Information Processing Systems , articleno =. 2020 , isbn =
2020
-
[154]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. doi:10.18653/v1/2024.acl-long.681
2024 doi
-
[155]
2024 , eprint=
2D Matryoshka Sentence Embeddings , author=. 2024 , eprint=
2024
-
[157]
2021 , journal=
A Mathematical Framework for Transformer Circuits , author=. 2021 , journal=
2021
-
[160]
2026 , eprint=
Tevatron Meets Megatron: Expert-Parallel LLM Reranker Training on an Academic Budget , author=. 2026 , eprint=
2026
-
[161]
Transactions on Machine Learning Research , issn=
Rethinking On-policy Optimization for Query Augmentation , author=. Transactions on Machine Learning Research , issn=. 2026 , url=
2026
-
[162]
2026 , eprint=
RECON: Reasoning with Condensation for Efficient Retrieval-Augmented Generation , author=. 2026 , eprint=
2026
-
[163]
Transactions on Machine Learning Research , issn=
Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation , author=. Transactions on Machine Learning Research , issn=. 2026 , url=
2026
-
[164]
Learning to rank: from pairwise approach to listwise approach
Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li. Learning to rank: from pairwise approach to listwise approach. In Proceedings of the 24th international conference on Machine learning, pages 129--136, 2007
2007
-
[165]
BERT : Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North A merican Chapter ...
2019
-
[166]
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dari...
2021
-
[167]
Layerskip: Enabling early exit inference and self-speculative decoding, August 2024
Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, Bram Wasti, Liangzhen Lai, Anas Mahmoud, Bilge Acun, Saurabh Agarwal, Ahmed Roman, Ahmed A Aly, Beidi Chen, and Carole-Jean Wu. Layerskip: Enabling early exit inference and self-speculative decoding, August 2...
2024
-
[168]
Splade v2: Sparse lexical and expansion model for information retrieval, 2021 a
Thibault Formal, Carlos Lassance, Benjamin Piwowarski, and Stéphane Clinchant. Splade v2: Sparse lexical and expansion model for information retrieval, 2021 a . URL https://arxiv.org/abs/2109.10086
2021 arXiv
-
[169]
SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking, page 2288–2292
Thibault Formal, Benjamin Piwowarski, and St\' e phane Clinchant. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking, page 2288–2292. Association for Computing Machinery, New York, NY, USA, 2021 b . ISBN 9781450380379. URL https://doi.org/10.1145/3404835.3463098
2021
-
[170]
Unsupervised corpus aware language model pre-training for dense passage retrieval
Luyu Gao and Jamie Callan. Unsupervised corpus aware language model pre-training for dense passage retrieval. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1...
2022 doi
-
[171]
P o WER - BERT : Accelerating BERT inference via progressive word-vector elimination
Saurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan Chakaravarthy, Yogish Sabharwal, and Ashish Verma. P o WER - BERT : Accelerating BERT inference via progressive word-vector elimination. In Hal Daumé III and Aarti Singh, editors, Proceedings of the 37th Internati...
2020
-
[172]
The llama 3 herd of models
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024
2024 arXiv
-
[173]
Distilling the knowledge in a neural network, 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network, 2015. URL https://arxiv.org/abs/1503.02531
2015 arXiv
-
[174]
Mistral 7b
Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L'elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, T...
-
[175]
Billion-scale similarity search with gpus, 2017
Jeff Johnson, Matthijs Douze, and Hervé Jégou. Billion-scale similarity search with gpus, 2017. URL https://arxiv.org/abs/1702.08734
2017 arXiv
-
[176]
Matryoshka representation learning
Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, et al. Matryoshka representation learning. Advances in Neural Information Processing Systems, 35: 0 30233--30249, 2022
2022
-
[177]
Splade-v3: New baselines for splade
Carlos Lassance, Herv \'e D \'e jean, Thibault Formal, and St \'e phane Clinchant. Splade-v3: New baselines for splade. arXiv preprint arXiv:2403.06789, 2024
2024 arXiv
-
[178]
2d matryoshka sentence embeddings, 2024
Xianming Li, Zongxi Li, Jing Li, Haoran Xie, and Qing Li. 2d matryoshka sentence embeddings, 2024. URL https://arxiv.org/abs/2402.14776
2024 arXiv
-
[179]
Pretrained transformers for text ranking: Bert and beyond
Jimmy Lin, Rodrigo Nogueira, and Andrew Yates. Pretrained transformers for text ranking: Bert and beyond. Springer Nature, 2022
2022
-
[180]
F ast BERT : a self-distilling BERT with adaptive inference time
Weijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao, Haotang Deng, and Qi Ju. F ast BERT : a self-distilling BERT with adaptive inference time. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Co...
2020 doi
-
[181]
Tevatron 2.0: Unified document retrieval toolkit across scale, language, and modality
Xueguang Ma, Luyu Gao, Shengyao Zhuang, Jiaqi Samantha Zhan, Jamie Callan, and Jimmy Lin. Tevatron 2.0: Unified document retrieval toolkit across scale, language, and modality. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Informa...
2025
-
[182]
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018
2018 arXiv
-
[183]
R ocket QA : An optimized training approach to dense passage retrieval for open-domain question answering
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. R ocket QA : An optimized training approach to dense passage retrieval for open-domain question answering. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, D...
2021
-
[184]
BEIR : A heterogeneous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas R \"u ckl \'e , Abhishek Srivastava, and Iryna Gurevych. BEIR : A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks ...
2021
-
[185]
Hard negatives, hard lessons: Revisiting training data quality for robust information retrieval with LLM s
Nandan Thakur, Crystina Zhang, Xueguang Ma, and Jimmy Lin. Hard negatives, hard lessons: Revisiting training data quality for robust information retrieval with LLM s. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors, Findings of the As...
2025 doi
-
[186]
Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference
Benjamin Warner, Antoine Chaffin, Benjamin Clavi \'e , Orion Weller, Oskar Hallstr \"o m, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, et al. Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long con...
2024 arXiv
-
[187]
Listwise approach to learning to rank: theory and algorithm
Fen Xia, Tie-Yan Liu, Jue Wang, Wensheng Zhang, and Hang Li. Listwise approach to learning to rank: theory and algorithm. In Proceedings of the 25th International Conference on Machine Learning, ICML '08, page 1192–1199, New York, NY, USA, 2008. Association for Computing Machi...
2008
-
[188]
D ee BERT : Dynamic early exiting for accelerating BERT inference
Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. D ee BERT : Dynamic early exiting for accelerating BERT inference. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computatio...
2020 doi
-
[189]
Bennett, Junaid Ahmed, and Arnold Overwijk
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. Approximate nearest neighbor negative contrastive learning for dense text retrieval. In International Conference on Learning Representations, 2021. URL https://open...
2021
-
[190]
CSPLADE : Learned sparse retrieval with causal language models
Zhichao Xu, Aosong Feng, Yijun Tian, Haibo Ding, and Lin Lee Cheong. CSPLADE : Learned sparse retrieval with causal language models. In Kentaro Inui, Sakriani Sakti, Haofen Wang, Derek F. Wong, Pushpak Bhattacharyya, Biplab Banerjee, Asif Ekbal, Tanmoy Chakraborty, and Dhirend...
2025
-
[191]
Distillation versus contrastive learning: How to train your rerankers
Zhichao Xu, Zhiqi Huang, Shengyao Zhuang, and Vivek Srikumar. Distillation versus contrastive learning: How to train your rerankers. In Kentaro Inui, Sakriani Sakti, Haofen Wang, Derek F. Wong, Pushpak Bhattacharyya, Biplab Banerjee, Asif Ekbal, Tanmoy Chakraborty, and Dhirend...
2025
-
[192]
State space models are strong text rerankers
Zhichao Xu, Jinghua Yan, Ashim Gupta, and Vivek Srikumar. State space models are strong text rerankers. In Vaibhav Adlakha, Alexandra Chronopoulou, Xiang Lorraine Li, Bodhisattwa Prasad Majumder, Freda Shi, and Giorgos Vernikos, editors, Proceedings of the 10th Workshop on Rep...
2025 doi
-
[193]
Tevatron meets megatron: Expert-parallel llm reranker training on an academic budget, 2026 a
Zhichao Xu, Xueguang Ma, Shengyao Zhuang, Luyu Gao, Wenqian Ye, Yu Wang, Jamie Callan, and Jimmy Lin. Tevatron meets megatron: Expert-parallel llm reranker training on an academic budget, 2026 a . URL https://arxiv.org/abs/2608.00916
2026 arXiv
-
[194]
A survey of model architectures in information retrieval
Zhichao Xu, Fengran Mo, Zhiqi Huang, Crystina Zhang, Puxuan Yu, Bei Wang Phillips, Jimmy Lin, and Vivek Srikumar. A survey of model architectures in information retrieval. Transactions on Machine Learning Research, 2026 b . ISSN 2835-8856. URL https://openreview.net/forum?id=x...
2026
-
[195]
Recon: Reasoning with condensation for efficient retrieval-augmented generation, 2026 c
Zhichao Xu, Minheng Wang, Yawei Wang, Wenqian Ye, Yuntao Du, Yunpu Ma, and Yijun Tian. Recon: Reasoning with condensation for efficient retrieval-augmented generation, 2026 c . URL https://arxiv.org/abs/2510.10448
2026 arXiv
-
[196]
Beyond correctness: Rewarding faithful reasoning in retrieval-augmented generation
Zhichao Xu, Zongyu Wu, Yun Zhou, Aosong Feng, Kang Zhou, Sangmin Woo, Kiran Ramnath, Yijun Tian, Xuan Qi, Weikang Qiu, Lin Lee Cheong, and Haibo Ding. Beyond correctness: Rewarding faithful reasoning in retrieval-augmented generation. Transactions on Machine Learning Research,...
2026
-
[197]
Rethinking on-policy optimization for query augmentation
Zhichao Xu, Shengyao Zhuang, Xueguang Ma, Bingsen Chen, Yijun Tian, Fengran Mo, Tao Li, Jie Cao, and Vivek Srikumar. Rethinking on-policy optimization for query augmentation. Transactions on Machine Learning Research, 2026 e . ISSN 2835-8856. URL https://openreview.net/forum?i...
2026
-
[199]
Qwen3 technical report, 2025
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jia...
2025 arXiv
-
[200]
Jasper-token-compression-600m technical report, 2025
Dun Zhang, Ziyang Zeng, Yudong Zhou, and Shuyang Lu. Jasper-token-compression-600m technical report, 2025. URL https://arxiv.org/abs/2511.14405
2025
-
[201]
Starbucks: Improved training for 2d matryoshka embeddings
Shengyao Zhuang, Shuai Wang, Fabio Zheng, Bevan Koopman, and Guido Zuccon. Starbucks: Improved training for 2d matryoshka embeddings. In Advances in Information Retrieval: 48th European Conference on Information Retrieval, ECIR 2026, Delft, The Netherlands, March 29 – April 2,...
2026 doi
-
[202]
Layer-wise token compression for efficient document reranking
Shengyao Zhuang, Zhichao Xu, and Ivano Lauriola. Layer-wise token compression for efficient document reranking. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '26, page 4426–4432, New York, NY, USA, 202...
2026
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.