REVIEW 3 major objections 5 minor 1 cited by
Rethinking the Privacy of Text Embeddings: A Reproducibility Study of "Text Embeddings Reveal (Almost) As Much As Text"
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This reproducibility study confirms that text embeddings can be inverted back to the original text—including passwords—and that 8-bit quantization is a simple defense that keeps retrieval quality.
desk verdict Useful reproduction of Vec2Text with real extensions, but the abstract's 'successfully replicate' is stronger than the paper's own tables support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is Vec2Text's iterative correction loop. At each step an inversion encoder-decoder (initialized from T5-base) receives a concatenated input of three MLP-projected vectors—the target embedding, the embedding of the text generated so far, and the difference between them $e - \hat{e}^{(t)}$—together with the tokens from the previous step; the decoder emits a new text whose embedding is closer to the target. The correction term is the identity doing the work: it tells the generator which direction in embedding space its current reconstruction is wrong, so repeated steps and sequence-level beam search drive the output toward the original text. The defense experiments target exactly this signal: Gaussian noise and 8-bit quantization perturb the fine-grained embedding details the correction loop depends on, while the coarse geometry used for retrieval remains intact.
What would settle it
Publish the missing 32-token model checkpoint and the exact evaluation splits, then rerun the 32-token condition of the API-based encoder: if the reproduced BLEU stays near 43 rather than the original roughly 83, the exact-replication claim for that condition is disproved; reaching about 83 would confirm it.
Extended reading notes
Core claim
On the terms of the original paper, the core discovery of this study is confirmation with scope: Vec2Text works as advertised when the released checkpoints are evaluated at their trained text lengths, and its reach extends beyond fluent sentences to secret-like strings. In-domain, gtr-nq-32 matches or slightly exceeds the reported numbers (e.g., 50-step beam search BLEU 98.5 vs 97.3; exact match 94.0 vs 92.0), and ada-ms-128 with an 81-token corpus reproduces the reported scores closely; the apparent large drop in the OpenAI-32 condition (BLEU ~43 vs ~83) is explained by the unavailable ada-ms-32 model being replaced by ada-ms-128. Out-of-domain, several BEIR datasets reproduce within 10%, while others fall short, which the paper attributes to unreleased model versions and sample sizes between 90 and 200. The extensions show that the same mechanism reconstructs password-like inputs at non-trivial rates and that both absolute-max and zeropoint 8-bit quantization reduce BLEU from roughly 36–63 to 16–27 on five BEIR datasets while nDCG@10 stays within about 0.005. The paper also establishes that inversion quality degrades sharply when input length leaves the model's training range, and that for a fixed time budget, adding iterative steps buys more BLEU than widening the beam.
Load-bearing premise
The comparison assumes the released checkpoints and evaluation protocol are the same models that produced the original reported numbers; the paper shows this is false for the 32-token condition of the API-based encoder, where the unavailable 32-token model is replaced by the 128-token one, and sample sizes and model versions also differ in the out-of-domain tests.
Editorial extensions
If this is right
- If the reproduced attack is as strong as reported, any system that stores embeddings instead of raw text should treat those embeddings as sensitive data, especially for user-generated content.
- Because both inversion models degrade on out-of-length texts, the practical risk is highest for pipelines whose text length matches the attacker's assumed training length; deliberately mismatched length distributions could reduce exposure.
- The Gaussian-noise defense works but only with tuned noise levels, whereas 8-bit quantization cuts reconstruction substantially with no hyperparameter search and almost no retrieval loss, making it the more deployable mitigation.
- Password reconstruction at 36% Easy, 22% Medium, and 4% Hard means even semantically opaque secrets are partially recoverable from embeddings, so services handling login data should not assume that secrecy or randomness protects embedded content.
- The Pareto front gives operators a concrete rule: for a fixed compute budget, increase iterative steps before increasing beam width when reconstruction quality is the goal, or invert that ordering when defending against the attack.
Reading between the lines
- The paper's own evidence implies the strong 'successfully replicate' claim should be read as scoped to the released checkpoints: the missing 32-token model means the most impressive OpenAI-32 number from the original paper is not independently verified here.
- An adaptive adversary who knows the quantization scheme could train a quantization-aware inversion model or denoise the quantized embeddings; the paper names this risk but does not test it, so 8-bit quantization is better treated as risk reduction than as a guaranteed defense.
- The length-sensitivity result suggests a testable defense: deploy retrieval with an embedding protocol that hides or shifts the true token-length distribution, forcing an attacker to guess the training length, which the paper shows can sharply lower BLEU.
- The same machinery could be pointed at recommendation-system user embeddings; the paper flags reconstructing user behavior history as future work, and the password results suggest non-linguistic embedding targets are plausible attack surfaces.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a reproducibility study of Vec2Text (Morris et al., EMNLP 2023). Using the released inversion checkpoints gtr-nq-32 and ada-ms-128 together with the original codebase, the authors compare in-domain reconstruction results on NQ and MSMARCO and out-of-domain results on BEIR datasets against the numbers in the original paper. They additionally run three extensions: a hyperparameter sensitivity analysis (iterations and beam size) with a Pareto front, a password-reconstruction experiment, and an 8-bit embedding quantization defense (Absolute Maximum and Zeropoint). The abstract and conclusion claim successful replication with only minor discrepancies; the reported tables show large gaps in several conditions, most notably the OpenAI-MSMARCO 32-token setting and multiple out-of-domain datasets.
Significance. If the hedged conclusions are treated as the contribution, this is a useful independent evaluation: it shows that the released ada-ms-128 checkpoint does not reproduce the 32-token OpenAI results, that out-of-domain performance can differ substantially from the original report, and that length sensitivity is a real limitation. The extensions (quantization defense, password attack, Pareto-optimal hyperparameters) are practical, and the code is publicly released. However, the current wording of the headline claim overstates the evidence; the main value is in documenting partial or failed reproduction rather than confirming the original numbers. With an appropriately recalibrated abstract and a controlled test of the proposed explanations, the paper would be a solid contribution to the reproducibility literature.
major comments (3)
- [Abstract; §5.1.1 (Table 1); §5.1.2 (Table 2); §4.2.2] The statement that the authors 'successfully replicate the original key results in both in-domain and out-of-domain settings, with only minor discrepancies' (Abstract) is contradicted by the paper's own tables. In Table 1, the OpenAI-MSMARCO 32-token condition reports BLEU 83.4 (original) vs 43.4 (ours) and Exact-match 60.9 vs 4.8; §4.2.2 acknowledges that the ada-ms-32 checkpoint was unavailable and ada-ms-128 was used instead, so this column is not a replication of the original condition. In Table 2, out-of-domain gaps exceed 10% on Signal1M (80.7 vs 57.7), NQ (32.7 vs 14.7), BioASQ (22.8 vs 8.6), SciFact (16.6 vs 9.1), and TREC-News (14.5 vs 7.9). The abstract's characterization of these as 'minor discrepancies' is not supported by the data.
- [Table 1, gtr-nq-32 rows; §5.1.1] The exact-match discrepancy for gtr-nq-32 at 20 iterative steps is substantial and in the opposite direction of the length-mismatch explanation: the original reports 40.2 while the reproduced result is 58.2, even though the BLEU scores are close (83.9 vs 83.6). This large gap likely indicates a different test sample or a different decoding protocol, yet §5.1.1 states that the gtr-nq-32 results are 'nearly identical performance' across all metrics. The authors should either identify the cause (e.g., test-set mismatch) or soften the claim to reflect the exact-match divergence.
- [§5.1.2] The suggested explanations for the out-of-domain gaps (multiple versions of ada-ms-128 and sample sizes of 90 vs 200) are not tested. The paper does not run any experiment that measures the effect of sample size on the reported metrics, nor does it compare different checkpoints where multiple versions exist. Until such a controlled comparison is provided, the claim that the discrepancies are due to these artifacts remains a hypothesis, not a finding. The authors should add an ablation varying sample size on at least one dataset, or explicitly state that the cause is unresolved.
minor comments (5)
- [Figure 3] The figure content and caption are garbled in the submitted manuscript, with /uni... sequences appearing in place of the expected plot labels and text; please regenerate the figure so that the comparison is legible.
- [Table 4] The column headers for gtr-nq-32 and ada-ms-128 are not clearly separated in the table layout; the current formatting makes it difficult to determine which Exact-match and Token F1 columns belong to which model.
- [§5.1.2] There is a typo in the sentence 'the out-of-domain performance of Vex2Text on several datasets'; 'Vex2Text' should be 'Vec2Text'.
- [§5.1.1] The speculation that the poor 32-token performance of ada-ms-128 'is due to specific adjustments made by Morris et al. during the training of ada-ms-128' should be either supported with evidence from the released code or training logs, or removed.
- [§4.3] The paper does not state whether BLEU is computed with the same tokenizer as the original paper; if a different tokenizer is used, that alone could produce non-trivial score differences and would be important to document.
Circularity Check
No circularity: this is an empirical reproducibility comparison against externally reported Vec2Text numbers, with no fitted parameter renamed as a prediction and no load-bearing self-citation.
full rationale
The paper's derivation chain is an empirical benchmark comparison. The central claims are: (1) the released checkpoints gtr-nq-32 and ada-ms-128 reproduce Morris et al.'s reported numbers; (2) Gaussian noise degrades reconstruction while preserving retrieval; (3) 8-bit quantization degrades reconstruction while preserving retrieval; and (4) Vec2Text can recover some passwords. None of these is defined in terms of another claim. The paper explicitly does not train inversion models, stating in Section 4.4: 'We do not train our inversion models; instead, we utilize the provided weights for the reproducibility experiments,' so no parameter is fitted to a subset of data and then renamed a prediction. The comparison target is the externally reported results in Morris et al. [30], not a quantity constructed from this paper's own outputs. The paper honestly flags its own limitations, including 'we lack crucial experimental details, including the exact models used, sample sizes, and hyperparameters' (Section 4.4) and 'we face reproduction challenges due to an underspecified experimental setup' (Section 5.1.2). The unresolved gaps visible in Tables 1 and 2, such as the OpenAI-MSMARCO 32-token condition where BLEU is 83.4 vs 43.4 because ada-ms-32 was unavailable and ada-ms-128 was used instead, are correctness and reporting risks that undercut the abstract's 'successfully replicate' wording, but they are not circularity. Similarly, the quantization extension's BLEU drop is directionally plausible because quantization injects error, yet the retrieval-preservation component is a separate empirical measurement, and the paper does not derive the BLEU decrease from the quantization equations by definition. The only self-citations are related-work references [22] and [23] by co-author Yongkang Li, and they are not load-bearing for any reproduction or extension claim. No self-definitional step, fitted-input-called-prediction step, or author-imported uniqueness argument appears anywhere in the paper. The strongest defensible finding is therefore no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Released checkpoints (gtr-nq-32 and ada-ms-128) are the same models whose scores Morris et al. reported, or are close enough for direct comparison.
- domain assumption Evaluation metrics, tokenization, and sample selection match Morris et al.'s pipeline.
- domain assumption Sampling of 1000 items per dataset or password difficulty class is representative.
- domain assumption Quantization preserves retrieval-relevant coarse structure while destroying inversion-relevant fine detail, and the adversary does not adapt to the quantization scheme.
Cite this review
Pith. "Pith review of Rethinking the Privacy of Text Embeddings: A Reproducibility Study of "Text Embeddings Reveal (Almost) As Much As Text"." pith.science (2026). https://pith.science/paper/6GQHNANM
@misc{pith2026250707700,
author = {Pith},
title = {Pith review of: Rethinking the Privacy of Text Embeddings: A Reproducibility Study of "Text Embeddings Reveal (Almost) As Much As Text"},
year = {2026},
howpublished = {\url{https://pith.science/paper/6GQHNANM}},
note = {Machine review of arXiv:2507.07700}
}
read the original abstract
Text embeddings are fundamental to many natural language processing (NLP) tasks, extensively applied in domains such as recommendation systems and information retrieval (IR). Traditionally, transmitting embeddings instead of raw text has been seen as privacy-preserving. However, recent methods such as Vec2Text challenge this assumption by demonstrating that controlled decoding can successfully reconstruct original texts from black-box embeddings. The unexpectedly strong results reported by Vec2Text motivated us to conduct further verification, particularly considering the typically non-intuitive and opaque structure of high-dimensional embedding spaces. In this work, we reproduce the Vec2Text framework and evaluate it from two perspectives: (1) validating the original claims, and (2) extending the study through targeted experiments. First, we successfully replicate the original key results in both in-domain and out-of-domain settings, with only minor discrepancies arising due to missing artifacts, such as model checkpoints and dataset splits. Furthermore, we extend the study by conducting a parameter sensitivity analysis, evaluating the feasibility of reconstructing sensitive inputs (e.g., passwords), and exploring embedding quantization as a lightweight privacy defense. Our results show that Vec2Text is effective under ideal conditions, capable of reconstructing even password-like sequences that lack clear semantics. However, we identify key limitations, including its sensitivity to input sequence length. We also find that Gaussian noise and quantization techniques can mitigate the privacy risks posed by Vec2Text, with quantization offering a simpler and more widely applicable solution. Our findings emphasize the need for caution in using text embeddings and highlight the importance of further research into robust defense mechanisms for NLP systems.
Figures
Forward citations
Cited by 1 Pith paper
-
SHARD: cell-keyed residual splitting for alignment-resistant private dense retrieval
SHARD shards private embedding residuals into cell-local keyed groups to raise the anchor requirement for alignment attacks by a factor of C while preserving full-dimensional nDCG@10 via encrypted reranking.
Reference graph
Works this paper leans on
-
[1]
Mohamed Abdalla, Moustafa Abdalla, Graeme Hirst, and Frank Rudzicz. 2020. Exploring the Privacy-Preserving Properties of Word Embeddings: Algorithmic Validation Study. J Med Internet Res 22, 7 (15 Jul 2020), e18055. https://doi.org/ 10.2196/18055
-
[2]
Bhavik B. 2019. Password Strength Classifier Dataset. Kaggle. https://www. kaggle.com/datasets/bhavikbb/password-strength-classifier-dataset Accessed: 2025-05-08
work page 2019
-
[3]
Vera Boteva, Demian Gholipour Ghalandari, Artem Sokolov, and Stefan Riezler
-
[4]
Yiyi Chen, Heather C. Lent, and Johannes Bjerva. 2024. Text Embedding Inversion Security for Multilingual Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024 , Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association fo...
-
[5]
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale. In Ad- vances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 30318–30332. https://proceedings.neurips.cc/paper_file...
work page 2022
-
[6]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Associa- tion for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, ...
2019
-
[7]
Boyd-Graber, Jannis Bulian, Massimiliano Cia- ramita, and Markus Leippold
Thomas Diggelmann, Jordan L. Boyd-Graber, Jannis Bulian, Massimiliano Cia- ramita, and Markus Leippold. 2020. CLIMATE-FEVER: A Dataset for Verification of Real-World Climate Claims. CoRR abs/2012.00614 (2020). arXiv:2012.00614 https://arxiv.org/abs/2012.00614
arXiv 2020
-
[8]
Alexey Dosovitskiy and Thomas Brox. 2016. Inverting Visual Representations with Convolutional Networks. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016 . IEEE Com- puter Society, 4829–4837. https://doi.org/10.1109/CVPR.2016.522
Show all 60 references
-
[9]
Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2024. The Faiss library. (2024). arXiv:2401.08281 [cs.LG]
2024 arXiv
-
[10]
Shankar Iyer, Nikhil Dandekar, and Kornél Csernai. 2017. First Quora Dataset Release: Question Pairs. https://quoradata.quora.com/First-Quora-Dataset- Release-Question-Pairs
2017
-
[11]
Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. 2016. MIMIC-III, a freely accessible critical care database.Scientific data 3, 1 (2016), 1–9. https://doi.org/10...
2016 doi
-
[12]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, ...
2020 doi
-
[13]
Dimitris Kastaniotis. 2022. Introduction to Weight Quantization. Towards Data Science. https://towardsdatascience.com/introduction-to-weight-quantization- 2494701b9c0c/ Accessed: September 8, 2025
2022
-
[14]
Omar Khattab and Matei Zaharia. 2020. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2020. ACM, 39–48. http...
2020
-
[15]
Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler
Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S. Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Skip-Thought Vectors. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Infor- mation Processing Systems 2015, December 7...
2015
-
[16]
Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob De- vlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob De- vlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and...
2019 doi
-
[17]
Le and Tomás Mikolov
Quoc V. Le and Tomás Mikolov. 2014. Distributed Representations of Sentences and Documents. In Proceedings of the 31th International Conference on Machine Learning, ICML 2014, Beijing, China, 21-26 June 2014 (JMLR Workshop and Con- ference Proceedings, Vol. 32). JMLR.org, 1188...
2014
-
[18]
Wal- lace
Eric Lehman, Sarthak Jain, Karl Pichotta, Yoav Goldberg, and Byron C. Wal- lace. 2021. Does BERT Pretrained on Clinical Notes Reveal Sensitive Data?. In Proceedings of the 2021 Conference of the North American Chapter of the Associa- tion for Computational Linguistics: Human L...
2021
-
[19]
Yibin Lei, Tao Shen, Yu Cao, and Andrew Yates. 2025. Enhancing Lexicon-Based Text Embeddings with Large Language Models. CoRR abs/2501.09749 (2025). https://doi.org/10.48550/ARXIV.2501.09749 arXiv:2501.09749
2025 doi
-
[20]
Yibin Lei, Di Wu, Tianyi Zhou, Tao Shen, Yu Cao, Chongyang Tao, and Andrew Yates. 2024. Meta-Task Prompting Elicits Embeddings from Large Language Mod- els. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL ...
2024 doi
-
[21]
Haoran Li, Mingshi Xu, and Yangqiu Song. 2023. Sentence Embedding Leaks More Information than You Expect: Generative Embedding Inversion Attack to Recover the Whole Sentence. In Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 20...
2023 doi
-
[22]
Yongkang Li, Panagiotis Eustratiadis, and Evangelos Kanoulas. 2025. Reproducing HotFlip for Corpus Poisoning Attacks in Dense Retrieval. In Advances in Infor- mation Retrieval - 47th European Conference on Information Retrieval, ECIR 2025, Lucca, Italy, April 6-10, 2025, Proce...
2025 doi
-
[23]
Yongkang Li, Panagiotis Eustratiadis, Simon Lupart, and Evangelos Kanoulas
-
[24]
Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2024. Fine- Tuning LLaMA for Multi-Stage Text Retrieval. In Proceedings of the 47th Inter- national ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2024, Washington DC, USA, July 14-18...
2024
-
[25]
Aravindh Mahendran and Andrea Vedaldi. 2015. Understanding deep image representations by inverting them. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015 . IEEE Computer Society, 5188–5196. https://doi.org/10.1109/CVPR....
2015
-
[26]
Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018. WWW’18 Open Challenge: Fi- nancial Opinion Mining and Question Answering. In Companion of the The Web Conference 2018 on The Web Conference 2018, WWW 2018,...
2018
-
[27]
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2019. Putting Evaluation in Context: Contextual Embeddings Improve Machine Translation Evaluation. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- Augus...
2019 doi
-
[28]
Haitao Mi, Baskaran Sankaran, Zhiguo Wang, and Abe Ittycheriah. 2016. Cov- erage Embedding Models for Neural Machine Translation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, Texas, USA, November 1-4, 2016 , Jia...
2016 doi
-
[29]
Tomás Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient Estimation of Word Representations in Vector Space. In 1st International Con- ference on Learning Representations, ICLR 2013, Scottsdale, Arizona, USA, May 2-4, 2013, Workshop Track Proceedings , Yoshua B...
2013 arXiv
-
[30]
Morris, Volodymyr Kuleshov, Vitaly Shmatikov, and Alexander M
John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, and Alexander M. Rush
-
[31]
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. CoRR abs/1611.09268 (2016). arXiv:1611.09268 http://arxiv.org/abs/1611.09268
2016 arXiv
-
[32]
Zhao, Yi Luan, Keith B
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y. Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, and Yinfei Yang. 2022. Large Dual Encoders Are Generalizable Retrievers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Lan...
2022 doi
-
[33]
Shumpei Okura, Yukihiro Tagami, Shingo Ono, and Akira Tajima. 2017. Embedding-based News Recommendation for Millions of Users. In Proceed- ings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Halifax, NS, Canada, August 13 - 17, 2017 . A...
2017
-
[34]
OpenAI. 2022. New and improved embedding model. https://openai.com/index/ new-and-improved-embedding-model/. Accessed: 2025-05-06
2022
-
[35]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6-12, 2002, Philadelphia, PA, USA . ACL, 311–318....
2002 doi
-
[36]
Rahil Parikh, Christophe Dupuy, and Rahul Gupta. 2022. Canary Extraction in Natural Language Understanding Models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022 , ...
2022
-
[37]
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. Glove: Global Vectors for Word Representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, October 25-29, 2014, Doha, Qatar, A meeting of SIGDAT, a ...
2014 doi
-
[38]
Qdrant Team. 2023. Qdrant: High-performance, massive-scale vector database and vector search engine. https://qdrant.tech/. Accessed: 2025-05-06
2023
-
[39]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.J. Mach. Learn. Res. 21 (2020), 140:1–140:67. https://jmlr.org/...
2020
-
[40]
Ian Soboroff, Shudong Huang, and Donna Harman. 2018. TREC 2018 News Track Overview. In Proceedings of the Twenty-Seventh Text REtrieval Conference, TREC 2018, Gaithersburg, Maryland, USA, November 14-16, 2018 (NIST Special Publication, Vol. 500-331), Ellen M. Voorhees and Ange...
2018
-
[41]
Congzheng Song and Ananth Raghunathan. 2020. Information Leakage in Embedding Models. In CCS ’20: 2020 ACM SIGSAC Conference on Computer and Communications Security, Virtual Event, USA, November 9-13, 2020 , Jay Lig- atti, Xinming Ou, Jonathan Katz, and Giovanni Vigna (Eds.). ...
2020
-
[42]
Axel Suarez, Dyaa Albakour, David P. A. Corney, Miguel Martinez-Alvarez, and José Esquivel. 2018. A Data Collection for Evaluating the Retrieval of Related Tweets to News Articles. In Advances in Information Retrieval - 40th European Conference on IR Research, ECIR 2018, Greno...
2018 doi
-
[43]
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Tra...
2021
-
[44]
George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R. Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, Yannis Almirantis, John Pavlopou- los, Nicolas Baskiotis, Patrick Gallinari,...
2015
-
[45]
Rafael Veras, Christopher Collins, and Julie Thorpe. 2014. On Semantic Patterns of Passwords and their Security Impact. In 21st Annual Network and Distributed System Security Symposium, NDSS 2014, San Diego, California, USA, February 23-26, 2014. The Internet Society. https://...
2014
-
[46]
Voorhees
Ellen M. Voorhees. 2003. Overview of the TREC 2003 Robust Retrieval Track. In Proceedings of The Twelfth Text REtrieval Conference, TREC 2003, Gaithersburg, Maryland, USA, November 18-21, 2003 (NIST Special Publication, Vol. 500-255) , Ellen M. Voorhees and Lori P. Buckland (E...
2003
-
[47]
Henning Wachsmuth, Shahbaz Syed, and Benno Stein. 2018. Retrieval of the Best Counterargument without Prior Topic Knowledge. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: ...
2018
-
[48]
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or Fiction: Verifying Scientific Claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, No...
2020 doi
-
[49]
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2023. SimLM: Pre-training with Representation Bottleneck for Dense Passage Retrieval. In Proceedings of the 61st Annual Meeting of the Association for Computational Lin...
2023 doi
-
[50]
Yangde Wang, Weidong Qiu, Peng Tang, Hao Tian, and Shujun Li. 2025. SE#PCFG: Semantically Enhanced PCFG for Password Analysis and Cracking. IEEE Trans- actions on Dependable and Secure Computing (2025), 1–14. https://doi.org/10. 1109/TDSC.2025.3547773
2025
-
[51]
Chuhan Wu, Fangzhao Wu, Tao Qi, and Yongfeng Huang. 2021. Empowering News Recommendation with Pre-trained Language Models. InSIGIR ’21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event, Canada, July 11-15, 2021 , F...
2021
-
[52]
Haoran Xu, Benjamin Van Durme, and Kenton W. Murray. 2021. BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine Translation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta ...
2021 doi
-
[53]
Fuzheng Zhang, Nicholas Jing Yuan, Defu Lian, Xing Xie, and Wei-Ying Ma. 2016. Collaborative Knowledge Base Embedding for Recommender Systems. InProceed- ings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, August...
2016
-
[54]
Shengyao Zhuang, Bevan Koopman, Xiaoran Chu, and Guido Zuccon. 2024. Understanding and Mitigating the Threat of Vec2Text to Dense Retrieval Systems. In Proceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the...
2024
- [55]
-
[560]
https://doi.org/10.18653/V1/2022.ACL-SHORT.61
2022 doi
-
[1656]
https://doi.org/10.1145/3404835.3463069
-
[2016]
In Advances in Information Retrieval - 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20-23, 2016
A Full-Text Learning to Rank Dataset for Medical Information Retrieval. In Advances in Information Retrieval - 38th European Conference on IR Research, ECIR 2016, Padua, Italy, March 20-23, 2016. Proceedings (Lecture Notes in Computer Science, Vol. 9626), Nicola Ferro, Fabio C...
2016 doi
-
[2023]
Text Embeddings Reveal (Almost) As Much As Text. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics,...
2023 doi
-
[2025]
In Proceedings of the 48th International ACM SIGIR Conference on Re- search and Development in Information Retrieval, SIGIR 2025
Unsupervised Corpus Poisoning Attacks in Continuous Space for Dense Retrieval. In Proceedings of the 48th International ACM SIGIR Conference on Re- search and Development in Information Retrieval, SIGIR 2025 . https://doi.org/10. 1145/3726302.3730110
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.