REVIEW 1 major objections 1 minor 45 references
Think Harder and Don't Overlook Your Options: Revisiting Issue-Commit Linking with LLM-Assisted Retrieval
T0 review · 1 major / 1 minor · reviewed 2026-07-01 · grok-4.3
Pith's one-line read Dense retrieval methods outperform sparse ones and traditional machine learning rerankers outperform LLM-based methods for recovering issue-commit links.
desk verdict Dense retrieval plus traditional ML rerankers beat the LLM options on issue-commit linking, and the paper makes a reasonable case that simpler pipelines are still worth using. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The two-stage retrieval and reranking pipeline, where initial candidates are retrieved using methods like BM25, SBERT, ANNOY, LSH, or HNSW, and then refined by rerankers including traditional ML models or LLMs.
What would settle it
A controlled experiment on a new set of projects showing that an LLM-based reranker achieves higher precision and recall than the best traditional ML reranker would falsify the performance comparison.
Extended reading notes
Core claim
The central discovery is that in issue-commit link recovery, dense retrieval methods outperform sparse retrieval approaches in identifying relevant commits, with combinations improving recall, while traditional machine learning-based reranking techniques achieve higher performance than LLM-based approaches such as those using ChatGPT, Qwen, Gemma, or Llama. Retrieval-based pipelines remain a practical solution for large-scale linking.
Load-bearing premise
The study assumes that the compared retrieval methods and rerankers are representative of current best practices and that results on the evaluated projects apply more generally.
Editorial extensions
If this is right
- Dense retrieval can efficiently reduce the search space for commits.
- Hybrid dense-sparse retrieval enhances recall in candidate identification.
- Traditional ML rerankers provide better precision than current LLM methods for this task.
- Simpler models should be evaluated before adopting computationally expensive LLM approaches in traceability tools.
Reading between the lines
- The superiority of traditional methods may encourage more focus on optimizing existing techniques rather than defaulting to LLMs.
- These findings could inform similar linking tasks in other software artifacts beyond issues and commits.
- Performance differences might vary with project size or language, suggesting need for broader testing.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript revisits issue-commit linking by evaluating retrieval methods such as BM25, BM25L, SBERT-Semantic Search, ANNOY, LSH, HNSW for candidate retrieval and reranking techniques including traditional machine learning models, cross-encoders, and LLMs (ChatGPT, Qwen, Gemma, Llama) in comparison to established approaches like BTLink, EasyLink, FRLink, RCLinker, and Hybrid-Linker. It concludes that dense retrieval outperforms sparse retrieval, combining them improves recall, and traditional ML reranking outperforms LLM-based methods.
Significance. If the empirical findings are supported by rigorous experimental details, this study would be significant for software engineering practice by demonstrating that computationally efficient retrieval and traditional ML methods can be more effective than advanced LLM approaches for issue-commit linking, thereby informing the design of traceability tools and encouraging careful consideration of simpler solutions.
major comments (1)
- [Abstract] The abstract reports that dense retrieval methods outperform sparse ones and that traditional machine learning-based reranking achieves higher performance than LLM-based approaches, but supplies no details on the datasets used, the projects evaluated, statistical tests performed, effect sizes, or experimental controls. This absence makes it impossible to verify whether the reported performance differences support the central claims.
minor comments (1)
- The abstract could more explicitly state the evaluation metrics used (e.g., precision, recall, F1) to contextualize the results.
Simulated Author's Rebuttal
We thank the referee for the detailed review and constructive feedback. We address the major comment point by point below and commit to revisions that improve clarity without altering the core findings.
read point-by-point responses
-
Referee: [Abstract] The abstract reports that dense retrieval methods outperform sparse ones and that traditional machine learning-based reranking achieves higher performance than LLM-based approaches, but supplies no details on the datasets used, the projects evaluated, statistical tests performed, effect sizes, or experimental controls. This absence makes it impossible to verify whether the reported performance differences support the central claims.
Authors: We agree that the abstract as currently written is too high-level and omits essential experimental context. The full manuscript (Sections 3 and 4) details the datasets (standard issue-commit linking benchmarks), the specific projects evaluated, the use of statistical significance testing (paired t-tests with effect sizes reported via Cohen’s d), and controls for retrieval candidate size and model hyperparameters. In the revised manuscript we will expand the abstract to concisely include these elements—mentioning the projects, primary metrics, and confirmation of statistical testing—so that the claims can be evaluated from the abstract alone. revision: yes
Circularity Check
No significant circularity; purely empirical evaluation
full rationale
The paper conducts an empirical comparison of retrieval methods (BM25 variants, SBERT, ANNOY, etc.) and rerankers (ML models, cross-encoders, LLMs) on issue-commit linking tasks across selected projects. No derivations, equations, fitted parameters renamed as predictions, or self-referential definitions appear. Claims rest on experimental results rather than any internal reduction to inputs or self-citation chains. The listed prior techniques (BTLink, FRLink, etc.) are external baselines, and the evaluation pipeline is standard and falsifiable via replication on the same datasets.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Think Harder and Don't Overlook Your Options: Revisiting Issue-Commit Linking with LLM-Assisted Retrieval." pith.science (2026). https://pith.science/paper/VKEDAWWC
@misc{pith2026260500447,
author = {Pith},
title = {Pith review of: Think Harder and Don't Overlook Your Options: Revisiting Issue-Commit Linking with LLM-Assisted Retrieval},
year = {2026},
howpublished = {\url{https://pith.science/paper/VKEDAWWC}},
note = {Machine review of arXiv:2605.00447}
}
read the original abstract
Linking issue reports to the commits that resolve them is essential for software traceability, maintenance, and evolution. Accurate issue-commit links help developers to understand system changes and the rationale behind them. While numerous automated techniques have been proposed, ranging from heuristic and feature-based approaches to modern deep learning and large language model approaches, our goal is to evaluate these techniques to determine which are most effective and efficient. In this study, we revisit several established issue-commit link recovery techniques, including BTLink, EasyLink, FRLink, RCLinker, and Hybrid-Linker, and assess their performance for reranking issue-commit links. We first evaluate different retrieval methods (BM25, BM25L, SBERT-Semantic Search, ANNOY, LSH, HNSW) for their ability to efficiently retrieve relevant commits, reducing the candidate set that must be considered by more computationally expensive models. Using the best retrieval methods, we then investigate the reranking effectiveness of different machine learning-based techniques, including traditional machine learning models, a cross-encoder, and large language models (ChatGPT, Qwen, Gemma, Llama), to refine the reranking of candidate commits and improve precision. Finally, we compare the effectiveness of these techniques. Our results show that dense retrieval methods outperform sparse retrieval approaches in identifying relevant commits and that combining dense and sparse retrieval can improve recall. Additionally, we find that traditional machine learning-based reranking techniques achieve higher performance than LLM-based approaches. Our results highlight that retrieval-based pipelines remain a practical and effective solution for large-scale issue-commit linking, and that simpler models should be carefully considered before adopting computationally expensive LLM-based approaches.
Figures
Reference graph
Works this paper leans on
-
[1]
Martin Aumüller, Erik Bernhardsson, and Alexander Faithfull. 2020. ANN- Benchmarks: A benchmarking tool for approximate nearest neighbor algorithms. Information Systems (2020), 13 pages
work page 2020
-
[2]
Thazin Win Win Aung, Huan Huo, and Yulei Sui. 2020. A Literature Review of Automatic Traceability Links Recovery for Software Change Impact Analysis. In Proceedings of the 28th International Conference on Program Comprehension. 14–24
work page 2020
-
[3]
Adrian Bachmann, Christian Bird, Foyzur Rahman, Premkumar Devanbu, and Abraham Bernstein. 2010. The missing links: bugs and bug-fix commits. In Proceedings of the Eighteenth ACM SIGSOFT International Symposium on Foundations of Software Engineering. 97–106
work page 2010
-
[4]
Erik Bernhardsson. [n. d.]. Approximate Nearest Neighbors Oh Yeah (ANNOY). https://github.com/spotify/annoy Retrieved April 30, 2026
work page 2026
-
[5]
Christian Bird, Adrian Bachmann, Eirik Aune, John Duffy, Abraham Bernstein, Vladimir Filkov, and Premkumar Devanbu. 2009. Fair and balanced? bias in bug-fix datasets. In Proceedings of the 7th Joint Meeting of the European Software Engineering Conference and the ACM SIGSOFT Symposium on The Foundations of Software Engineering. 121–130
work page 2009
-
[6]
D. Brown. 2020. Rank-BM25: A Collection of BM25 Algorithms in Python. https://github.com/dorianbrown/rank_bm25 Retrieved April 30, 2026
work page 2020
-
[7]
Jane Cleland-Huang, Orlena C. Z. Gotel, Jane Huffman Hayes, Patrick Mäder, and Andrea Zisman. 2014. Software traceability: trends and future directions. In Future of Software Engineering Proceedings. 55–69
work page 2014
-
[8]
Jacob Cohen. 1960. A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement (1960), 37–46
work page 1960
Show all 45 references
-
[9]
Liming Dong, He Zhang, Wei Liu, Zhiluo Weng, and Hongyu Kuang. 2022. Semi-supervised pre-processing for learning-based traceability framework on real-world software projects. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Fou...
2022
-
[10]
Google. 2024. Gemma-7B-IT. https://huggingface.co/google/gemma-7b-it Re- trieved April 30, 2026
2024
-
[11]
Pengfei He, Shaowei Wang, Shaiful Chowdhury, and Tse-Hsun Chen. 2025. Eval- uating the Effectiveness and Efficiency of Demonstration Retrievers in RAG for Coding Tasks. InProceedings of 2025 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER...
2025
-
[12]
Huihui Huang, Ratnadira Widyasari, Ting Zhang, Ivana Clairine Irsan, Jieke Shi, Han Wei Ang, Frank Liauw, Eng Lieh Ouh, Lwin Khin Shar, Hong Jin Kang, and David Lo. 2025. Back to the Basics: Rethinking Issue-Commit Linking with LLM-Assisted Retrieval. InProceedings of the IEEE...
2025
-
[13]
Omid Jafari, Preeti Maurya, Parth Nagarkar, Khandker Mushfiqul Islam, and Chidambaram Crushev. 2021. A Survey on Locality Sensitive Hashing Algorithms and their Applications. arXiv (2021)
2021
-
[14]
Masanari Kondo, Yutaro Kashiwa, Yasutaka Kamei, and Osamu Mizuno. 2022. An empirical study of issue-link algorithms: which issue-link algorithms should we use? Empirical Software Engineering (2022), 50 pages
2022
-
[15]
Jinpeng Lan, Lina Gong, Jingxuan Zhang, and Haoxiang Zhang. 2023. BTLink : automatic link recovery between issues and commits based on pre-trained BERT model. Empirical Software Engineering (2023), 55 pages
2023
-
[16]
Le, Mario Linares-Vasquez, David Lo, and Denys Poshyvanyk
Tien-Duy B. Le, Mario Linares-Vasquez, David Lo, and Denys Poshyvanyk. 2015. RCLinker: Automated Linking of Issue Reports and Commits Leveraging Rich Contextual Information. In Proceedings of the 23rd International Conference on Program Comprehension. 36–47
2015
-
[17]
Jinfeng Lin, Yalin Liu, Qingkai Zeng, Meng Jiang, and Jane Cleland-Huang. 2021. Traceability Transformed: Generating more Accurate Links with Pre-Trained BERT Models. In Proceedings of the 43rd International Conference on Software Engineering. 324–335
2021
-
[18]
Yuanhua Lv and ChengXiang Zhai. 2011. When documents are very long, BM25 fails!. In Proceedings of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1103–1104
2011
-
[19]
Malkov and D
Yu A. Malkov and D. A. Yashunin. 2020. Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs. IEEE Trans. Pattern Anal. Mach. Intell. (2020), 824–836
2020
-
[20]
Pooya Rostami Mazrae, Maliheh Izadi, and Abbas Heydarnoori. 2021. Automated Recovery of Issue-Commit Links Leveraging Both Textual and Non-textual Data. In Proceedings of the 37th International Conference on Software Maintenance and Evolution (ICSME). 263–273
2021
-
[21]
Meta. 2024. Llama-3.1-8B-Instruct. https://huggingface.co/meta-llama/Llama- 3.1-8B-Instruct Retrieved April 30, 2026
2024
-
[22]
Anh Tuan Nguyen, Tung Thanh Nguyen, Hoan Anh Nguyen, and Tien N. Nguyen
-
[23]
In Proceedings of the ACM SIGSOFT 20th International Symposium on the Foundations of Software Engineering
Multi-layered approach for recovering links between bug reports and fixes. In Proceedings of the ACM SIGSOFT 20th International Symposium on the Foundations of Software Engineering. 11 pages
-
[24]
Replication Package. [n. d.]. Think Harder and Don’t Overlook Your Options: Revisiting Issue-Commit Linking with LLM-Assisted Retrieval. https://figshare. com/s/3b9673176afe92929398 Retrieved April 30, 2026
2026
-
[25]
Michael C. Panis. 2010. Successful Deployment of Requirements Traceability in a Commercial Engineering Organization...Really. In Proceedings of the 18th IEEE International Requirements Engineering Conference. 303–307
2010
-
[26]
Qwen. 2025. Qwen3-32B. https://huggingface.co/Qwen/Qwen3-32B Retrieved April 30, 2026
2025
-
[27]
Michael Rath, Jacob Rendall, Jin L. C. Guo, Jane Cleland-Huang, and Patrick Mäder. 2018. Traceability in the wild: automatically augmenting incomplete trace links. In Proceedings of the 40th International Conference on Software Engineering. 834–845
2018
-
[28]
Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embed- dings using siamese bert-networks. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-I...
2019
-
[29]
Rodriguez, Jane Cleland-Huang, and Davide Falessi
Alberto D. Rodriguez, Jane Cleland-Huang, and Davide Falessi. 2021. Lever- aging Intermediate Artifacts to Improve Automated Trace Link Retrieval. In Proceedings of the 37th International Conference on Software Maintenance and Evolution (ICSME). 81–92
2021
-
[30]
Hang Ruan, Bihuan Chen, Xin Peng, and Wenyun Zhao. 2019. DeepLink: Recov- ering issue-commit links based on deep learning. J. Syst. Softw. (2019), 13 pages
2019
-
[31]
Kunal Sawarkar, Abhilasha Mangal, and Shivam Raj Solanki. 2024. Blended RAG: Improving RAG (Retriever-Augmented Generation) Accuracy with Se- mantic Search and Hybrid Query-Based Retrievers. In Proceedings of the 7th International Conference on Multimedia Information Processin...
2024
-
[32]
Gerald Schermann, Martin Brandtner, Sebastiano Panichella, Philipp Leitner, and Harald Gall. 2015. Discovering Loners and Phantoms in Commit and Issue Data. In Proceedings of 23rd International Conference on Program Comprehension. 4–14
2015
-
[33]
Davide Spadini, Maurício Aniche, and Alberto Bacchelli. 2018. PyDriller: Python framework for mining software repositories. In Proceedings of the 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 908–911
2018
-
[34]
Yan Sun, Celia Chen, Qing Wang, and Barry Boehm. 2017. Improving missing issue-commit link recovery using positive and unlabeled data. In Proceedings of the 32nd IEEE/ACM International Conference on Automated Software Engineering (ASE). 147–152
2017
-
[35]
Yan Sun, Qing Wang, and Ye Yang. 2017. FRLink: Improving the recovery of missing issue-commit links by revisiting file relevance.Information and Software Technology (2017), 33–47
2017
-
[36]
Gemma Team. 2024. Gemma: Open Models Based on Gemini Research and Technology. arXiv (2024)
2024
-
[37]
Qwen Team. 2025. Qwen3 Technical Report. arXiv (2025)
2025
-
[38]
Andrew Trotman, Antti Puurula, and Blake Burgess. 2014. Improvements to BM25 and Language Models Examined. In Proceedings of the 19th Australasian Document Computing Symposium. 58–65
2014
-
[39]
Rongxin Wu, Hongyu Zhang, Sunghun Kim, and Shing-Chi Cheung. 2011. Re- Link: recovering links between bugs and changes. In Proceedings of the 19th ACM SIGSOFT Symposium and the 13th European Conference on Foundations of Software Engineering. 15–25
2011
-
[40]
Zhaonan Wu, Yanjie Zhao, Chen Wei, Zirui Wan, Yue Liu, and Haoyu Wang
-
[41]
In Proceedings of 47th International Conference on Software Engineering: Companion Proceedings (ICSE-Companion)
COmmitSHield: Tracking Vulnerability Introduction and Fix in Version Control Systems. In Proceedings of 47th International Conference on Software Engineering: Companion Proceedings (ICSE-Companion). 279–290
-
[42]
Rui Xie, Long Chen, Wei Ye, Zhiyu Li, Tianxiang Hu, Dongdong Du, and Shikun Zhang. 2019. DeepLink: A Code Knowledge Graph Based Deep Learning Ap- proach for Issue-Commit Link Recovery. InProceedings of the 26th International Conference on Software Analysis, Evolution and Reeng...
2019
-
[43]
Chenyuan Zhang, Yanlin Wang, Zhao Wei, Yong Xu, Juhong Wang, Hui Li, and Rongrong Ji. 2023. EALink: An Efficient and Accurate Pre-trained Frame- work for Issue-Commit Link Recovery. In Proceedings of the 38th IEEE/ACM International Conference on Automated Software Engineering ...
2023
-
[44]
Jianfei Zhu, Guanping Xiao, Zheng Zheng, and Yulei Sui. 2022. Enhancing Traceability Link Recovery with Unlabeled Data. In Proceedings of the 33rd International Symposium on Software Reliability Engineering (ISSRE). 446–457
2022
-
[45]
Jianfei Zhu, Guanping Xiao, Zheng Zheng, and Yulei Sui. 2024. Deep semi- supervised learning for recovering traceability links between issues and commits. Journal of Systems and Software (2024), 19 pages
2024
Reviewed July 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.