REVIEW 3 major objections 4 minor 2 cited by
Can Large Language Models Improve SE Active Learning via Warm-Starts?
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read For low- and medium-dimensional software engineering tasks, warm-starting active learning with LLM-generated examples sharply cuts the number of expensive labels needed, but for high-dimensional tasks Bayesian Gaussian-process methods…
desk verdict Interesting question and a large benchmark, but the LLM warm-start mechanism as written cannot produce the claimed gains because the synthetic rows are mapped back onto the same four random seed rows. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a two-step warm-start pipeline. An LLM is shown four randomly chosen rows labeled 'best' or 'rest' by Chebyshev distance to the ideal objective values, and is prompted to invent two better and two worse input vectors. Because invented rows carry no observed labels, each is mapped to its nearest neighbor in the initial four-row set, measured by Euclidean distance on the independent variables, and those real rows become the warm start. This mapping turns synthetic suggestions into labelable examples; afterward, standard acquisition functions—UCB, PI, and EI for Gaussian Process Models; Explore and Exploit for Tree of Parzen Estimators—drive the active-learning loop. The Chebyshev best/rest split is what gives the LLM its notion of 'better' and 'worse'.
What would settle it
Run the same 49-task experiments with the nearest-neighbor mapping replaced by a random row drawn from the initial set; if the LLM warm start's advantage over random cold starts disappears, or is matched by randomly chosen warm starts, then the Euclidean matching rule—not the LLM's synthesis—carries the result. A cheaper check is to compare the Chebyshev labels of the nearest-neighbor rows against the labels of the initial random rows on low-dimensional datasets where LLM/exploit wins; if the matched rows are not systematically better, the mechanism is not transferring quality.
Extended reading notes
Core claim
The central empirical claim is that warm-starting an active learner with LLM-synthesized examples gives large, statistically significant gains for low- and medium-dimensional multi-objective SE tasks, and that this gain vanishes in high-dimensional tasks. In the paper's statistical rankings, LLM warm starts followed by the exploit acquisition function reached the top rank in 100% of low-dimensional datasets and 50% of medium-dimensional datasets, compared with 27% and 21% for random starts with the same acquisition function. For high-dimensional datasets, random starts followed by UCB_GPM were top in 50% of datasets and outperformed every LLM-based combination. The authors interpret this as evidence that LLMs can extrapolate 'better' and 'worse' directions in simple search spaces, but lose that ability when many independent variables interact.
Load-bearing premise
The result depends on the assumption that when the LLM invents a promising new input combination, the closest real row already in the initial random sample is also promising; if Euclidean distance in the input space does not track solution quality, the warm start becomes no better than random.
Editorial extensions
If this is right
- For SE tasks with fewer than about a dozen independent variables, practitioners can expect active learning with LLM warm starts to reach good optima within a few dozen labels, reducing labeling effort compared with random starts.
- Starting with just four LLM-guided labels can produce large improvements: LLM/exploit reached the top ranking in 100% of low-dimensional and 50% of medium-dimensional datasets, versus 27% and 21% for random/exploit.
- For high-dimensional tasks, random starts plus UCB_GPM remain the strongest configuration, so using LLM warm starts there is not supported by these results.
- The acquisition function matters: LLM warm starts pair well with exploit but not with explore, suggesting the LLM already covers the boundary regions that explore would target.
- Overall, the paper supports a regime-dependent recipe—LLM warm starts for simple spaces and Bayesian learners for complex ones—rather than a single method that wins everywhere.
Reading between the lines
- The paper leaves implicit that the nearest-neighbor transfer step ties the LLM's value to the density of the initial random sample; denser initial sampling could extend LLM warm starts to higher dimensions without retraining the model.
- A testable extension is to apply dimensionality reduction before prompting, so the LLM extrapolates in a lower-dimensional space; if the high-dimensional failure is a geometry problem, this could push LLM warm starts into the regime where Bayesian methods currently win.
- Although the paper measures runtimes, it does not make cost the headline; those measurements hint that LLM few-shot warm starts can be cheaper than fitting Gaussian processes, which could make LLM warm starts attractive even when their accuracy gain is modest.
- One could benchmark the same pipeline on non-software tabular optimization problems to see whether the dimensionality cutoff is a property of LLM reasoning or of the SE data distribution.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes using large language models (LLMs) to generate warm starts for active learning in multi-objective software engineering optimization. The authors use Gemini to synthesize candidate x-vectors from a few labeled examples, map those candidates back to labeled rows via nearest-neighbor matching, and compare the resulting warm starts against random starts under Gaussian Process (UCB, PI, EI) and TPE (explore, exploit) acquisition functions. Experiments on 49 MOOT datasets are summarized with Scott-Knott rankings, and the paper's central claim (RQ3) is that LLM warm starts followed by exploit-guided acquisition perform best on low- and medium-dimensional tasks, while random starts followed by UCB_GPM perform best on high-dimensional tasks.
Significance. If the central claim were supported, the paper would fill a clear gap in the active-learning-for-SE literature: it tests many more datasets than most prior work, it considers multi-objective problems, and it restricts the labeling budget to a few dozen examples, which is well motivated by the reported SME labeling rates. The authors also release data and scripts, which is a genuine strength for reproducibility. However, the significance is conditional because, as written, the mechanism described for transferring LLM-synthesized information to the labeled seed set cannot introduce any new labeled point beyond the initial random seed, so the reported RQ3 gains cannot be explained by the described method.
major comments (3)
- [§3.3.3, steps 5–6; §3.6] The warm-start construction as written cannot transmit any new information from the LLM into the labeled seed set. Step 5 says that for each invented row r in E1, the nearest neighbor is found in E0 and the result forms E2⊂E0. Since §3.6 fixes B0=4, E0 has only four rows, so E2 contains no point outside E0; if the four nearest-neighbor matches are distinct, E2=E0 and the LLM condition and the random condition begin from the same labeled rows. If the matches collide, E2 has fewer than four distinct rows, which would likely hurt rather than help. The reported gains in Tables 9–11 therefore cannot be explained by the mechanism described in the manuscript. The authors must either correct the description of the actual experimental procedure (e.g., if the code maps invented rows to the full training set rather than to E0) or rerun the experiments with a method that actually allows LLM-synthesized coordinates to enter the warm-start set.
- [§3.1 and Table 1; Tables 9–11] The dimensionality stratification is internally inconsistent. §3.1 defines low as |x|<6, medium as 6≤|x|≤11, and high as |x|>11, but Table 1 labels SS-L (11 x-variables) as low and SS-O and SS-P (also 11 x-variables) as high. In addition, §3.1 states there are 14 medium datasets, while Table 10 reports results for 16 medium datasets. Since the paper's headline conclusion is specifically about low-, medium-, and high-dimensional behavior, these inconsistencies affect the interpretation of Tables 9–11 and must be reconciled.
- [§3.3.3, step 5] The nearest-neighbor matching rule relies on an unsupported assumption that Euclidean distance in x-space preserves the quality ranking intended by the LLM. This assumption is load-bearing if the method is corrected to map invented rows to the full training set: even a perfect LLM suggestion could be mapped to a row whose objective quality is unrelated to the LLM's intent. The paper provides no validation of this mapping (e.g., comparing the Chebyshev scores of the matched rows with the scores of the LLM-invented coordinates), and it should either justify the assumption or test it directly.
minor comments (4)
- [Table 1] The column header '𝑥/𝑦' is described as 'independent𝑥 vales and dependent𝑦 values', but the table entries are formatted inconsistently (e.g., '3/2' vs '9/1'); this is understandable but should be cleaned up.
- [§3.5, Eq. (5)] Equation (5) is incomplete as printed: it ends with a dangling '+' and the terms inside the expected difference are not fully defined. Please provide a complete, self-contained definition of the Scott-Knott splitting criterion.
- [§4.1] The sentence 'In no case in Table 9, 10 or 9 did a purely random method score appear most often in the rank “0” column' contains a typo ('Table 9, 10 or 9' should presumably be 'Tables 9, 10, or 11') and is ambiguous about whether 'random' refers to the 'Acquire=random' rows or to the treatment labeled 'random' in the tables.
- [§3.3.3, Table 7 paragraph] There is a typo: 'few-short learner' should be 'few-shot learner'. Similar small typos elsewhere (e.g., 'all the reasons pf the last paragraph' in §3.2, 'sections' for 'seconds' in §2.4.3) should be corrected in a revision.
Circularity Check
The LLM warm-start set E2 is defined as a subset of the random seed E0 (with B0=4), so RQ3's claimed LLM advantage cannot be carried by any LLM-synthesized row.
-
self definitional
[§3.3.3 (steps 5-6), §3.6, §4.3 (RQ3)]
"(5) Since all the items in E1 were invented by the LLM, they have no y column labels. To generate labeled examples, we then went back to the training data and for each row r∈E1, we found its nearest neighbor from E0 (where “near” was measured using the Euclidean distance of the independent x column values). This formed the set E2⊂E0. (6) E2 was then used to warm-start the active learners. ... We ran all our active learners with B0 evaluations for the warm starts and B1 total evaluations where B0 = 4"
By construction, step 5 restricts E2 to E0, the initial randomly selected seed rows, and §3.6 fixes |E0|=4. The LLM's invented rows E1 are therefore never used as labeled points; they can at most select or reorder their nearest neighbors inside E0. If the four nearest-neighbor matches are distinct, E2=E0 and the LLM and random conditions begin with the same labeled set; if matches collide, E2 is a strict subset with fewer distinct rows. RQ3's claim that LLM-generated warm starts outperform random starts is thus not supported by the described pipeline: no LLM-synthesized x-vector enters the active learner, so the reported advantage in Tables 9-11 cannot be attributed to the LLM. The comparison reduces by construction to a subset-selection artifact defined over the random seed.
full rationale
Apart from this constructional collapse, the study is largely self-contained and empirically organized: the Chebyshev distance is used consistently both to rank seed rows and to score final solutions, and the low/medium/high split is imported from Di Fiore et al. as a categorization rather than as a derived result. The many self-citations are background, data provenance, or incidental SME-rate claims, not load-bearing uniqueness theorems. The nearest-neighbor matching could have been a legitimate channel if E1 rows were matched against the full training data and labeled with their actual y values, but as written E2 is defined as a subset of E0, so the central RQ3 comparison collapses by definition. This is an internal-consistency and constructional-circularity problem, not a dispute with community consensus; score 6 reflects that the central prediction reduces by construction while the supporting RQ1/RQ2 analyses and the dataset comparison retain independent content.
Assumptions & free parameters
free parameters (4)
- Warm-start seed size B0 =
4
- Best/rest split threshold =
sqrt(N) rows
- Dimension cutoffs =
6 and 11
- TPE explore epsilon =
unspecified small positive
assumptions (4)
- ad hoc to paper Euclidean distance in x-space is a valid measure of similarity for nearest-neighbor matching
- domain assumption Chebyshev distance to the ideal y-vector is a valid scalarization of multi-objective quality
- domain assumption Labeling y-values is expensive, so evaluating with at most 30 labels is the right regime
- standard math Scott-Knott with bootstrapping and Cliff's Delta is an appropriate significance test
Cite this review
Pith. "Pith review of Can Large Language Models Improve SE Active Learning via Warm-Starts?." pith.science (2026). https://pith.science/paper/VNVILF7T
@misc{pith2026250100125,
author = {Pith},
title = {Pith review of: Can Large Language Models Improve SE Active Learning via Warm-Starts?},
year = {2026},
howpublished = {\url{https://pith.science/paper/VNVILF7T}},
note = {Machine review of arXiv:2501.00125}
}
read the original abstract
When SE data is scarce, "active learners" use models learned from tiny samples of the data to find the next most informative example to label. In this way, effective models can be generated using very little data. For multi-objective software engineering (SE) tasks, active learning can benefit from an effective set of initial guesses (also known as "warm starts"). This paper explores the use of Large Language Models (LLMs) for creating warm-starts. Those results are compared against Gaussian Process Models and Tree of Parzen Estimators. For 49 SE tasks, LLM-generated warm starts significantly improved the performance of low- and medium-dimensional tasks. However, LLM effectiveness diminishes in high-dimensional problems, where Bayesian methods like Gaussian Process Models perform best.
Figures
Forward citations
Cited by 2 Pith papers
-
When to Use Which? Benchmarking Optimisers for Configurable Systems under Varying Budgets
Across 22 configurable systems and budgets from 100 to 10,000 evaluations, FLASH is the most consistently effective optimiser, while GA and IRACE catch up only at large budgets.
-
BINGO! Simple Optimizers Win Big if Problems Collapse to a Few Buckets
SE optimization data clusters into a tiny fraction of possible buckets, so simple stochastic samplers can match DEHB with far less computation.
Reference graph
Works this paper leans on
-
[1]
Toufique Ahmed and Premkumar Devanbu. 2023. Few-shot training LLMs for project-specific code-summarization. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering (<conf-loc>, <city>Rochester</city>, <state>MI</state>, <coun- try>USA</country>, </conf-loc>) (ASE ’22). Association for Computing Machinery, New York, N...
arXiv 2023
-
[2]
Waad Alhoshan, Liping Zhao, Alessio Ferrari, and Keletso J Letsholo. 2022. A zero-shot learning approach to classifying requirements: A preliminary study. In International Working Conference on Requirements Engineering: Foundation for Software Quality . Springer, 52–59
2022
- [3]
-
[4]
Jordan Ash and Ryan P Adams. 2020. On warm-starting neural network training. Advances in neural information processing systems 33 (2020), 3884–3894
2020
-
[5]
Yikun Ban, Yuheng Zhang, Hanghang Tong, Arindam Banerjee, and Jingrui He. 2022. Improved algorithms for neural active learning. Advances in Neural Information Processing Systems 35 (2022), 27497–27509
2022
-
[6]
James Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl. 2011. Algorithms for hyper-parameter optimization. Advances in neural information processing systems 24 (2011)
2011
-
[7]
Muhammad Bilal, Marco Serafini, Marco Canini, and Rodrigo Rodrigues. 2020. Do the best cloud configurations grow on trees? an experimental evaluation of black box algorithms for optimizing cloud workloads. (2020)
2020
-
[8]
Barry W Boehm and Richard Turner. 2004. Balancing agility and discipline: A guide for the perplexed . Addison-Wesley Professional
2004
Show all 88 references
-
[9]
Klaus Brinker. 2003. Incorporating diversity in active learning with support vector machines. In Proceedings of the 20th international conference on machine learning (ICML-03). 59–66
2003
-
[10]
Eric Brochu, Vlad M Cora, and Nando De Freitas. 2010. A tutorial on Bayesian optimization of expensive cost functions, with application to active user modeling and hierarchical reinforcement learning. arXiv preprint arXiv:1012.2599 (2010)
2010 arXiv
-
[11]
Tom B Brown. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 (2020)
2020 arXiv
-
[12]
Gemma Catolino. 2017. Just-in-time bug prediction in mobile applications: the domain matters!. In 2017 IEEE/ACM 4th International Conference on Mobile Software Engineering and Systems (MOBILESoft) . IEEE, 201–202
2017
-
[13]
Sampling
Jianfeng Chen, Vivek Nair, Rahul Krishna, and Tim Menzies. 2018. “Sampling” as a baseline optimizer for search-based software engineering. IEEE Transactions on Software Engineering 45, 6 (2018), 597–614
2018
-
[14]
Jianfeng Chen, Vivek Nair, and Tim Menzies. 2017. Beyond Evolutionary Algorithms for Search-based Software Engineering. Information and Software Technology 2017 (2017)
2017
-
[15]
Liangyu Chen, Yutong Bai, Siyu Huang, Yongyi Lu, Bihan Wen, Alan Yuille, and Zongwei Zhou. 2024. Making your first choice: to address cold start problem in medical active learning. In Medical Imaging with Deep Learning . PMLR, 496–525
2024
-
[16]
Tsong Yueh Chen, Fei-Ching Kuo, Huai Liu, Pak-Lok Poon, Dave Towey, TH Tse, and Zhi Quan Zhou. 2018. Metamorphic testing: A review of challenges and opportunities. ACM Computing Surveys (CSUR) 51, 1 (2018), 1–27
2018
-
[17]
Dennis D Cox and Susan John. 1992. A statistical method for global optimization. In [Proceedings] 1992 IEEE international conference on systems, man, and cybernetics. IEEE, 1241–1246
1992
-
[18]
Arnav Mohanty Das, Gantavya Bhatt, Megh Manoj Bhalerao, Vianne R Gao, Rui Yang, and Jeff Bilmes. 2023. Continual Active Learning. (2023)
2023
-
[19]
K. Deb, L. Thiele, M. Laumanns, and E. Zitzler. 2002. Scalable multi-objective optimization test problems. In Proceedings of the 2002 Congress on Evolutionary Computation. CEC’02 (Cat. No.02TH8600) , Vol. 1. 825–830 vol.1. https://doi.org/10.1109/CEC.2002.1007032
2002 arXiv
-
[20]
Di Fiore, M
F. Di Fiore, M. Nardelli, and L. Mainini. 2024. Active Learning and Bayesian Optimization: A Unified Perspective to Learn with a Goal. Archives of Computational Methods in Engineering (2024), 1–29
2024
-
[21]
Manqing Dong, Feng Yuan, Lina Yao, Xiwei Xu, and Liming Zhu. 2020. Mamo: Memory-augmented meta-optimization for cold-start recommendation. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining . 688–697
2020
-
[22]
Mark Easterby-Smith. 1980. The design, analysis and interpretation of repertory grids. International Journal of Man-Machine Studies 13, 1 (1980), 3–24. https://doi.org/10.1016/S0020-7373(80)80032-0
1980 doi
-
[23]
Feather and Tim Menzies
Martin S. Feather and Tim Menzies. 2002. Converging on the Optimal Attainment of Requirements. In 10th Anniversary IEEE Joint International Conference on Requirements Engineering (RE 2002), 9-13 September 2002, Essen, Germany . IEEE Computer Society, 263–272. https://doi.org/1...
2002 arXiv
-
[24]
Ioannis Giagkiozis and Peter J Fleming. 2015. Methods for multi-objective optimization: An analysis. Information Sciences 293 (2015), 338–350
2015
-
[25]
Phillip Green, Tim Menzies, Steven Williams, and Oussama El-Rawas. 2009. Understanding the value of software engineering technologies. In 2009 IEEE/ACM International Conference on Automated Software Engineering . IEEE, 52–61
2009
-
[26]
Guy Hacohen, Avihu Dekel, and Daphna Weinshall. 2022. Active learning on a budget: Opposite strategies suit high and low budgets. arXiv preprint arXiv:2202.02794 (2022)
2022 arXiv
-
[27]
Mark Harman, S Afshin Mansouri, and Yuanyuan Zhang. 2012. Search-based software engineering: Trends, techniques and applications. ACM Computing Surveys (CSUR) 45, 1 (2012), 11
2012
-
[28]
Abram Hindle, Daniel M German, and Ric Holt. 2008. What do large commits tell us?: a taxonomical study of large commits. In Proceedings of the 2008 international working conference on Mining software repositories . ACM, 99–108
2008
-
[29]
Hartwig H Hochmair, Levente Juhász, and Takoda Kemp. 2024. Correctness Comparison of ChatGPT-4, Gemini, Claude-3, and Copilot for Spatial Tasks. Transactions in GIS (2024)
2024
-
[30]
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2024. Large Language Models for Software Engineering: A Systematic Literature Review. ACM Trans. Softw. Eng. Methodol. 33, 8, Article 220 (Dec. 2024), 79 pages. h...
2024 doi
-
[31]
Peiyun Hu, Zachary C Lipton, Anima Anandkumar, and Deva Ramanan. 2018. Active learning with partial feedback. arXiv preprint arXiv:1802.07427 (2018)
2018 arXiv
-
[32]
Donald R Jones, Matthias Schonlau, and William J Welch. 1998. Efficient global optimization of expensive black-box functions. Journal of Global optimization 13 (1998), 455–492
1998
-
[33]
Yasutaka Kamei, Emad Shihab, Bram Adams, Ahmed E Hassan, Audris Mockus, Anand Sinha, and Naoyasu Ubayashi. 2012. A large-scale empirical study of just-in-time quality assurance. IEEE Transactions on Software Engineering 39, 6 (2012), 757–773
2012
-
[34]
Hong Jin Kang, Khai Loong Aw, and David Lo. 2022. Detecting False Alarms from Automatic Static Analysis Tools: How Far Are We?. InProceedings of the 44th International Conference on Software Engineering (Pittsburgh, Pennsylvania) (ICSE ’22). Association for Computing Machinery...
2022
-
[35]
Sunghun Kim, E James Whitehead Jr, and Yi Zhang. 2008. Classifying software changes: Clean or buggy? IEEE Transactions on Software Engineering 34, 2 (2008), 181–196
2008
-
[36]
Alison Kington. 2009. Defining Teachers’ Classroom Relationships. (2009)
2009
-
[37]
Ksenia Konyushkova, Raphael Sznitman, and Pascal Fua. 2017. Learning active learning from data. Advances in neural information processing systems 30 (2017)
2017
-
[38]
Ksenia Konyushkova, Raphael Sznitman, and Pascal Fua. 2018. Discovering general-purpose active learning strategies.arXiv preprint arXiv:1810.04114 (2018)
2018 arXiv
-
[39]
Harold J Kushner. 1964. A new method of locating the maximum point of an arbitrary multipeak curve in the presence of noise. (1964)
1964
-
[40]
Van-Hoang Le and Hongyu Zhang. 2023. Log parsing with prompt-based few-shot learning. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 2438–2449
2023
-
[41]
Yan-Hui Lin, Ze-Qi Ding, and Yan-Fu Li. 2023. Similarity based remaining useful life prediction based on Gaussian Process with active learning. Reliability Engineering & System Safety 238 (2023), 109461
2023
-
[42]
Ming Liu, Wray Buntine, and Gholamreza Haffari. 2018. Learning how to actively learn: A deep imitation learning approach. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 1874–1883
2018
-
[43]
Tennison Liu, Nicolás Astorga, Nabeel Seedat, and Mihaela van der Schaar. 2024. Large language models to enhance bayesian optimization. arXiv preprint arXiv:2402.03921 (2024)
2024 arXiv
-
[44]
David Lowell, Zachary C Lipton, and Byron C Wallace. 2018. Practical obstacles to deploying active learning. arXiv preprint arXiv:1807.04801 (2018)
2018 arXiv
-
[45]
Andre Lustosa and Tim Menzies. 2024. Learning from Very Little Data: On the Value of Landscape Analysis for Predicting Software Project Health. ACM Transactions on Software Engineering and Methodology 33, 3 (2024), 1–22
2024
-
[46]
Rafid Mahmood, Sanja Fidler, and Marc T Law. 2021. Low budget active learning via wasserstein distance: An integer programming approach. arXiv preprint arXiv:2106.02968 (2021)
2021 arXiv
-
[47]
co-training
Suvodeep Majumder, Joymallya Chakraborty, and Tim Menzies. 2024. When less is more: on the value of “co-training” for semi-supervised software defect predictors. Empirical Software Engineering 29, 2 (2024), 1–33
2024
-
[48]
George Mathew, Tim Menzies, Neil A Ernst, and John Klein. 2017. “SHORT” er Reasoning About Larger Requirements Models. In 2017 IEEE 25th International Requirements Engineering Conference (RE) . IEEE, 154–163
2017
-
[49]
Tim Menzies. 1999. Critical success metrics: evaluation at the business level. International journal of human-computer studies 51, 4 (1999), 783–799
1999
-
[50]
Tim Menzies, John Black, Joel Fleming, and Murray Dean. 1992. An expert system for raising pigs. In The first Conference on Practical Applications of Prolog
1992
-
[51]
Tim Menzies, Alex Dekhtyar, Justin Distefano, and Jeremy Greenwald. 2007. Problems with Precision. IEEE Transactions on Software Engineering (September 2007). http://menzies.us/pdf/07precision.pdf
2007
-
[52]
Tim Menzies, Oussama Elrawas, Jairus Hihn, Martin Feather, Ray Madachy, and Barry Boehm. 2007. The business case for automated software engineering. In Proceedings of the twenty-second IEEE/ACM international conference on Automated software engineering . ACM, 303–312
2007
-
[53]
Tim Menzies, Steve Williams, Oussama El-Rawas, Barry Boehm, and Jairus Hihn. 2009. How to avoid drastic software process change (using stochastic stability). In 2009 IEEE 31st International Conference on Software Engineering . IEEE, 540–550
2009
-
[54]
Wiem Mkaouer, Marouane Kessentini, Adnan Shaout, Patrice Koligheu, Slim Bechikh, Kalyanmoy Deb, and Ali Ouni. 2015. Many-objective software remodularization using NSGA-III. ACM Transactions on Software Engineering and Methodology (TOSEM) 24, 3 (2015), 1–45
2015
-
[55]
Audris Mockus and Lawrence G Votta. 2000. Identifying Reasons for Software Changes using Historic Databases.. In icsm. 120–130
2000
-
[56]
Vivek Nair, Tim Menzies, and Jianfeng Chen. 2016. An (accidental) exploration of alternatives to evolutionary algorithms for sbse. In International Symposium on Search Based Software Engineering . Springer, 96–111
2016
-
[57]
Vivek Nair, Zhe Yu, and Tim Menzies. 2017. FLASH: A Faster Optimizer for SBSE Tasks. arXiv preprint arXiv:1705.05018, under review (2017)
2017 arXiv
-
[58]
Vivek Nair, Zhe Yu, Tim Menzies, Norbert Siegmund, and Sven Apel. 2018. Finding faster configurations using flash. arXiv preprint arXiv:1801.02175 (2018)
2018 arXiv
-
[59]
Jaechang Nam and Sunghun Kim. 2015. CLAMI: Defect Prediction on Unlabeled Datasets. In Proceedings of the 30th IEEE/ACM International Conference on Automated Software Engineering (ASE 2015)
2015
-
[60]
Yoshihiko Ozaki, Yuki Tanigaki, Shuhei Watanabe, and Masaki Onishi. 2020. Multiobjective tree-structured parzen estimator for computationally expensive optimization problems. In Proceedings of the 2020 genetic and evolutionary computation conference . 533–541. Manuscript submi...
2020
-
[61]
Dan Port, Alexy Olkov, and Tim Menzies. 2008. Using simulation to investigate requirements prioritization strategies. In 2008 23rd IEEE/ACM International Conference on Automated Software Engineering . IEEE, 268–277
2008
-
[62]
Christopher Schröder, Andreas Niekler, and Martin Potthast. 2021. Revisiting uncertainty-based query strategies for active learning with transformers. arXiv preprint arXiv:2107.05687 (2021)
2021 arXiv
-
[63]
Andrew Jhon Scott and Martin Knott. 1974. A cluster analysis method for grouping means in the analysis of variance. Biometrics (1974), 507–512
1974
-
[64]
Md Kamrul Siam, Huanying Gu, and Jerry Q Cheng. 2024. Programming with AI: Evaluating ChatGPT, Gemini, AlphaCode, and GitHub Copilot for Programmers. arXiv preprint arXiv:2411.09224 (2024)
2024 arXiv
-
[65]
Aditya Siddhant and Zachary C Lipton. 2018. Deep bayesian active learning for natural language processing: Results of a large-scale empirical study. arXiv preprint arXiv:1808.05697 (2018)
2018 arXiv
-
[66]
Avi Singh, John D Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J Liu, James Harrison, Jaehoon Lee, Kelvin Xu, et al. 2023. Beyond human data: Scaling self-training for problem-solving with language models. arXiv preprint arXiv:2312.06585 (2023)
2023 arXiv
-
[67]
Yuan Sui, Mengyu Zhou, Mingjie Zhou, Shi Han, and Dongmei Zhang. 2024. Table meets llm: Can large language models understand structured table data? a benchmark and empirical study. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining . 645–654
2024
-
[68]
Lilia Tang, Chaitanya Bhandari, Yongle Zhang, Anna Karanika, Shuyang Ji, Indranil Gupta, and Tianyin Xu. 2023. Fail through the Cracks: Cross-System Interaction Failures in Modern Cloud Systems. In Proceedings of the Eighteenth European Conference on Computer Systems (Rome, It...
2023
-
[69]
Vali Tawosi, Salwa Alamir, and Xiaomo Liu. 2023. Search-Based Optimisation of LLM Learning Shots for Story Point Estimation. In International Symposium on Search Based Software Engineering . Springer, 123–129
2023
-
[70]
Vali Tawosi, Rebecca Moussa, and Federica Sarro. 2023. Agile Effort Estimation: Have We Solved the Problem Yet? Insights From a Replication Study. IEEE TSE 49, 4 (2023), 2677–2697
2023
-
[71]
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al
-
[72]
Ricardo Valerdi. 2010. Heuristics for systems engineering cost estimation. IEEE Systems Journal 5, 1 (2010), 91–98
2010
-
[73]
B Vasilescu. 2018. Personnel communication at fse’18. Found. Softw. Eng (2018)
2018
-
[74]
Bogdan Vasilescu, Yue Yu, Huaimin Wang, Premkumar Devanbu, and Vladimir Filkov. 2015. Quality and productivity outcomes relating to continuous integration in GitHub. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering . ACM, 805–816
2015
-
[75]
Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, and Marius Hobbhahn. [n. d.]. Position: Will we run out of data? Limits of LLM scaling based on human-generated data. In Forty-first International Conference on Machine Learning
-
[76]
Rui Wang, Wubin Ma, Mao Tan, Guohua Wu, Ling Wang, Dunwei Gong, and Jian Xiong. 2021. Preference-inspired coevolutionary algorithm with active diversity strategy for multi-objective multi-modal optimization. Information Sciences 546 (2021), 1148–1165
2021
-
[77]
Shuhei Watanabe, Noor Awad, Masaki Onishi, and Frank Hutter. 2022. Speeding up multi-objective hyperparameter optimization by task similarity- based meta-learning for the tree-structured parzen estimator. arXiv preprint arXiv:2212.06751 (2022)
2022 arXiv
-
[78]
Cody Watson, Nathan Cooper, David Nader Palacio, Kevin Moran, and Denys Poshyvanyk. 2022. A Systematic Literature Review on the Use of Deep Learning in Software Engineering Research.ACM Trans. Softw. Eng. Methodol.31, 2, Article 32 (March 2022), 58 pages. https://doi.org/10.11...
2022 doi
-
[79]
Jiarong Wei, Yancong Lin, and Holger Caesar. 2024. BaSAL: Size-Balanced Warm Start Active Learning for LiDAR Semantic Segmentation. In2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 18258–18264
2024
-
[80]
Christopher Williams and Carl Rasmussen. 1995. Gaussian processes for regression. Advances in neural information processing systems 8 (1995)
1995
-
[81]
Xiaoxue Wu, Wei Zheng, Xin Xia, and David Lo. 2022. Data Quality Matters: A Case Study on Data Label Correctness for Security Bug Report Prediction. IEEE Transactions on Software Engineering 48, 7 (2022), 2541–2556. https://doi.org/10.1109/TSE.2021.3063727
2022
-
[82]
Rahul Yedida, Hong Jin Kang, Huy Tu, Xueqi Yang, David Lo, and Tim Menzies. 2023. How to find actionable static analysis warnings: A case study with FindBugs. IEEE Transactions on Software Engineering 49, 4 (2023), 2856–2872
2023
-
[83]
Ofer Yehuda, Avihu Dekel, Guy Hacohen, and Daphna Weinshall. 2022. Active learning through a covering lens. Advances in Neural Information Processing Systems 35 (2022), 22354–22367
2022
-
[84]
Zhe Yu, Fahmid Morshed Fahid, Huy Tu, and Tim Menzies. 2022. Identifying self-admitted technical debts with jitterbug: A two-step approach. IEEE Transactions on Software Engineering 48, 5 (2022), 1676–1691
2022
-
[85]
Michelle Yuan, Hsuan-Tien Lin, and Jordan Boyd-Graber. 2020. Cold-start active learning through self-supervised language modeling. arXiv preprint arXiv:2010.09535 (2020)
2020 arXiv
-
[86]
Guofu Zhang, Zhaopin Su, Miqing Li, Feng Yue, Jianguo Jiang, and Xin Yao. 2017. Constraint handling in NSGA-II for solving optimal testing resource allocation problems. IEEE Transactions on Reliability 66, 4 (2017), 1193–1212
2017
-
[87]
Qingfu Zhang and Hui Li. 2007. MOEA/D: A multiobjective evolutionary algorithm based on decomposition. IEEE Transactions on evolutionary computation 11, 6 (2007), 712–731. Manuscript submitted to ACM
2007
-
[2024]
arXiv preprint arXiv:2403.05530 (2024)
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530 (2024)
2024 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.