REVIEW 6 cited by
Benchmarking in Optimization: Best Practice and Open Issues
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This survey compiles ideas and recommendations from more than a dozen researchers with different backgrounds and from different institutes around the world. Promoting best practice in benchmarking is its main goal. The article discusses eight essential topics in benchmarking: clearly stated goals, well-specified problems, suitable algorithms, adequate performance measures, thoughtful analysis, effective and efficient designs, comprehensible presentations, and guaranteed reproducibility. The final goal is to provide well-accepted guidelines (rules) that might be useful for authors and reviewers. As benchmarking in optimization is an active and evolving field of research this manuscript is meant to co-evolve over time by means of periodic updates.
Forward citations
Cited by 6 Pith papers
-
Large-scale benchmarking of multi-objective soft-computing metaheuristics for redundancy allocation in repairable k-out-of-n systems
A 65-algorithm benchmark on a bi-objective repairable redundancy-allocation problem shows that algorithm rankings are budget-dependent and that Scaled Binomial Initialization changes relative performance.
-
Adaptive Estimation of the Number of Algorithm Runs in Stochastic Optimization
A skewness-based stopping rule estimates the required number of algorithm runs online and, in large COCO benchmarks, achieves 82 to 95 percent estimation accuracy while saving roughly 50 percent of runs, with a 5 to 2...
-
From Heuristic Selection to Automated Algorithm Design: LLMs Benefit from Strong Priors
Prompting LLMs with strong benchmark algorithm code, rather than relying on linguistic instructions, improves LLM-driven black-box optimization; the proposed BAG method outperforms five baselines on pbo and bbob.
-
Evolutionary Computation and Large Language Models: A Survey of Methods, Synergies, and Applications
A survey that maps bidirectional synergies between evolutionary computation and large language models and proposes a taxonomy plus research gaps.
-
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
A meta-review of about 110 critical studies finds nine systemic weaknesses in AI benchmarking and concludes that benchmarks are receiving disproportionate trust in AI governance.
-
Time-Fair Benchmarking for Metaheuristics: A Restart-Fair Protocol for Fixed-Time Comparisons
A fixed-time, restart-fair benchmarking protocol for metaheuristics, using ERT and time-based performance profiles, is proposed without empirical validation.
Discussion (0). Continue with ORCID to comment.