Pith. sign in

REVIEW 6 cited by

Benchmarking in Optimization: Best Practice and Open Issues

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.03488 v2 pith:IGHFFH2U submitted 2020-07-07 cs.NE cs.PFmath.OCstat.AP

classification cs.NEcs.PFmath.OCstat.AP
keywords benchmarkingbestdifferentgoaloptimizationpracticeactiveadequate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This survey compiles ideas and recommendations from more than a dozen researchers with different backgrounds and from different institutes around the world. Promoting best practice in benchmarking is its main goal. The article discusses eight essential topics in benchmarking: clearly stated goals, well-specified problems, suitable algorithms, adequate performance measures, thoughtful analysis, effective and efficient designs, comprehensible presentations, and guaranteed reproducibility. The final goal is to provide well-accepted guidelines (rules) that might be useful for authors and reviewers. As benchmarking in optimization is an active and evolving field of research this manuscript is meant to co-evolve over time by means of periodic updates.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 81 citations worldwide. Full citation record

  1. Large-scale benchmarking of multi-objective soft-computing metaheuristics for redundancy allocation in repairable k-out-of-n systems

    cs.NE 2025-12 conditional novelty 6.0 of 10

    A 65-algorithm benchmark on a bi-objective repairable redundancy-allocation problem shows that algorithm rankings are budget-dependent and that Scaled Binomial Initialization changes relative performance.

  2. Adaptive Estimation of the Number of Algorithm Runs in Stochastic Optimization

    cs.NE 2025-07 conditional novelty 5.0 of 10

    A skewness-based stopping rule estimates the required number of algorithm runs online and, in large COCO benchmarks, achieves 82 to 95 percent estimation accuracy while saving roughly 50 percent of runs, with a 5 to 2...

  3. From Heuristic Selection to Automated Algorithm Design: LLMs Benefit from Strong Priors

    cs.LG 2026-03 conditional novelty 4.0 of 10

    Prompting LLMs with strong benchmark algorithm code, rather than relying on linguistic instructions, improves LLM-driven black-box optimization; the proposed BAG method outperforms five baselines on pbo and bbob.

  4. Evolutionary Computation and Large Language Models: A Survey of Methods, Synergies, and Applications

    cs.NE 2025-05 conditional novelty 4.0 of 10

    A survey that maps bidirectional synergies between evolutionary computation and large language models and proposes a taxonomy plus research gaps.

  5. Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

    cs.AI 2025-02 conditional novelty 4.0 of 10

    A meta-review of about 110 critical studies finds nine systemic weaknesses in AI benchmarking and concludes that benchmarks are receiving disproportionate trust in AI governance.

  6. Time-Fair Benchmarking for Metaheuristics: A Restart-Fair Protocol for Fixed-Time Comparisons

    cs.NE 2025-09 conditional novelty 3.0 of 10

    A fixed-time, restart-fair benchmarking protocol for metaheuristics, using ERT and time-based performance profiles, is proposed without empirical validation.

Pith tools