EXPObench: Benchmarking Surrogate-based Optimisation Algorithms on Expensive Black-box Functions

Laurens Bliek; Arthur Guijt; Rickard Karlsson; Sicco Verwer; Mathijs de Weerdt

EXPObench: Benchmarking Surrogate-based Optimisation Algorithms on Expensive Black-box Functions

Laurens Bliek, Arthur Guijt, Rickard Karlsson, Sicco Verwer, Mathijs de Weerdt

TL;DR

EXPObench addresses the lack of standardised benchmarking for surrogate-based optimisation on expensive black-box functions by introducing a public benchmark library that evaluates six algorithms on four real-world problems (Windwake, Pitzdaily, ESP, HPO). It provides a coherent experimental framework, a public dataset of evaluation points and runtimes, and practical rules of thumb for algorithm selection. The study reveals that exploration and objective evaluation time often dominate algorithm performance, sometimes outweighing surrogate model accuracy, and highlights cross-domain effectiveness of discrete models on continuous problems. This work enables more uniform benchmarking, accelerates method development with a reusable data resource, and guides practitioners toward informed algorithm choices under different cost and budget constraints.

Abstract

Surrogate algorithms such as Bayesian optimisation are especially designed for black-box optimisation problems with expensive objectives, such as hyperparameter tuning or simulation-based optimisation. In the literature, these algorithms are usually evaluated with synthetic benchmarks which are well established but have no expensive objective, and only on one or two real-life applications which vary wildly between papers. There is a clear lack of standardisation when it comes to benchmarking surrogate algorithms on real-life, expensive, black-box objective functions. This makes it very difficult to draw conclusions on the effect of algorithmic contributions and to give substantial advice on which method to use when. A new benchmark library, EXPObench, provides first steps towards such a standardisation. The library is used to provide an extensive comparison of six different surrogate algorithms on four expensive optimisation problems from different real-life applications. This has led to new insights regarding the relative importance of exploration, the evaluation time of the objective, and the used model. We also provide rules of thumb for which surrogate algorithm to use in which situation. A further contribution is that we make the algorithms and benchmark problem instances publicly available, contributing to more uniform analysis of surrogate algorithms. Most importantly, we include the performance of the six algorithms on all evaluated problem instances. This results in a unique new dataset that lowers the bar for researching new methods as the number of expensive evaluations required for comparison is significantly reduced.

EXPObench: Benchmarking Surrogate-based Optimisation Algorithms on Expensive Black-box Functions

TL;DR

Abstract

EXPObench: Benchmarking Surrogate-based Optimisation Algorithms on Expensive Black-box Functions

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (2)