ATLASResearch
methods
Quantitative/ Evaluation design

Benchmarking

Compare systems under disclosed workloads

Benchmarking evaluates implementations against shared workloads, datasets and measurement protocols. It requires an appropriate baseline, controlled execution conditions and repeated observations at relevant sources of variability. Microbenchmarks isolate components, while application benchmarks measure interacting systems. A benchmark score has meaning through its workload coverage and protocol, rather than through the prestige of a leaderboard.

WHEN IT FITS

Use it to compare runtime, throughput, resource use or predictive performance when competing systems can run comparable tasks and the target workload can be specified and defended.

Strengths

  • Creates a common basis for comparison
  • Exposes trade-offs across workloads and resource budgets

Limitations

  • Benchmark-specific tuning can undermine generalisation
  • Hardware, warm-up and workload choices can dominate results

Know the boundary

A benchmark win does not establish superiority for workloads or hardware outside the evaluated conditions.

USED ACROSS
Computer science