ATLASResearch
methods
Quantitative/ Measurement and analysis

Classification metrics

Make classification errors visible and decision-relevant

A confusion matrix tabulates true and predicted classes at a chosen threshold. Precision is the proportion of positive predictions that are correct; recall is the proportion of actual positives found. F1 combines precision and recall harmonically but ignores true negatives. Multiclass micro, macro and weighted averaging answer different questions and must be named explicitly.

WHEN IT FITS

Choose these measures when evaluating classification decisions and the costs of missed cases and false alarms differ, with credible reference labels and an evaluation sample relevant to deployment prevalence.

Strengths

  • Exposes error types hidden by overall accuracy
  • Connects thresholds to operational false-alarm and missed-case trade-offs

Limitations

  • Precision depends on prevalence
  • F1 can hide calibration and asymmetric error costs

Know the boundary

An F1 score is not a probability of correctness, and a confusion matrix is not threshold-free.

USED ACROSS
Computer scienceBanking & financeEducationPsychology