t-test vs Bootstrap: How to Choose a Significance Test for ML Results

Which statistical test fits which machine-learning comparison: two models on one test set, several seeds, or many datasets. Includes a working paired-bootstrap example.

By Dr Raktim Mondol · 29 September 2026 · 3 min read

“Our method is better” is a claim about the world, not just about one run. Reviewers increasingly ask whether an improvement is real or noise. The right test depends on what you are comparing and where the randomness comes from, so start there, not with a favourite test.

First: where does the randomness come from?

There are two different sources of variation in an ML result, and they need different tools:

  • Test-set sampling. Your test set is a sample from a larger population. A different sample would give slightly different scores. This affects every result, even for a deterministic model.
  • Training randomness. Different random seeds, initialisations, and data orders give different trained models, even on the same data.

Running several seeds captures the second source but says nothing about the first. Resampling the test set (a bootstrap) captures the first but not the second. A careful comparison thinks about both.

Which test for which situation

SituationA sensible choiceNotes
Two classifiers, same test set, one training run eachMcNemar's test on the paired right/wrong outcomesDietterich (1998) recommends it when each model can only be trained once. It uses only the examples where the two models disagree.
Two models, same test set, a metric such as F1, AUC, BLEU, or accuracyPaired bootstrap over test examplesWorks for almost any metric. Report the confidence interval of the difference.
Two methods, several random seedsReport mean ± standard deviation (and ideally a confidence interval)A paired t-test across seeds has little power with 3 to 5 seeds, and only reflects training randomness. Do not over-interpret a p-value from three runs.
Two methods across many datasetsWilcoxon signed-rank test over per-dataset scoresDemšar (2006) recommends non-parametric tests for comparing classifiers over multiple datasets.
Several methods across many datasetsFriedman test with a post-hoc test (for example Nemenyi)Also from Demšar (2006). It tests ranks, not raw score differences.

Why a plain t-test is often the wrong tool

A t-test assumes that the numbers you feed it are independent and roughly normally distributed. Per-example scores from a test set are not normal (they are often just zeros and ones). Scores across seeds are few. Scores across datasets are not comparable in scale. That does not make the t-test forbidden, but it is a poor default. Choose the test whose assumptions match your data.

A paired bootstrap you can copy

The idea: resample the test examples with replacement many times, recompute the metric for both systems on each resample, and look at the spread of the difference. Both systems are evaluated on the same resampled examples each time, which is what makes it “paired”.

import numpy as np

def paired_bootstrap(metric, y_true, pred_a, pred_b, n_boot=10_000, seed=0):
    """Compare systems A and B on the same test set.

    metric(y_true, y_pred) -> float. Arrays must be NumPy arrays of equal length.
    Returns the observed difference (A - B) and a 95% confidence interval.
    """
    rng = np.random.default_rng(seed)
    n = len(y_true)
    observed = metric(y_true, pred_a) - metric(y_true, pred_b)
    diffs = np.empty(n_boot)
    for i in range(n_boot):
        idx = rng.integers(0, n, n)          # resample test examples with replacement
        diffs[i] = metric(y_true[idx], pred_a[idx]) - metric(y_true[idx], pred_b[idx])
    low, high = np.percentile(diffs, [2.5, 97.5])
    return observed, (low, high)
If the 95% interval excludes zero, the improvement is unlikely to be explained by test-set sampling alone.

This assumes the test examples are independent. If they are not (several samples per patient, or per user), resample at the level of the group, or the interval will be too narrow.

Habits that matter more than the test

  1. Decide the test before you run the experiments, and name your primary metric in advance. Choosing a test or metric after seeing the results is a form of cherry-picking.
  2. Run enough seeds. Three is the bare minimum and five to ten is much better when you can afford it. Report all of them, not the best one.
  3. Give baselines the same tuning budget as your method. An unfair comparison is not fixed by any test.
  4. Correct for multiple comparisons. If you test many pairs, the chance of a false positive grows. Adjust (for example with the Holm-Bonferroni method) or state which comparisons were planned.
  5. Report effect sizes and intervals, not just p-values. “Significant” with a tiny effect can be irrelevant, and a large effect with a wide interval can be inconclusive.
  6. Keep the test set untouched until the end. Tune on validation data only.

References

  • Dietterich, T. G. (1998). Approximate statistical tests for comparing supervised classification learning algorithms. Neural Computation, 10(7), 1895-1923.
  • Demšar, J. (2006). Statistical comparisons of classifiers over multiple data sets. Journal of Machine Learning Research, 7, 1-30.

Free download

Experiment Protocol & Ablation Planner

Fix your evaluation plan before you run experiments — the fastest way to avoid reviewer objections.

Want guidance, not just a guide?

Statistics & Experiment Design for ML Research

Design experiments and prove your results actually hold up.

3 Hours live online · Third Saturday of every month · $299 USD

View the course