Free tool · runs in your browser
Comparing two models on one test set, several seeds, or many datasets? This walks you to a sensible test, with the Python call to run it and the pitfalls to avoid. Based on Dietterich (1998) and Demšar (2006).
Read the guide
Which statistical test fits which machine-learning comparison: two models on one test set, several seeds, or many datasets. Includes a working paired-bootstrap example.
ReadGo deeper
Design experiments and prove your results actually hold up.
$299 · View the courseFree download
Fix your evaluation plan before you run experiments — the fastest way to avoid reviewer objections.