Free tool · runs in your browser

Significance Test Chooser for ML Results

Comparing two models on one test set, several seeds, or many datasets? This walks you to a sensible test, with the Python call to run it and the pitfalls to avoid. Based on Dietterich (1998) and Demšar (2006).

QUESTION 1 OF 2

What are you comparing?