How to Evaluate a Machine Learning Model

Choose the split, the metric, and the baseline before you report a number. A single seed, a leaked test set, or the wrong metric will not survive a careful reader.

By Dr Raktim Mondol · 29 September 2026 · 3 min read

A model is not evaluated by the number it produced on the data it was tuned on. Evaluation is the argument that the number means what you say it means, on data the model was not allowed to learn from, against a comparison a reader accepts. If that argument is missing, a high score is a debugging output.

Keep a test set you have not spent

Every time you look at the test number and change the model, the features, or the threshold, you use the test set as a training signal. Do that often enough and the number describes your search, not the model’s behaviour on new data. Tune on the validation set. Touch the test set when the protocol in How to design a machine learning experiment says the comparison is finished, and report that run.

  • Leakage is not only a shuffled row. It is a feature computed with the test labels, a patient’s later visit in the training set, a duplicate copied into both sides, or preprocessing fit on the full table before the split.
  • The split unit has to match the claim. If you claim the model works on a new hospital, a new user, or a later year, the test set has to be a new hospital, a new user, or a later year.
  • A public benchmark with a public test set has often been read by the whole field. Treat a gain on it as a comparison with published numbers, and say so, rather than as a surprise held-out trial.

Pick a metric that can show the failure you care about

Accuracy is a poor summary when one class is rare. A single threshold hides the tradeoff a user would actually face. A loss that fell during training is not an evaluation of the decision. Write the primary metric because of the decision, then add the plot or the second number that would reveal a useless model with a flattering average. If you report area under a curve, also say what happens at the operating point someone would use.

SituationA metric that can misleadWhat to show beside it
One class is rareAccuracyPrecision and recall, or the rate of the missed cases, at a stated threshold.
The user faces a tradeoffA single F-scoreThe curve, and the point you recommend, with the costs you assumed.
Scores are used as probabilitiesAccuracy or AUROC aloneA calibration check. A sharp ranking can still be a badly scaled probability.
The claim is about a new site or a later yearA random split of the pooled dataThe number on the held-out site or year, even if it is worse.

Compare against a baseline, and show the spread

A number without a baseline has no meaning to a reviewer who was not in the room. Report the baseline on the same split and the same metric. Report more than one run when the training is stochastic, and show the spread, not only the best seed. A mean that moved by less than the spread is not a finding, even if the arrow in the table points up. Whether you then want a formal test is a separate decision, covered in How to choose a significance test for ML results.

Look at the errors before you write the claim

Open a sample of the failures. If they share a cause your average hides — a site, a language, a sensor, a demographic group, a long document — the average is the wrong sentence for the abstract. Either measure that slice and narrow the claim, or say you did not measure it. “The model works” is not a result. “On this split, this metric moved by this much against this baseline, and it failed on this slice” is a result.

Want guidance, not just a guide?

Statistics & Experiment Design for ML Research

Design experiments and prove your results actually hold up.

3 Hours live online · Third Saturday of every month · $299 USD

View the course