How to Design a Machine Learning Experiment

Fix the dataset, the split, the baseline, the metric, and the number of runs before you train, so the result is something a reviewer can check.

By Dr Raktim Mondol · 29 September 2026 · 3 min read

A machine-learning experiment is a comparison you could lose. If every choice is still free when the first accuracy number appears, you will keep changing the setup until the number looks right. Write the protocol first. Then run it. The training code is the easy part to redo. A contaminated comparison is not.

Name the dataset and the split in the protocol

Write the dataset’s name, version, licence, and the exact split. If the community already published a split, use it, unless your question is about the split itself. If you make a split, write the rule: random, by patient, by time, by site, or by speaker. A random split of rows that belong to the same person is a different, weaker claim than a split that keeps a person entirely on one side. Say which one you did.

  • Train, validation, and test are three jobs. Validation is for choices you are still allowed to change. Test is for the number you will report. Using the test set to pick a checkpoint spends it.
  • Write the unit of splitting. Image, patient, document, user, or time window. The unit is part of the claim.
  • Record the seed and the sizes. Another person should be able to rebuild the same sets, or you should ship the id lists.

Pick the baseline a sceptical reader expects

The baseline is the thing your reader already believes. Often that is the strongest published method on the same split, a simple model, and an obvious non-learning rule where one exists. Reimplement or rerun it when you can. Citing a number from a paper that used a different split, a different metric, or extra data is not a baseline. It is a different experiment.

One primary metric, chosen now

Choose the metric that matches the decision in the question. Accuracy on a rare event can hide a model that never finds the event. A leaderboard metric can be the wrong one for the cost of a mistake. Write the primary metric in the protocol, and list any secondary metrics as secondary. Reporting five metrics and discussing the one that moved is metric shopping. How that choice interacts with leakage and calibration is in How to evaluate a machine learning model.

Decide the runs before you see them

One seed is an anecdote. Write how many independent runs you will do, what you will report besides the mean, and which difference you would treat as worth a claim. If the runs are expensive, a smaller number with a pre-written reason is better than a single dramatic seed. The test you will use, if you use one, should be chosen here too. t-test versus bootstrap is the companion for that line of the protocol.

Plan the ablations as questions, not as a grid

Each ablation should remove or swap one piece you intend to take credit for. An exhaustive grid of every hyperparameter is a search, not an explanation, and it spends the test set if you pick the winner on it. Write the few comparisons that, if they failed, would change the sentence in your abstract.

Protocol lineWrite this before trainingLeave it blank and this happens
DataName, version, split unit, id lists or seed.The split quietly includes the test cases.
BaselineWhich system, whose code, which split.You compare against a number you cannot reconstruct.
MetricOne primary metric and why.The write-up features whichever number moved.
RunsHow many seeds, what you will report.The best seed becomes “the result”.
Stop ruleThe date or budget where tuning ends.The project never reaches a paper.

Free download

Experiment Protocol & Ablation Planner

Fix your evaluation plan before you run experiments — the fastest way to avoid reviewer objections.

Want guidance, not just a guide?

Statistics & Experiment Design for ML Research

Design experiments and prove your results actually hold up.

3 Hours live online · Third Saturday of every month · $299 USD

View the course