Common Mistakes in AI Research
The mistakes that get an AI paper rejected, or published and then disbelieved: leakage, a weak baseline, one seed, metric shopping, and a claim the experiment never tested.
By Dr Raktim Mondol · 29 September 2026 · 3 min read
Most of these are not subtle statistical arguments. They are ordinary shortcuts that make a number look finished before the comparison is. A reviewer who runs experiments for a living has seen each one often enough to look for it in the first pass. Fix them in the protocol, not in the rebuttal.
The test set is used as a tuning knob
You try a change, look at the test score, and keep the change when the score rises. After enough tries, the test score is no longer an estimate of new data. It is the best of a search. Tune on validation. Report the test number from the protocol you wrote in How to design a machine learning experiment. If you already spent the test set, say that you did, and treat the number as exploratory.
The split is easier than the claim
Random rows, when the real question is a new patient, a new site, a new year, or a new user, train the model on near-copies of the test cases. The score is then a score on neighbours, not on the setting named in the introduction. Split on the unit you claim to generalise to. How to evaluate a machine learning model goes through the cases.
The baseline is missing, tired, or unfair
- No baseline. A table with only your model cannot show that the work did anything.
- A baseline you under-tuned while your own model got a long search. Give the comparison a fair budget, or state the unequal budget as a limitation.
- A number from a paper with a different split or extra data. That is not the same race. Rerun it, or do not call it a win.
- A baseline that is not what practitioners use, when your claim is that practitioners should switch.
One seed, then the best seed
Training is noisy. Reporting the seed that looked best, or reporting one seed as if it were the method, overstates the result. Run the number of seeds you wrote down. Show the spread. If the spread swallows the gain, the honest sentence is that you did not establish a gain. A formal comparison, when you need one, is in How to choose a significance test for ML results.
The metric is chosen after the runs
Five metrics in a table and a discussion of the one that moved is not a primary result. Choose the metric because of the decision, before the runs. Secondary metrics can explain a failure. They should not replace a primary metric that did not move.
The claim is larger than the experiment
A gain on a benchmark is a gain on that benchmark. It is not evidence that the model understands a domain, that it is safe to deploy, that it is fair across groups you did not measure, or that it will survive a new hospital. Write the claim so that a reader can point at a row in the results and see it. Cut every adjective that the row does not pay for.
Generated text is treated as a source
A writing tool can draft a related-work paragraph that cites papers it did not read, or invent a paper that fits the sentence. It can also state a number that was never in your logs. Every citation in the submitted paper has to be a paper you opened, and every number has to come from a run you can point to. The checklist linked below is that pass: citations, claims, and figures, one line at a time. A tools workshop will not replace it. The pass still has to happen before you upload.
Free download
AI-Assisted Research Verification Checklist
Use AI tools in your research without hallucinated citations, wrong numbers, or policy violations.
Want guidance, not just a guide?
Statistics & Experiment Design for ML Research
Design experiments and prove your results actually hold up.
3 Hours live online · Third Saturday of every month · $299 USD
View the course