How to Conduct an AI Research Project

A working order for an AI research project: question, data, baseline, experiment, evaluation, and the paper. Skip a step and the GPU time usually does not come back.

By Dr Raktim Mondol · 29 September 2026 · 3 min read

An AI research project is a research project whose evidence happens to be a model. The order that holds up is the same as any empirical project: decide what would count as an answer, make sure the data can support that answer, compare against something a reader already understands, then write only the claim the comparison supports. Training a model first and looking for a question afterwards produces a demo. Demos are useful. They are a weak paper.

Freeze the question before you freeze the architecture

Write the question with a dataset, a comparison, and a metric, using How to formulate a research question. “Try a transformer” is a plan for a weekend, not a project. If the interesting part is a new architecture, the question is still about a task: on which data, against which baseline, under which constraint of data, compute, or latency. Architecture is the intervention. It is not the question.

Spend the first week on data and the prior work

  1. Find the closest papers, not the most famous ones. You need the ones that ran your comparison, or nearly did. How to find a research gap is the check.
  2. Open the data before you promise a result. Licence, size, label quality, what a split already published looks like, and whether you are allowed to train on it.
  3. Write down what you will not use. Test data that leaked into training, a benchmark whose labels you do not trust, a scrape you cannot describe.

Put a baseline in the plan

A reader believes a gain only against a comparison they understand. That might be the published number on the same split, a simpler model you reimplement, or the current system a practitioner uses. If you cannot run the baseline yourself, say whose number you are citing and on which split. A project with only the new model has nothing to put in a results table. The design of that comparison is How to design a machine learning experiment.

Run a small version before the full one

A subset of the data, one seed, and a cheap model will tell you whether the pipeline drops labels, leaks the test set, or measures the wrong thing. Fix that before you spend the GPU budget. Log the command, the commit, the data version, and the metric in one place. A result you cannot rerun from those four things is a screenshot.

Evaluate the claim you wrote, not the claim the plot suggests

When the runs finish, go back to the sentence you froze. Report that comparison, including runs that did not help. A higher average on one metric is not a new claim about fairness, clinical use, or “understanding” unless you measured those. How to choose the split and the metric is How to evaluate a machine learning model. The failures that show up at this stage are listed in Common mistakes in AI research.

Write the paper from the comparison, not from the model

The introduction’s contribution bullets should be sentences your results section already supports. Related work should contain the baselines you compared against, not a tour of the field. The method chapter describes what you ran. If a detail is needed to rerun the result, it is not optional colour. The section order is in How to structure a research paper.

Want guidance, not just a guide?

Pathway to Research & PhD

The whole pipeline — literature review to PhD offer — in one session.

3 Hours live online · First Sunday of every month · $490 USD

View the course