Insights · 5 min read

Machine learning development: start with the data

Most machine learning development projects fail before any model is trained. The data is not there. Here is how to scope a first ML project that survives.

By GGP Editorial

Machine learning development has an odd reputation. People either expect magic or expect it to be a scam, and both views come from the same place: projects that started with a model and skipped the data. In practice, most ML work fails before a single model is trained, because the data is not there or the target is not measurable. This guide is about the other 90 percent of the work: the part before and after the model, where the project is actually won or lost.

Is your problem even an ML problem

The first thing we ask a client is whether they need machine learning at all. A surprising number of problems people call ML are really just rules. If you can write down the logic, write it down. A rule is easier to build, test, and explain than a model, and it will be right more often on day one.

Machine learning earns its keep when the rules are too many to write or too hard to express: predicting churn from usage patterns, matching listings to buyers, spotting fraud in a stream of transactions. Those are pattern problems, not rule problems. If your problem is a pattern problem and you have historical examples of the outcome, you are in ML territory.

The three things that decide success

Every ML project stands or falls on three things, and none of them is the model:

FactorWhat it meansWhat kills it
Data qualityClean, labeled examples of the outcomeMissing labels, biased samples, tiny datasets
A measurable targetA number you are trying to moveA vague goal like "smarter" with no metric
A deployment pathA way to put the model in front of usersA model that never leaves a notebook

The measurable target is the one people skip. "Make recommendations better" is not a target. "Raise click-through from 2 percent to 3 percent on the recommendation feed" is. If you cannot write the second sentence, do not start the project, because you will not know whether it worked.

A related trap is judging a model by training accuracy. A model can score well on the data it was trained on and still be useless on new data. You need a held-out set it has never seen, and ideally a live test on a small slice of traffic. The number that matters is how it does in production, not in the notebook.

What ML development actually costs

Most of the budget goes to data, not to the model. Cleaning records, fixing labels, and building the pipeline that feeds the model often takes three or four times the hours spent training it. The labeling itself is often the bottleneck: somebody has to go through historical records and mark which ones are examples of the outcome you care about, and that work is slow and usually done by people, not by a clever script. Budget real weeks for it, because rushed labels poison everything downstream.

Then there is the part nobody plans for: after launch, the model degrades as the world changes, and someone has to watch it and retrain it. That ongoing work is the real cost of ML. A model is not a one-time build the way a login page is. It is a living part of the system, and it needs a plan for monitoring and refresh. Clients who treat ML as a one-off are usually surprised when accuracy drops six months later and no one knows why.

A sane way to start

Start small and measurable. Pick one prediction that has historical data behind it, define the metric, and ship a first version even if it is behind a feature flag. Ship early so you learn whether the model behaves in production, not just in a notebook.

Pick the prediction you would actually act on. A model that flags which customers are about to churn is worth building only if your team will call or email those customers. A prediction nobody acts on is a science project, and science projects get cancelled.

GlobeSoft builds AI and ML features inside larger systems rather than as science projects: a fraud check inside a payment flow, a recommendation list inside an app, a forecast inside a reporting tool. We work across finance, IoT, and property management, with clients in Brazil, South Africa, Singapore, and the United States, and we run overlap hours and a dedicated group so the work does not stall on time zones. If you are unsure whether your idea is a rules problem or a pattern problem, start with that conversation. It is the cheapest part of the project and the one that saves the most money.

Talk to us about your project

Need help applying this?

Tell us what you are building and where you are today. We typically reply within 24 hours.