Insights · 5 min read
Most AI projects stall between the demo and the day a real user touches the feature. How to plan an AI build around data, integration, and evaluation so it ships.
By GGP Editorial
The demo is the easy part. You wire a model to a form, it gives a decent answer, and the room is impressed. The hard part starts the next week, when that same feature has to handle real data, real users, and a budget that actually matters.
We have shipped enough AI features to know where these projects stall. It is rarely the model. It is almost always the plumbing around it: getting clean data in, getting the output somewhere useful, and being honest about what the thing can and can't do. The team that understands this gap is the one that gets you to a working product.
A working AI feature has five parts, and only one of them is the model itself.
Data comes first. If the source data is messy, no amount of tuning fixes it. One client wanted a document-processing tool and assumed the model was the hard part. The real work was normalising ten years of inconsistently named PDFs. That step took longer than everything else combined.
Then integration. The model's answer has to land inside an app, a dashboard, or a workflow someone already uses. A chatbot that sits in a separate tab gets opened once and forgotten.
Then evaluation. You need a way to tell whether the output is getting better or worse after each change. A small test set you can score by hand saves you from shipping a fix that made everything worse.
Then cost control. Every call to a hosted model bills you something. A feature that is cheap with ten users can get expensive with ten thousand. Most teams only notice this after the invoice arrives.
Last, guardrails. What should the system do when the model is confidently wrong? That is a product decision, not a technical one, and somebody has to make it before launch.
Most projects should start with a hosted API. It gets you most of the value for a fraction of the effort, and you can swap models later without rewriting the feature.
Fine-tuning earns its keep when you have a specific domain vocabulary or a tone the base model does not match, and when you have enough labelled examples to justify it. Without those examples, fine-tuning is just expensive guessing.
Training a model from scratch is almost never the right call for a business feature. It is a research budget with a research timeline, and it rarely beats a well-integrated API for a product that needs to ship this quarter.
Pick one narrow task and ship it. The temptation is to build an assistant that does everything, and that is how projects end up six months in with nothing live. A single feature that saves someone ten minutes a day is worth more than an ambitious demo that never finishes.
Once that one task works, you have real usage data, you have user feedback, and you know whether to keep going or change direction. The second feature gets easier because the first one paid for the plumbing.
Clients are usually surprised that training the model is a small slice of the total cost. Here is a rough split we see across projects:
| Work area | Share of effort | Notes |
|---|---|---|
| Data cleaning and preparation | 30-40% | The unglamorous part nobody budgets for |
| Integration and UI | 25-30% | Getting the output into the real product |
| Model selection and tuning | 15-20% | Usually an API call, not a training run |
| Evaluation and testing | 10-15% | Scoring outputs, edge cases, regressions |
| Monitoring and maintenance | 5-10% | Ongoing, and easy to forget |
The numbers move from project to project, but the pattern holds: data and integration eat the time. Budget for that and your timeline stops looking naive.
AI work is iterative, so a team that overlaps your working hours matters more than it does on a fixed-scope build. We are a China-based team of 40-plus engineers, founded in 2018, and we schedule overlapping hours with clients in the US, Singapore, and Brazil. Questions get answered in the same conversation instead of the next morning. We work in English and Portuguese, and each client gets a dedicated group so nothing gets lost between time zones.
We have built for those markets directly, not just served them. A client in South Africa needed a recommendation feature that respected local data rules. A Singapore fintech wanted a risk-scoring module that had to run fast enough for their compliance review. Different constraints, same lesson: the AI part was quick, the shipping part was the project.
If you are stuck at the prototype stage, the problem is probably not your model. It is the plan around it. Map the data, the integration, and the evaluation first, and the model becomes the smallest decision you have to make.
Tell us what you are building and where you are today. We typically reply within 24 hours.