Insights · 9 min read

How to build an AI application

An AI application is a normal application with one part that makes a decision a person used to make. The model is usually the cheapest piece. The data, scope, and evaluation are what decide whether it ships.

By GGP Editorial

How to build an AI application

An AI application is a normal application with one unusual part: somewhere inside it, a model makes a decision that a person used to make. Everything around that part, the interface, the data, the integrations, the deployment, is ordinary software engineering. Founders get this backwards all the time. They fixate on the model and skip the parts that decide whether the product ships.

I build these for a living, and I have put AI into customer service, document processing, and back-office products for clients in Brazil, South Africa, Singapore, and the US. The model is now the cheapest and easiest piece, because the major labs sell it over an API and you do not need to train anything to get started. The hard parts are the data, the scope, and being honest about what good enough looks like.

Start with the job, not the model

Before anyone writes code, write down exactly what the application should do, in one sentence. Not "uses AI", the actual job. Classify incoming emails. Extract fields from invoices. Draft a first reply to a support ticket. Recommend a product. Summarise a document.

The verb matters because it tells you what kind of AI you need. Classification, extraction, generation, and recommendation are different jobs with different failure modes and different ways to measure success. If you cannot name the job in one sentence, you cannot build the thing, because a model cannot be aimed at a task nobody has defined.

The second thing to define is the failure mode: what a wrong answer costs and who catches it. A recommendation that is slightly off is cheap to live with. A financial figure that is wrong is not. Knowing the cost of a mistake decides how much human review you build in, and that decision drives the whole architecture.

The three ways to get the AI part

There are three ways to get intelligence into your application, and almost every business product should use the first one.

ApproachWhat it isWhen it fits
Hosted model APICall a model from OpenAI, Anthropic, or similarAlmost every business application
Fine-tune a modelTrain an existing model further on your dataWhen you need consistent tone or domain behaviour at scale
Train your own modelBuild and train a model from scratchResearch and niche cases, rarely a product launch

The API is the answer for nearly everyone starting out. You pay per use, you ship in days, and you can swap models later. Fine-tuning earns its keep when you have a specific, consistent behaviour you need repeated, like a style of writing or a narrow domain vocabulary, and you have the labelled examples to teach it. Training your own model is not how business applications launch; it is how a lab spends a year and a lot of money, and it is almost never the right call for a first product.

The architecture most apps actually use

Most real AI applications share the same shape, and it has more moving parts than people expect.

A retrieval step. If the model needs to answer from your documents or your data, you retrieve the relevant pieces and hand them to the model with the question. This is retrieval-augmented generation, and it is how you keep the model grounded in facts instead of letting it make things up.

An orchestration layer. The code that decides when to call the model, what to send it, what to do with the answer, and when to call it again. This is ordinary backend work in Java, Python, Go, or Node.

A human review loop. For most business applications you keep a person in the loop for the edge cases. That means a screen where someone approves or corrects the model's output, not a black box nobody can audit.

Evaluation. You need a way to measure whether the model is doing the job well, which means collecting examples where you know the right answer and scoring the model against them. Almost nobody budgets for this, and it is the difference between a product that improves and one that quietly drifts.

If that sounds like a lot of ordinary software around a small AI part, that is the point. The AI is the easy fifth.

Where the budget goes

The cost of an AI application is driven by the same things as any software project, plus a few AI-specific lines.

Integrations. Pulling data from your existing systems into the model and pushing the result back out is usually the biggest surprise, because every internal system has its own quirks.

Data quality. A model reflects the data you feed it. If your data is a mess, the model's output is a mess, and the cleanup work is real.

The human review loop. Building the screen where someone approves the model's output is engineering work people forget to price.

Model usage. This is priced per call and has been falling, so it is rarely the big line. Budget for it, but do not let it scare you off the project.

I will not quote a fixed price here, because the range is too wide to be honest and any number I name would be a guess. The pattern that holds across projects is that the integrations and the data cost more than the model. If you want the full picture of what AI development costs, we have written a longer breakdown.

Build it or buy it

A lot of an AI application is commodity, and you should buy the commodity rather than build it. Host the model through an API provider. Use an off-the-shelf embedding and retrieval service if one fits. Buy the parts that are not your product.

The part you build is where your judgement and your business live: the orchestration, the data pipeline, the review screen, and the specific behaviour that makes your product yours. This is the same shape as any software decision, and our article on adding AI to existing business software walks through how to buy the edges and build the core. Most teams should start there rather than with a big greenfield build.

The parts that fail

Three things go wrong more than anything else.

Dirty data. The model reflects what you feed it. Duplicated records and inconsistent formats produce output you cannot trust, and no model fixes that on its own. Clean the data before you point the model at it.

No evaluation. If you are not measuring whether the model is right, you are guessing, and the product drifts quietly. Build the scoring set early, before you tune anything.

Overbuilding. Teams ship a research project when the customer needed one well-defined job done. The fix is scope discipline: one job, a human in the loop, and a way to measure it. Our MVP article makes the same point for products generally, and AI products fail for the same reason.

How to start without overbuilding

Pick one job. Not ten. The one that is high-volume, expensive today, and specific enough that you can say when the model got it right.

Write down the job in a page, including examples of correct answers. If you cannot write the examples, you cannot evaluate the model, and you are not ready to build.

Prototype with a hosted API before you build any infrastructure. You can test whether the model does the job at all in a day, and that answer is worth more than a month of planning. If the prototype works, build the real thing around it. If it does not, you saved yourself a project.

Keep a human in the loop for the first version. This is not a weakness. It is how you gather the examples that later let you automate more of the loop.

How we build it

GlobeSoft builds AI applications on the same stack we use for everything else: Java and Spring Boot on the backend, with language-model integrations where the judgement lives, and Vue or React for the screens your team and customers use. We have shipped AI into customer service, document processing, and back-office products for clients across Brazil, South Africa, Singapore, and the US.

If you can name the job you want the model to do, send it to us. We will tell you whether it is an API job, a fine-tuning job, or something that does not need AI at all, and what a working first version would take to ship. That first read costs nothing and it is the cheapest way to avoid building the wrong thing.

Frequently asked questions

Do I need a data science team to build an AI application?

Usually not. The models are available over APIs now, so the work is software engineering: wiring the model into your systems, building the review loop, and measuring the output. You need engineers who understand the models, not a research lab.

Should I fine-tune a model or just use the API?

Start with the API. Fine-tune only when you have a specific, consistent behaviour you need repeated and enough labelled examples to teach it. Training your own model is almost never the right call for a first product.

How much does it cost to build an AI application?

It depends on the integrations, the state of your data, and how much you build versus buy. The integrations and data work usually cost more than the model itself. A first version around one well-defined job is a small project; a full product is a large one.

What is RAG and do I need it?

Retrieval-augmented generation means retrieving relevant pieces of your data and handing them to the model with the question, so the model answers from facts instead of memory. You need something like it whenever the model must answer from your documents or data.

How do I know if the AI is working?

You need examples where you know the right answer, and you score the model against them. Build that evaluation set early. Without it you are guessing, and the product will drift without you noticing.

An AI application is mostly a normal application, and the teams that succeed treat the model as one component among several. Name the job, clean the data, buy the commodity, keep a human in the loop, and measure the result. The model is the easy part. The discipline around it is the product.

Talk to us about your project

Need help applying this?

Tell us what you are building and where you are today. We typically reply within 24 hours.