Almost every product team is being asked the same question right now: "What are we doing with AI?" The pressure leads to two common mistakes. Some teams bolt a chatbot onto the homepage and hope it works. Others commission a large AI project with no clear goal and discover six months later that nobody uses it.
There is a calmer way to do it. Treat AI like any other feature: pick a real problem, test a small version with real data, put safeguards around it and only then expand. At Creuto we work with teams that already have a product, and this is the sequence we follow.
The principle
Start from a workflow that hurts, not from a model that excites you. A good AI feature is the one your customers or team stop noticing because the task just got easier.
Where AI actually helps in an existing product
AI is strongest where there is a lot of language, documents or messy input and where a person currently does the sorting, searching or first draft. Typical examples:
- Search and answers over your own content: staff or customers ask questions and get answers with links to the source document.
- Support triage: incoming tickets are labelled, routed and given a suggested reply for an agent to approve.
- Document extraction: invoices, contracts or forms become structured data without manual typing.
- Drafting and summarising: first drafts of emails, reports or meeting notes that a person edits.
- Recommendations and tagging: suggesting the next step, product or category based on past behaviour.
It is weaker where the answer must be exactly right every time with no review, such as final financial calculations or legal decisions. For those, ordinary rule-based software is usually safer and cheaper.
The 6-step plan
Step 1: Pick one workflow and one measure
Choose a single task that is repetitive, slow or expensive today. Write down how you measure it now, for example "support agents spend a long time finding the right policy" or "ops staff retype invoice details". That measure becomes your success test later. If you cannot name a number or behaviour to improve, choose a different workflow.
Step 2: Check your data
AI is only as useful as the information it can use. Ask:
- Where does the relevant information live: documents, a database, a help centre, email?
- Is it current, consistent and permitted to be used this way?
- Who is allowed to see what? A good feature respects the same permissions your product already has.
Messy or outdated data is the most common reason a promising pilot disappoints. Fixing it is often the real first task.
Step 3: Choose the right approach
There is more than one way to add AI, and the simplest that works is usually the right one:
| Approach | Good for | Watch out for |
|---|---|---|
| Prompting a hosted model | Drafting, summarising, simple classification | Consistency and sending sensitive data out |
| Retrieval over your documents (RAG) | Answers grounded in your own content with sources | Quality of documents and permissions |
| Fine-tuning a model | A very specific style or format at scale | Needs good examples, rarely the first step |
| Classic machine learning or rules | Predictions on structured data, exact logic | Not suited to free text |
We run a short feasibility check to decide which approach fits, and sometimes the answer is "this does not need AI". That honest answer saves money. You can read more about this in our post on pragmatic enterprise AI workflows.
Step 4: Design guardrails and a human check
This step separates a demo from a feature you can trust. Plan these before you build:
- Grounding and citations: answers come from your documents and show where they came from.
- Permissions: the feature only retrieves what the current user may see.
- Validation: outputs are checked against rules, such as required fields or allowed values.
- Human approval: anything risky, such as finance, health or legal content, goes to a person before it reaches a customer.
- Spend limits and logging: usage costs are capped and every request is recorded for review.
Not sure where AI fits in your product?
Book a free 45-minute call. We look at your workflows, tell you honestly where AI helps and where it does not, and outline a small first pilot.
Book a free AI strategy call →Step 5: Run a small pilot and measure it
Release to a small group, such as one team or a slice of customers. Compare the measure you picked in step 1 before and after. Collect examples where the feature failed, because those teach you the most. Decide in advance what result counts as success so the decision is not driven by excitement.
Before launch, test on your own real questions or documents, not only on clean samples. Only ship what passes.
Step 6: Ship, monitor and improve
After launch, watch quality, usage and cost. Models and data change, so check results regularly, keep a way for users to flag bad answers and update your documents as the business changes. Expand to the next workflow only once the first one is stable and clearly paying for itself.
Three example pilots
To make this concrete, here are three hypothetical first pilots. They show how small and specific a good starting point is.
- Internal knowledge search: a support team asks questions in plain language and gets answers drawn from the help centre and policy documents, each with a link to the source. Measure the time agents spend searching before and after.
- Ticket triage: incoming requests are labelled by topic and urgency and routed to the right queue with a suggested reply. A person approves the reply. Measure first-response time and misrouted tickets.
- Invoice capture: uploaded invoices are read and turned into structured fields for review. Staff correct exceptions instead of retyping every field. Measure minutes per invoice and error rate.
Notice the pattern: one workflow, one measure, a person in the loop and a clear before-and-after comparison.
How to measure whether it paid off
Judge an AI feature the same way you judge any investment. Compare what you measured in step 1 against the pilot group, and count the full cost: build time, running costs such as model usage, and the time people spend reviewing results.
- Time saved: minutes per task, multiplied by how often it happens.
- Quality: error rates, rework, customer satisfaction or resolution speed.
- Adoption: how many people use it voluntarily after the first week. Low adoption usually means it is not helping.
- Running cost: monthly usage against the value created.
If the numbers do not justify scaling, that is still a good outcome. You learned cheaply, and you can try a different workflow.
What to prepare before you talk to an AI team
- The one workflow you want to improve, and how it is done today.
- Examples of real inputs and the correct outputs, such as past tickets with the replies that worked.
- Where the data lives and who is allowed to access it.
- Any rules about customer data, regions or industries that limit what can be sent to an outside provider.
- A rough view of how you will decide the pilot succeeded.
Having these ready makes the first conversation far more useful, and lets a team give you an honest view on feasibility much faster.
The risks to plan for
- Privacy: know what data leaves your systems, which provider handles it and under what agreement. Send only what the task needs.
- Wrong answers: plan for them with sources, validation and review rather than hoping they will not happen.
- Cost creep: usage-based pricing can grow quietly. Cache repeated work and set limits.
- Over-promising: tell users what the feature can and cannot do, and show when something is AI-generated.
Should you build it or buy a ready-made tool?
If a packaged tool already does the job and your needs are common, buy it. Build when the feature depends on your own data, workflows or permissions, or when it should feel like a native part of your product. A first pilot is a good way to find out which side you are on without a large commitment.
If you are still working out what you want built, our guide to scoping a product before you build applies to AI features too. Start with a narrow first release and a written plan. You can also read about our OpenAI Select Partner status and what it means for the work.
Frequently asked questions
Do I need to rebuild my product to add AI?
No. Most useful AI features are added to an existing product as a new service that your current system calls, such as search over your documents, a drafting assistant or automatic sorting of incoming requests. Your core product stays as it is.
How do I know if AI is the right tool for my problem?
Ask whether the task involves language, documents, images or fuzzy judgement that rules cannot capture, and whether an occasional mistake is acceptable or can be caught by a person. If the answer needs to be exactly right every time, a normal rule-based feature is often better.
Will my company data be used to train public AI models?
It does not have to be. Enterprise agreements with major providers and private deployment options keep your data out of public model training. Confirm this in writing for any provider you use and limit what you send to what the task needs.
How do I stop an AI feature from making things up?
You reduce it, rather than remove it. Ground answers in your own documents, require the feature to cite its sources, validate outputs against rules, and send risky answers to a person for approval. Test on your real questions before launch.
How long does a first AI pilot take?
A narrow pilot on one workflow can often be built and tested in a few weeks. The timeline depends on how clean your data is and how many systems it touches. We scope it after looking at your workflow and data.
Ready to scope a first AI feature?
Creuto is an OpenAI Select Partner. We start with one workflow, test it on your real data and ship only what passes.