Skip to main content
aienterprise-aiai-agents+2

What to Ask an AI Implementation Partner Before You Sign

The license is the cheap part. What an AI implementation partner owes you, what really moves the price of an AI build, and what to ask before you sign.

Buying the license is the easy part of this purchase. An AI seat arrives working: the model answers well enough that the demo lands in the room, and nobody has to be talked into it. Then someone asks it to close the ticket it just read, or post the credit note it just recommended, and the real project starts. That second thing is what an AI implementation partner is for, and it gets priced by what the software is allowed to do.

Read-only is the cheap half. An agent that answers questions over your own documents, staff-facing, behind your SSO, runs ₹4L – ₹6L ($4.5k – $6.8k) and ships in five to eight weeks, per our cost of AI agent development guide. That is an indicative range for a scoping conversation, not a quote. Letting the same agent change a record puts it in the tier above — ₹14L – ₹18L ($16k – $20k), three to four months. Same model, different permissions.

Why AI pilots stall after the license is signed

The model is the most replaceable component in the stack. Providers ship new ones on a schedule you do not control, and switching between them is a config change plus an evaluation run.

Everything around it is yours alone. Which system holds the authoritative version of a customer. What “closed” means in your workflow versus your client’s. Who is allowed to read the document the model just summarized. What the software does when the ERP accepts a write and the CRM rejects the same one.

None of that arrives in a vendor’s product, because none of it is knowable from outside your company. Microsoft and EY have committed $1 billion over five years to helping customers adopt AI, and a large part of what the money buys is jointly trained forward-deployed engineers — people, sitting inside the customer’s business. A vendor spending at that scale on humans has already priced the gap between capability and value.

Take an assistant sitting over open orders. Answering questions about them is retrieval, and a competent version of that is a small piece of work. Letting it change a delivery date is a different program. The write path drags in a permission check, an audit record, defined behavior for the update that half succeeded, and a queue for the writes a person has to approve. It also drags in a ruling on what the software does when nobody approves them.

What an AI implementation partner actually does

Identity comes first. The same customer exists as a payment-processor ID, a CRM record and an email address, and until one of them is declared the winner the agent cannot count that customer’s orders, let alone act on them. The ruling has consequences — what happens when the email matches but the company name does not — and it cannot be made from a product roadmap written in another country.

Then the unsure case. A model that is out of its depth produces something confident anyway, so the system needs a confidence threshold and a review queue, or a written policy naming the actions that never run without a person. Skip that decision and you have shipped a feature whose worst behavior stays invisible until a customer finds it.

Someone has to write down the cases the system is measured against, using your real records, so that the next model upgrade can be called better or worse instead of merely different. Without that set, every upgrade is a coin toss you are not allowed to decline, because the provider will eventually retire the version you built on.

Run cost gets noticed after launch and should have been settled before it. Every model call has a price, and a feature that looks cheap against a pilot’s volume behaves differently once it runs on every inbound email. Somebody has to know that number before the feature goes live, and to have built the caching and the cheaper fallback path that keep it flat.

Write paths and approvals are what move the price

The model choice barely registers. What moves the figure is how many systems the feature writes to, whether those writes can be reversed, whether a person approves before it commits, whether the data it reads carries access rules that vary by user, and whether a regulator will eventually read the audit trail.

Each of those is a decision about what the software is permitted to do, and each either adds weeks or removes them. Deciding that a refund over a threshold waits for a human is one line of policy and a real block of build: a queue, a notification, a record of who approved, and behavior for the approval that never arrives.

Which is why the work gets scoped in a paid discovery before a price is fixed. Nobody can price an integration into a system they have not opened, and a quote given without opening it is either padded or about to be revised. If the choice between building the AI layer and buying a product that already does most of the job is still open, that comparison is worth settling before anyone scopes anything.

Who owns AI governance when a partner writes the code

A competent engineer in your workspace designing the agent’s decision rules creates a quiet temptation: they built it, so let them own it. Regulated industries do not permit that. The gap surfaces at review, and by then the person who made the call may have rotated off the account.

A reviewer asks who decided the agent could issue a refund without a person, on what basis, and where that decision is written down. Whoever wrote the code will not be in the room. A partner can build the queue. Your organization has to own the policy saying which items go into it, and the plan for running the feature if the partner disappears.

The model-risk policy, the escalation rules, the exit plan and the record of who signed each one are your artifacts — drafted with the partner, signed by someone inside your company. An engagement that leaves behind working software and nothing else has undercharged you on the part that mattered.

Questions to ask an AI vendor before you sign

Ask these in the first call:

  • What will you refuse to let this system do in release one, and why?
  • What does it cost to run per month at our volume, and what makes that number grow?
  • Who writes the evaluation set, and do we own it afterward?
  • When the provider retires the model version, what breaks and who fixes it?
  • What is on the handover list — repository, pipeline, environments, credentials, runbook?

A vendor who treats the handover list as a formality expects to still be billing you in year three. A proposal that parks compliance in a later phase has not priced compliance. AI embedded in the product has to be designed for HIPAA, GDPR or SOC 2 from the first sprint, because retrofitting data residency and access logging means reopening the architecture rather than adding a module.

Deliverables an AI vendor should hand over at the end

State them as things you can hand to another team. The product in production, with source code and deploy pipeline transferred. The model layer: prompts and any tuned weights, plus the evaluation cases and the scores they got. The integrations documented down to what the feature does when a dependency is down — the part of integration work that stays invisible until the dependency is actually down. Architecture notes recording what was chosen, what was rejected, the monthly run cost, and the first thing that will break as volume grows.

Capability transfer belongs on that list, and it has a test: your own engineers ship the next change to the feature without the partner writing any of it. Put that in the contract as a named deliverable, and identify the change before handover so nobody gets to pick an easy one.

The engagements we turn down

An agent with write access to money and no approval step in release one. The queue and the audit record are cheap to design in and expensive to retrofit, and a pilot without them has no answer when finance asks who authorized the payment.

An AI layer over data whose definitions nobody has agreed. If sales and finance count revenue differently, a model reading across that gap learns the disagreement and repeats it in a confident tone. We can chair that conversation. We cannot settle it, and building on an unsettled definition produces a system that is precisely, reproducibly wrong.

A model where a rules engine belongs. Eligibility checks and pricing tiers that are stable and written down should be code, testable and explainable line by line. Putting a model in that path adds cost and a failure mode, and buys nothing except the ability to say the feature uses AI.

And a build with no named decision behind it. If nobody can point at a recurring judgment call somebody currently makes on instinct, there is nothing to measure the system against and no way to tell whether it worked. Name the decision, name the person making it today, then price the tier that decision sits in. Our cost guide says most teams land further up that list than they expect.

Have a project in mind?

Fixed price after a paid discovery — no hourly billing. A real engineer reads every enquiry, and we reply within 24 hours.