The core risk
Not that an offshore team fails to deliver. That they deliver an impressive demo and call it a system — no evaluation, no guardrails, no cost control, no observability. Your entire evaluation should be designed to surface that gap before you sign.
AI has drawn a large number of firms into offering services they have not yet built the practice for. This is how to tell the difference, whether the vendor is in Bangalore, Warsaw or Austin.
Why AI vendor evaluation is unusually hard
| Normal software | AI systems |
|---|---|
| A demo shows real capability | A demo shows chosen examples |
| Bugs are visible | Wrong answers look plausible |
| Portfolio proves competence | Anyone can wire an API to a chat box |
| Quality is judged by users | Users cannot judge factual accuracy |
| Cost is predictable | Cost varies with usage patterns |
The consequence: the usual procurement signals — polished demo, confident answers, impressive client list — carry almost no information here. A weekend project can demo well. What you are trying to detect is whether the team has ever taken one of these systems into production and lived with it afterwards.
The twelve questions
Capability
- "How will you measure whether the agent is accurate?" You want to hear about an evaluation set of real queries with expected answers, scored automatically. Vagueness here is the strongest single negative signal.
- "What accuracy would you consider acceptable for go-live?" A number, agreed in advance. Teams who have shipped production AI answer immediately; teams who have not find the question uncomfortable.
- "What happens when it does not know the answer?" Confidence thresholds, refusal, human handoff. "It will say it does not know" without describing the mechanism is a prompt instruction, not engineering.
- "How do you handle instructions hidden in retrieved documents?" Prompt injection. A team that has thought about production security answers specifically; others are visibly surprised.
Cost and operations
- "What will this cost per month at our expected volume?" A projection based on your numbers, not a flat figure.
- "How will we see what it cost us, broken down?" Cost attribution per feature or conversation.
- "How will we know if quality degrades after launch?" Monitoring and regression evaluation. Models change; systems drift.
- "Who owns this after handover?" Someone on your side must. A vendor who has not raised this has not thought past delivery.
Commercial
- "What is explicitly excluded from this quote?" The most revealing question in any procurement.
- "Who specifically will build this?" Named engineers with relevant work, not a company capability statement.
- "Can we start with a small paid piece?" Willingness to be judged on contained scope says a great deal.
- "When would you tell us to buy an off-the-shelf tool instead?" A vendor who cannot name those conditions is selling, not advising.
If you only ask three: how accuracy will be measured, what is excluded from the quote, and when they would tell you not to build. Those three filter out most of the risk, and none of them require you to be technical.
Warning signs
- Accuracy discussed only qualitatively — "very accurate", "highly reliable", no numbers or method.
- No mention of evaluation until you raise it.
- Guardrails described as a prompt rather than as system design.
- Fixed price on genuinely uncertain scope — either padded or heading for a dispute.
- Claimed expertise across every AI domain. The field is too young for that to be true.
- Reluctance to name the engineers who will do the work.
- No questions about your data before quoting. Data quality determines everything downstream.
A vendor who asks hard questions about your documentation quality before quoting understands the work. One who quotes without seeing your data is pricing a demo.
Structuring the engagement
- Paid discovery, time-boxed. Ends with a written scope, an evaluation set built from your real queries, and a working proof against your actual data.
- Build, with evaluation, guardrails and observability as named deliverables — not assumptions.
- Pilot with limited users, instrumented for quality and cost.
- Decision gate against the accuracy threshold agreed in phase one.
- Handover — documentation, evaluation suite, and a named owner on your side.
Contract terms specific to AI work
- Full IP assignment including prompts, evaluation sets and configuration — not just application code.
- The evaluation suite as a deliverable. Without it you cannot safely change anything later.
- Documented model and provider choices, including data retention terms.
- Your own API accounts where practical, so you see spend directly and can revoke access.
- Your repository from day one.
- Data handling terms — what leaves your environment, to whom, and under what policy.
Evaluating AI development partners? Ask us these twelve questions — including what we would exclude and when we would tell you not to build. See our AI agent service, cost comparison, and why pilots fail.