The core risk

Not that an offshore team fails to deliver. That they deliver an impressive demo and call it a system — no evaluation, no guardrails, no cost control, no observability. Your entire evaluation should be designed to surface that gap before you sign.

AI has drawn a large number of firms into offering services they have not yet built the practice for. This is how to tell the difference, whether the vendor is in Bangalore, Warsaw or Austin.

Why AI vendor evaluation is unusually hard

Normal softwareAI systems
A demo shows real capabilityA demo shows chosen examples
Bugs are visibleWrong answers look plausible
Portfolio proves competenceAnyone can wire an API to a chat box
Quality is judged by usersUsers cannot judge factual accuracy
Cost is predictableCost varies with usage patterns

The consequence: the usual procurement signals — polished demo, confident answers, impressive client list — carry almost no information here. A weekend project can demo well. What you are trying to detect is whether the team has ever taken one of these systems into production and lived with it afterwards.

The twelve questions

Capability

  1. "How will you measure whether the agent is accurate?" You want to hear about an evaluation set of real queries with expected answers, scored automatically. Vagueness here is the strongest single negative signal.
  2. "What accuracy would you consider acceptable for go-live?" A number, agreed in advance. Teams who have shipped production AI answer immediately; teams who have not find the question uncomfortable.
  3. "What happens when it does not know the answer?" Confidence thresholds, refusal, human handoff. "It will say it does not know" without describing the mechanism is a prompt instruction, not engineering.
  4. "How do you handle instructions hidden in retrieved documents?" Prompt injection. A team that has thought about production security answers specifically; others are visibly surprised.

Cost and operations

  1. "What will this cost per month at our expected volume?" A projection based on your numbers, not a flat figure.
  2. "How will we see what it cost us, broken down?" Cost attribution per feature or conversation.
  3. "How will we know if quality degrades after launch?" Monitoring and regression evaluation. Models change; systems drift.
  4. "Who owns this after handover?" Someone on your side must. A vendor who has not raised this has not thought past delivery.

Commercial

  1. "What is explicitly excluded from this quote?" The most revealing question in any procurement.
  2. "Who specifically will build this?" Named engineers with relevant work, not a company capability statement.
  3. "Can we start with a small paid piece?" Willingness to be judged on contained scope says a great deal.
  4. "When would you tell us to buy an off-the-shelf tool instead?" A vendor who cannot name those conditions is selling, not advising.

If you only ask three: how accuracy will be measured, what is excluded from the quote, and when they would tell you not to build. Those three filter out most of the risk, and none of them require you to be technical.

Warning signs

  • Accuracy discussed only qualitatively — "very accurate", "highly reliable", no numbers or method.
  • No mention of evaluation until you raise it.
  • Guardrails described as a prompt rather than as system design.
  • Fixed price on genuinely uncertain scope — either padded or heading for a dispute.
  • Claimed expertise across every AI domain. The field is too young for that to be true.
  • Reluctance to name the engineers who will do the work.
  • No questions about your data before quoting. Data quality determines everything downstream.

A vendor who asks hard questions about your documentation quality before quoting understands the work. One who quotes without seeing your data is pricing a demo.

Structuring the engagement

  1. Paid discovery, time-boxed. Ends with a written scope, an evaluation set built from your real queries, and a working proof against your actual data.
  2. Build, with evaluation, guardrails and observability as named deliverables — not assumptions.
  3. Pilot with limited users, instrumented for quality and cost.
  4. Decision gate against the accuracy threshold agreed in phase one.
  5. Handover — documentation, evaluation suite, and a named owner on your side.

Contract terms specific to AI work

  • Full IP assignment including prompts, evaluation sets and configuration — not just application code.
  • The evaluation suite as a deliverable. Without it you cannot safely change anything later.
  • Documented model and provider choices, including data retention terms.
  • Your own API accounts where practical, so you see spend directly and can revoke access.
  • Your repository from day one.
  • Data handling terms — what leaves your environment, to whom, and under what policy.

Evaluating AI development partners? Ask us these twelve questions — including what we would exclude and when we would tell you not to build. See our AI agent service, cost comparison, and why pilots fail.

Frequently asked questions

Ask process questions rather than technical ones. "How will you measure accuracy?" and "What happens when it does not know?" are answerable by any non-technical buyer, and the specificity of the answer is highly informative regardless of your background.
Prefer a small paid one. Free POCs attract vendors optimising for a demo that wins the deal, not for an honest assessment. A paid discovery phase produces something you own and reveals how the team actually works.
Ask what adjacent work they have done and what they would be learning on your project. Everyone was new once. The problem is not inexperience — it is inexperience presented as expertise, which is what the questions in this article surface.