The AI services market has expanded faster than the pool of teams who have taken a system into production and lived with it. These twelve questions surface the difference, and none of them require you to be technical.

Why the usual signals fail here

SignalWhy it does not help
Impressive demoDemos run on chosen examples
Client logosSays nothing about production experience
Team sizeThis work is done by individuals
Years in businessFrequently years of unrelated work
Confident answersConfidence is free; specificity is not

Capability — four questions

  1. "How will you measure whether it is accurate?"
    Want to hear: an evaluation set of real queries with expected answers, scored automatically.
    Worrying: "We'll test it thoroughly."
  2. "What accuracy would you consider acceptable for launch?"
    Want to hear: a number, and a proposal to agree it with you in advance.
    Worrying: discomfort with the question.
  3. "What happens when it does not know?"
    Want to hear: confidence thresholds, refusal behaviour, human handoff with context.
    Worrying: "We'll instruct it to say it doesn't know" — that is a prompt, not a mechanism.
  4. "How do you handle instructions hidden in documents it reads?"
    Want to hear: specifics about prompt injection and tool scoping.
    Worrying: visible surprise at the question.

Question four is the single most efficient filter. Prompt injection is a known production concern. A team that has shipped AI systems has an answer ready; a team that has only built prototypes has usually never encountered it — and their reaction tells you which you are speaking to.

Operations — four questions

  1. "What will this cost per month at our volume?" A projection based on your numbers, not a flat figure. Follow up with "what would double it?"
  2. "How will we see cost broken down?" Attribution per feature or per conversation, not a single API bill.
  3. "How will we know if quality degrades?" Monitoring, regression evaluation, alerting. Models change; systems drift.
  4. "Who owns this after handover?" A vendor who has not raised this has not thought past delivery.

Commercial — four questions

  1. "What is explicitly excluded from this quote?" The most revealing question in any procurement.
  2. "Who specifically will build this, and what have they built?" Named engineers. This work is not fungible headcount.
  3. "Can we start with a small paid piece?" Willingness to be judged on contained scope says a great deal.
  4. "When would you tell us to buy something instead?" A vendor who cannot name those conditions is selling, not advising.

If you only ask three: how accuracy will be measured, what is excluded, and when they would tell you not to build. Those three filter most of the risk in about ten minutes.

Warning signs

  • Accuracy discussed only in adjectives — "highly accurate", no method.
  • Evaluation never mentioned until you raise it.
  • Guardrails described as prompt instructions rather than system design.
  • A fixed price on genuinely uncertain scope.
  • Claimed expertise across every AI domain — the field is too young.
  • Reluctance to name the engineers.
  • No questions about your data quality before quoting.

A vendor who asks hard questions about your documentation before quoting understands the work. One who quotes without seeing your data is pricing a demo.

What the proposal should contain

  • Scope broken into components, with effort against each.
  • Evaluation approach and the accuracy target.
  • Named guardrails and which actions require human approval.
  • A cost projection at your expected volume.
  • Explicit exclusions.
  • Handover contents — documentation, evaluation suite, runbook.
  • Who does the work.

Contract terms specific to AI

  • IP assignment covering prompts, evaluation sets and configuration — not just application code.
  • The evaluation suite as a deliverable. Without it you cannot safely change anything later.
  • Documented model and provider choices, including data retention terms.
  • Your own API accounts where practical, so spend is visible and access revocable.
  • Your repository from day one.

Evaluating partners? Put these twelve questions to us — including what we would exclude and when we would tell you not to build. See our AI agent service, build vs buy, and hiring offshore AI developers.

Frequently asked questions

Most of these questions are answerable by a non-technical buyer. You are listening for specificity versus reassurance. A vendor who names a methodology is different from one who says "we test thoroughly", and you do not need an engineering background to hear that difference.
A small paid pilot with two vendors on the same scope is genuinely informative, though it costs time on your side to manage. A cheaper alternative is asking each to walk through how they would approach it — the plans differ more revealingly than the demos do.
It varies widely, but it should be a small fraction of the projected build and produce something you own — a written specification, an evaluation set, and a working proof against your data. If discovery produces only a document, you paid for a sales exercise.