The AI services market has expanded faster than the pool of teams who have taken a system into production and lived with it. These twelve questions surface the difference, and none of them require you to be technical.
Why the usual signals fail here
| Signal | Why it does not help |
|---|---|
| Impressive demo | Demos run on chosen examples |
| Client logos | Says nothing about production experience |
| Team size | This work is done by individuals |
| Years in business | Frequently years of unrelated work |
| Confident answers | Confidence is free; specificity is not |
Capability — four questions
- "How will you measure whether it is accurate?"
Want to hear: an evaluation set of real queries with expected answers, scored automatically.
Worrying: "We'll test it thoroughly." - "What accuracy would you consider acceptable for launch?"
Want to hear: a number, and a proposal to agree it with you in advance.
Worrying: discomfort with the question. - "What happens when it does not know?"
Want to hear: confidence thresholds, refusal behaviour, human handoff with context.
Worrying: "We'll instruct it to say it doesn't know" — that is a prompt, not a mechanism. - "How do you handle instructions hidden in documents it reads?"
Want to hear: specifics about prompt injection and tool scoping.
Worrying: visible surprise at the question.
Question four is the single most efficient filter. Prompt injection is a known production concern. A team that has shipped AI systems has an answer ready; a team that has only built prototypes has usually never encountered it — and their reaction tells you which you are speaking to.
Operations — four questions
- "What will this cost per month at our volume?" A projection based on your numbers, not a flat figure. Follow up with "what would double it?"
- "How will we see cost broken down?" Attribution per feature or per conversation, not a single API bill.
- "How will we know if quality degrades?" Monitoring, regression evaluation, alerting. Models change; systems drift.
- "Who owns this after handover?" A vendor who has not raised this has not thought past delivery.
Commercial — four questions
- "What is explicitly excluded from this quote?" The most revealing question in any procurement.
- "Who specifically will build this, and what have they built?" Named engineers. This work is not fungible headcount.
- "Can we start with a small paid piece?" Willingness to be judged on contained scope says a great deal.
- "When would you tell us to buy something instead?" A vendor who cannot name those conditions is selling, not advising.
If you only ask three: how accuracy will be measured, what is excluded, and when they would tell you not to build. Those three filter most of the risk in about ten minutes.
Warning signs
- Accuracy discussed only in adjectives — "highly accurate", no method.
- Evaluation never mentioned until you raise it.
- Guardrails described as prompt instructions rather than system design.
- A fixed price on genuinely uncertain scope.
- Claimed expertise across every AI domain — the field is too young.
- Reluctance to name the engineers.
- No questions about your data quality before quoting.
A vendor who asks hard questions about your documentation before quoting understands the work. One who quotes without seeing your data is pricing a demo.
What the proposal should contain
- Scope broken into components, with effort against each.
- Evaluation approach and the accuracy target.
- Named guardrails and which actions require human approval.
- A cost projection at your expected volume.
- Explicit exclusions.
- Handover contents — documentation, evaluation suite, runbook.
- Who does the work.
Contract terms specific to AI
- IP assignment covering prompts, evaluation sets and configuration — not just application code.
- The evaluation suite as a deliverable. Without it you cannot safely change anything later.
- Documented model and provider choices, including data retention terms.
- Your own API accounts where practical, so spend is visible and access revocable.
- Your repository from day one.
Evaluating partners? Put these twelve questions to us — including what we would exclude and when we would tell you not to build. See our AI agent service, build vs buy, and hiring offshore AI developers.