Direct answer: an AI services contract should turn vendor promises into verifiable evidence. Beyond scope, fees and dates, the SOW should define outcomes, acceptance, data and IP, models and third parties, risk and security, operations and change, and transfer and exit. Legal language is jurisdiction-specific; this guide structures commercial and technical decisions for counsel to review.
Conventional software terms often treat delivery as deterministic. AI behavior can shift with data, models, prompts, policies and usage. “Working” is therefore not an adequate acceptance criterion. The contract should state how performance will be tested, under which conditions, who accepts residual risk and what happens when performance degrades.
AI Contract Evidence-7: seven proofs before signature
For each dimension, require an obligation, the evidence proving it and the resulting decision. Score each 0–3: absent, described, measurable or operationally tested. A low-priced proposal without evidence for acceptance, data or exit usually defers cost and risk until after signature.
1. Outcome: what business change is being purchased?
Name the process, audience, decision or task being changed, plus baseline, metric and observation window. Separate supplier deliverables — discovery, prototype, integrations, evaluations and runbooks — from business outcomes that also depend on client adoption. Assign the product owner, assumptions and dependencies. Avoid language such as “drive efficiency with AI.”
2. Acceptance: what test releases each phase and payment?
Attach the evaluation set, normal cases, exceptions, thresholds, human review and test environment. Cover quality, latency, availability, cost per task, security and user experience. State who runs and witnesses tests, how disputes are resolved and when retesting occurs. Accept evidence, not a demo.
Use progressive gates: continue after discovery; prove value with representative data; establish production readiness; and complete post-launch stabilization. Probabilistic systems can use ranges and tolerances, but the method, sample and decision rule must be reproducible.
3. Data and ownership: who may use what, and why?
Inventory inputs, derived data, feedback, logs, prompts, embeddings, outputs and personal data. Specify purpose, access, retention, location, deletion, subprocessors and training use. Distinguish pre-existing assets from newly created code, configurations, evaluations and content. Counsel should align the schedule with applicable privacy, IP and sector rules.
4. Models and third parties: which dependencies may change?
List critical models, APIs, libraries, clouds and versions. Require disclosure of material substitutions, pre-upgrade testing and impact on price, performance or rights. State who approves model changes, how unplanned lock-in is controlled and which components are replaceable. The supply chain must be visible enough to govern.
5. Risk and security: which controls become operating duties?
Map risk tier to access controls, environment separation, evaluation, human oversight, proportionate red teaming, incident response, logs and audit. NIST calls for lifecycle management of third-party technology and providers; SP 800-218A applies secure development practices to AI software. Set notice periods and responsibilities without promising zero risk.
6. Operations and change: who maintains quality after launch?
Define SLOs, observability, support, escalation, operating cost, evaluation maintenance and review cadence. Trigger revalidation when the model, data, purpose or autonomy changes, or after an incident. Separate defect, improvement and scope change to avoid disputes. NIST’s Playbook calls for continuous monitoring of third-party resources.
7. Transfer and exit: can the client continue or switch?
List code, repositories, documentation, configuration, prompts, evaluations, exportable data, credentials, runbooks and training. Define format, timing and a portability test. Cover transition help, data return or destruction and continuity during exit. Permanent dependency may be an economic choice, but it should never be a surprise.
Build the contract stack
- MSA: commercial terms, confidentiality, liability and general rules. SOW: outcome, phases, team, dependencies, price, acceptance and change. AI schedule: data, models, evaluations, risks, transparency and controls. Operations schedule: SLOs, support, incidents, monitoring and costs. Exit plan: portability, transfer and termination.
The pre-signature failure test
Choose one critical scenario and one plausible failure. Ask the supplier which obligation applies, what evidence will be produced, who decides, how long it takes and what remedy exists. If the answer depends on goodwill or a future meeting, the SOW is not ready. The EU model AI clauses are useful reference material, but the Commission explicitly notes that they are not a complete contract and require customization.
Next step
Start with the RFP guide at https://makinai.co/insights/en/how-to-write-rfp-ai-services, compare partner types at https://makinai.co/insights/en/ai-strategy-firm-or-implementation-partner and model investment at https://makinai.co/insights/en/how-much-ai-consulting-services-cost. When decision, build and operation must connect, see https://makinai.co/services/en/ai-strategy-transformation-consulting.