All insights
EN · AI Agents, Automation & Operations

How to define AI service levels, support, and incident response before hiring a provider

Turn uptime into an operating agreement for AI quality, severity, detection, response, recovery, evidence, and continuous improvement.

A production AI service flows through quality, monitoring, incident, recovery, and evidence-review controls.
The AI Service Reliability Compact-8 connects outcomes, SLIs, SLOs, severity, response, recovery, and contractual remedies. · Generated with OpenAI

Before hiring a company to operate production AI, define the service by the outcome users receive—not infrastructure uptime alone. Measure availability and latency, but also response or action quality, safety, permissions, cost, human queues, and recovery. For every measure, name the source, window, exclusions, owner, and consequence.

Separate three layers. An SLI is an observable measure; an SLO is an operating target or range; an SLA is a contractual commitment with a consequence. A dashboard without consequences is not an SLA. Credits without diagnosis, correction, and exit rights do not protect the business. The buyer should retain raw-data visibility and authority to degrade, pause, or stop automation.

The AI Service Reliability Compact-8: 32 points before signature

Score each dimension from zero to four: zero is absent; one is a promise; two is partial; three is a documented obligation with an owner and evidence; four is tested under a realistic scenario. For a material service, use 24 of 32 as a planning threshold, no zero, and at least three in quality, incidents, recovery, and evidence. Tailor it to criticality; this is not a universal benchmark.

  • Service outcome — users, critical journeys, permitted actions, volume, hours, and degraded operation.
  • Technical SLIs and SLOs — availability, success, percentile latency, throughput, queues, and dependencies.
  • AI quality and safety — evaluations, groundedness, harmful error, policy, autonomy, and human review.
  • Observability and detection — logs, traces, version, cost, drift, abuse, coverage, and actionable alerts.
  • Severity and communication — impact, data, customers, autonomy, timing, channel, audience, and updates.
  • Response and recovery — on-call, containment, fallback, rollback, RTO/RPO, restoration, and validation.
  • Change governance — release, regression, approval, exception, model supplier, and configuration.
  • Evidence and remedies — raw data, report, cause, corrective action, credit, holdback, audit, and exit.

Define the service boundary before the percentage

Map the full path: user input, data and permissions, context retrieval, model, tools, integrations, human review, response or action, and final record. State which third parties enter the calculation and which events qualify as exclusions. If the API responds but the action is wrong, unsafe, or sent to the wrong customer, the service is not healthy.

Use a small set of indicators that represent experience and risk

  • Useful availability: eligible attempts that complete the correct task.
  • End-to-end latency by percentile, including tools, human review, and queues when applicable.
  • Quality: pass rate on the evaluation set and risk-stratified production samples.
  • Safety: correctly blocked actions, policy violations, unauthorized access, and exposed data.
  • Economics: cost per accepted task, anomalous consumption, and the limit before controlled degradation.
  • Operations: fallback, human escalation, rework, reopened cases, and time to validated recovery.

Google SRE recommends a few objective, representative indicators and warns that averages conceal tails. AI aggregates can also hide failures by language, population, channel, or high-impact task. Define segment thresholds and preserve the denominator. A traffic mix change must not manufacture better performance.

Create AI-specific severity levels

  • Sev 1 — ongoing material harm, incorrect autonomous action at scale, data exposure, fraud, safety risk, or loss of control; contain immediately and activate crisis leadership.
  • Sev 2 — significant degradation in a critical outcome, group, or workflow without confirmed material harm; reduce autonomy, activate fallback, and investigate.
  • Sev 3 — localized, recoverable failure with a workaround and low impact; repair through the agreed operating process.
  • Sev 4 — cosmetic defect, question, or improvement without operating impact; route to the backlog.

Classify by impact, scope, data, reversibility, and autonomy—not the failed component. A technically fast response can be Sev 1 when it makes prohibited decisions. Name who may declare or lower severity, who must be consulted, and when customers, leadership, legal, security, affected people, regulators, or insurers are notified.

Make the incident timeline measurable

Separate time to detect, acknowledge, contain, update, recover, validate, and complete root-cause analysis. Start the clock at the earliest signal available to either party. Define coverage by time zone, language, and channel; named contacts; update cadence; minimum content; and an alternate channel if the primary path fails.

Tie remedies to recovery, not credits alone

Credits can discipline performance but rarely offset business harm. Use a ladder: incident report and corrective plan; added capacity; change freeze; holdback; service credit; focused audit; step-in or transition; and termination for repeated or critical failure. Define when related events form one incident and how recurrence is measured.

Do not accept broad exclusions for cloud, model, or subcontractor failure when the provider selects, integrates, or operates that dependency. Allocate responsibility to real control. Planned maintenance needs notice, a window, a cap, and an operating alternative; it cannot become an unlimited exclusion.

Test the provider before award

Give finalists the same scenario: an update increases confident but incorrect answers, an agent executes two unauthorized actions, and the model provider begins failing intermittently. In 75 minutes, request detection, severity, containment, communication, fallback, return-to-service decision, evidence, and post-incident plan. Watch who assumes command and whether the team protects users before discussing credits.

Require eight operating artifacts

  • Service catalog and dependency map.
  • SLI–SLO–SLA cards with formula, source, window, owner, and consequence.
  • Severity matrix and escalation tree.
  • Containment, fallback, rollback, and restoration runbooks.
  • Inventory of models, prompts, data, tools, and versions.
  • Evaluation set and regression thresholds.
  • Incident and root-cause report template.
  • Change, exercise, corrective-action, and accepted-risk logs.

U.S. context and connected decisions

Align the operating agreement with the contracting entity, delegated authority, state breach-notification rules, sector requirements, insurance, privacy commitments, and customer contracts. Do not assume one vendor clock satisfies every reporting obligation. Require early notice and cooperation while facts are incomplete, with later updates and preserved evidence. Local counsel should validate legal timing and content.

Use https://makinai.co/insights/en/how-to-choose-managed-ai-services-provider to select the operator, https://makinai.co/insights/en/how-to-choose-ai-evaluation-testing-company for the evaluation system, https://makinai.co/insights/en/security-due-diligence-ai-services-company for diligence, and https://makinai.co/insights/en/what-to-include-ai-services-contract-sow for the main agreement.

When to involve MAKINAI

MAKINAI can help map the service boundary, define indicators, severity, runbooks, incident tests, and remedies before selection or launch. Explore https://makinai.co/services/en/ai-agents-automation-development. The goal is not zero-failure theater; it is early detection, bounded harm, evidence-based recovery, and learning without losing control.

Sources and references

  1. NIST SP 800-61 Rev. 3 — Incident Response Recommendations · National Institute of Standards and Technology

    Integrates preparation, detection, response, recovery, communications, lessons learned, and continuous improvement into cybersecurity risk management.

    2026-09-07
  2. NIST AI 600-1 — Generative AI Profile · National Institute of Standards and Technology

    Treats generative-AI risk management as a lifecycle activity and connects measurement, monitoring, incident disclosure, and third-party risk.

    2026-09-07
  3. Google SRE — Service Level Objectives · Google

    Separates service indicators, objectives, and agreements, and requires objective measurement tied to explicit consequences.

    2026-09-07
  4. UK Government — Artificial Intelligence Playbook · UK Government

    Calls for continuous AI performance monitoring, managed releases, rollback, escalation routes, quality assurance, and fallback processes.

    2026-09-07
Making connections

Continue exploring

AI Agents, Automation & Operations

How to choose a managed AI services provider

Read insight
AI Agents, Automation & Operations

How to choose an AI integration partner for enterprise systems

Read insight
AI Agents, Automation & Operations

How to choose an AI customer service company

Read insight
Related capability

Products, agents & automation

Building an enterprise AI agent is not just connecting a model to a chat interface. It requires product design, context, tools, integrations, identity, evaluation, guardrails, observability and human operations. MAKINAI builds the complete experience and measures whether it improves capability, quality or speed.

Explore this capability