Direct answer: choose the provider that can prove how it will preserve business outcomes, response or action quality, reliability, security and cost as models, data and workflows change. A help desk and an availability SLA are necessary, but they are not enough. An AI service can remain online while becoming less useful: retrieval quality slips, a source becomes stale, an agent calls the wrong tool, cost per accepted task rises or human review absorbs the promised efficiency.
The core issue is the object of the contract. Traditional support restores components. Managed AI operations must preserve a business capability inside accepted boundaries. That requires a baseline, evaluation sets, telemetry, runbooks, authority to intervene and controlled change. NIST places operation and monitoring inside the AI lifecycle, and its Generative AI Profile connects risk management to organizational context and tolerance. Buyers should turn those principles into evidence that can be tested before a multiyear commitment.
Managed service, implementation project or internal team?
An implementation project builds or changes a system against a scoped acceptance point. An internal team provides control, institutional memory and fast trade-off decisions, but it needs coverage across product, data, engineering, security, evaluation and operations. A managed service fits when the capability already matters, changes frequently and needs continuous coverage the enterprise cannot yet sustain alone. The strongest structure is often hybrid: product and risk ownership stay with the client, while the provider operates an explicit layer and transfers capability throughout the term.
- Use a project when the main job is to build, migrate or integrate toward a defined completion milestone.
- Use managed operations when models, prompts, knowledge, tools and metrics need recurring supervision and improvement.
- Keep risk decisions, data ownership, change authorization and acceptance criteria inside the enterprise.
- Do not outsource accountability: name both an executive owner and an operational owner on the client side.
The Managed AI Operations Proof-7
MAKINAI uses seven proofs to compare finalists. Score each from zero to four: zero means absent; one means a promise; two means a documented process; three means evidence from a comparable environment; four means a demonstration using a controlled slice of your system. The maximum is 28. A proposal below 20 may still fit a narrow scope, but it should not receive broad operational authority. More important than the total are knockout gates: weak security, no tested rollback, unclear data ownership or inaccessible records cannot be offset by a polished presentation.
- 1. Outcomes: connects SLIs and SLOs to accepted results such as correct resolution, completed work or verified time saved.
- 2. Evaluation: maintains test sets, human sampling, regression checks and change approval.
- 3. Observability and cost: traces versions, latency, tokens, tools, errors, quality and cost per outcome.
- 4. Security and privacy: governs identity, permissions, retention, prompt injection, leakage and critical suppliers.
- 5. Incident and recovery: detects impact, constrains autonomy, rolls back, communicates and records corrective action.
- 6. Continuity: changes models or components, preserves data and operates through third-party disruption.
- 7. Transfer: provides documentation, runbooks, telemetry, decision history and training without manufactured lock-in.
Turn the seven proofs into an evidence-led RFP
Ask for a compact evidence room, not hundreds of pages. For each proof, require one artifact: an SLO map, a regression evaluation example, a telemetry schema, a threat model, an anonymized incident report, a continuity plan and a transition package. OpenTelemetry's work on generative AI makes signals such as model identity, token consumption and tool calls easier to discuss consistently. Google SRE supplies a useful distinction: an indicator measures, while an objective sets the target. The agreement needs both and must define what follows a breach.
Do not use generic uptime as the only SLA. An assistant can respond 99.9 percent of the time and still fail on the highest-value requests. Define a basket: technical availability, accepted completion rate, quality by segment, permission incidents, cost per task, percentile latency, human escalation rate and recovery time. Business KPIs should guide prioritization and improvement, even when the provider cannot guarantee them because adoption and process also matter. Separate operational SLOs, product targets and commercial KPIs so incentives remain intelligible.
Five knockout criteria
- The client cannot access logs, evaluations, costs or the change history.
- There is no tested rollback or safe mode to reduce autonomy during an incident.
- Responsibility across client, provider, model vendor, cloud and data sources is unclear.
- Exit terms do not cover code, prompts, data, vectors, documentation and third-party accounts.
- The buyer cannot test relevant security controls, including risks reflected in the OWASP GenAI LLM Top 10.
Run a transition drill before a long commitment
Use a paid 30-day test on one real system with constrained scope. In week one, the finalist reconstructs the service map and baseline. In week two, it instruments or validates telemetry and runs known evaluations. In week three, simulate three events: a model change, an unavailable knowledge source and a revoked permission. In week four, require analysis, remediation, executive communication and handoff. The point is not to stage a dramatic failure. It is to observe how the provider reasons when signals conflict and the fix crosses product, engineering, security and business teams.
What the first 90 days should produce
During days 1 to 30, freeze nonessential changes and confirm inventory, owners, risk tiers, SLOs and access. During days 31 to 60, close gaps in observability, evaluation, runbooks and incident response. During days 61 to 90, complete a cost-and-quality review, a continuity exercise and the first formal improvement cycle. A monthly steering forum should decide what to retain, repair, replace or scale. Managed AI operations are not an endless ticket queue; they are a decision system for the health and evolution of an operating capability.
Use this scorecard alongside MAKINAI's guides to choosing an enterprise AI integration partner, evaluating an AI agent development company, and defining an AI services contract and SOW. Those guides address construction and commercial terms; the Proof-7 addresses what happens after the system becomes important. If you are defining how to operate a copilot, agent or automation already in production, MAKINAI can help establish the baseline, test the operating model and structure a measurable transition.