Direct answer: choose an AI customer service company by its demonstrated ability to resolve real journeys safely, not by the fluency of a demo. The provider should prove seven capabilities: demand understanding, governed knowledge, identity and authorization, end-to-end resolution, human escalation, quality and risk, and operating economics. Make finalists run the same production-oriented conversations and decisions.
Bot deflection or containment is not the same as resolution. A customer may abandon, contact another channel or receive a plausible answer without completing the task. Build the business case around a verified outcome: the issue is resolved, the action is correct, context is preserved and a person remains reachable.
Service Resolution Proof-7: seven proofs for vendor selection
Score every dimension from 0 to 4: absent, promised, demonstrated, measured with representative data, or operated with a named owner. The maximum is 28, but four gaps should disqualify a proposal: no representative evaluation set, no human escalation, no action-level permission controls, or no incident trail and exit package.
1. Demand: does the system understand real contact reasons?
Require a taxonomy built from actual conversations, not a generic intent list. It must cover ambiguity, multiple requests, errors, emotion, language switching and cases that should not be automated. NIST’s AI Use Taxonomy starts from human activities and intended outcomes; that discipline helps keep technology from dictating the service journey.
The minimum proof is a map connecting contact reason, channel, audience, required data, possible action, risk tier and outcome. The supplier must explain what the system will answer, execute, route or refuse.
2. Knowledge: does every answer have authority and freshness?
Examine how policies, products, orders and procedures reach the system; who approves content; how stale versions are removed; and which users may access each source. Require traceable evidence when an answer depends on enterprise knowledge. A capable partner evaluates retrieval separately from answer quality and knows when the system should abstain.
MAKINAI’s RAG vendor guide covers this layer at https://makinai.co/insights/en/how-to-choose-rag-ai-knowledge-system-company. In customer service, knowledge must connect to channel, identity and action—not remain an isolated search experience.
3. Identity and action: can the agent do only what it should?
Answering FAQs is different from viewing an order, editing a profile, granting credit or canceling service. Ask for authentication, action-level authorization, value limits, confirmation, reversal and human approval. Credentials should follow least privilege; the model should not inherit broad access merely because an integration exposes it.
OWASP documents how prompt injection can cause unauthorized disclosure or action, and recommends least privilege, deterministic validation and human approval for high-risk operations. Test malicious instructions, contaminated documents, requests for another customer’s data and attempts to bypass policy.
4. Resolution: which outcome will be verified?
Define resolution per journey: correct information understood, order actually changed, payment restored or next step accepted. Also measure seven-day repeat contact, action correctness, total time and customer effort. Do not let the bot declare its own success without independent confirmation from the system of record or the customer.
5. Escalation: does the human receive useful context?
Handoff needs explicit triggers: low confidence, customer request, vulnerability, distress, regulated risk, disallowed action or repeated failure. Verify the right queue, faithful summary, history, collected data and a way for the human to correct the system. “Try again later” is not an escalation strategy.
Test accessibility as part of service acceptance. WCAG 2.2 provides testable criteria for digital interfaces; voice, chat and authentication flows also need usable alternatives for people with different access needs.
6. Quality and risk: how is safe behavior proven?
Require a versioned evaluation set, adversarial cases, risk-based thresholds, human review, regression tests before model changes and incident records. NIST’s Generative AI Profile frames lifecycle risk management. The proposal must identify who can block, roll back or narrow a capability when performance drops.
7. Operations and economics: who maintains service, and at what unit cost?
Compare cost per verified resolution, not only cost per conversation or token. Include models, retrieval, voice, integrations, observability, evaluations, supervision, human support and continuous improvement. Define SLOs, on-call response, model changes, portability of prompts and evals, log export and ownership of knowledge operations.
Run the replay test before a pilot
Give finalists 30–50 anonymized conversations covering ambiguity, an upset customer, sensitive data, an unsupported request, system outage, language switching, a reversible action and required human help. Run the same script and observe answer, action, refusal, handoff, logs and cost. This is far more diagnostic than a staged demo.
An 8–12 week pilot
- Choose one journey with volume and a verifiable outcome. Capture the human and digital baseline. Freeze the evaluation set before configuration. Integrate read access before write access. Run shadow mode, then constrained production. Measure verified resolution, action accuracy, handoff, repeat contact, time, effort, accessibility and cost. Scale only after business, operations, security and service owners pass one shared gate.
Next step
Use this scorecard with the RFP guide at https://makinai.co/insights/en/how-to-write-rfp-ai-services and the production-agent checklist at https://makinai.co/insights/en/how-to-evaluate-ai-agent-development-company. To design, test and operate customer service automation with real integrations and controls, see https://makinai.co/services/en/ai-agents-automation-development.