Direct answer: select the partner by its ability to prove a controlled end-to-end enterprise transaction, not an isolated model response. Require evidence across seven dimensions: outcome and boundary, data contracts, identity and authorization, orchestration and APIs, reliability and fallback, observability and economics, and ownership and transfer. Before rollout, run a test with representative data, an unavailable integration, a denied permission and a contract change. The company should demonstrate how it prevents unauthorized actions, records every material decision and keeps the workflow operable when either AI or a legacy system fails.
Most enterprise AI projects do not end at the model. To create value, the system must retrieve sources, interpret context, request approval, write to systems of record, alert people and handle exceptions. That chain is where delays, cost and risk accumulate. A capable partner treats legacy systems as part of the product: it understands their contracts, constraints, owners and failure modes, then proposes incremental modernization without hiding dependencies.
Integration Delivery Proof-7
Score each dimension from 0 to 4: absent, described, demonstrated, validated in a representative environment or operated in a comparable context. The maximum is 28. Identity, authorization, traceability and recovery should be non-compensable gates. Give every finalist the same evidence request.
1. Outcome and system boundary
Start with the decision or task that must end better: resolve a request, update CRM, reconcile a document, recommend an action or execute an operating step. Require baseline, volume, exceptions, owners and economic result. The design must show where AI recommends, a rule decides, a human approves and a system records. Without an explicit boundary, scope expands while accountability and acceptance remain vague.
2. Data and context contracts
Request an inventory of sources, owners, purpose, quality, freshness, sensitive fields and retention. Each integration needs a schema contract, validation, versioning and handling for missing, late or conflicting data. The partner should explain how unauthorized content is excluded from context and how outputs reconcile with the system of record. Available data is not automatically trustworthy or permitted data.
3. Identity, authorization and approval
The solution must act under the correct identity and least privilege. Evaluate authentication, object- and function-level authorization, separation of duties, machine credentials, consent and human approval. OWASP’s API Security Top 10 highlights broken authorization and authentication. Ask for negative tests: what happens when a user requests another customer’s object, attempts a privileged function or induces an agent to call a prohibited tool?
4. Orchestration, APIs and compatibility
Require an end-to-end diagram covering models, gateways, queues, APIs, connectors, rules, systems of record and channels. Review rate limits, idempotency, timeouts, retries, duplication, versioning and transaction compensation. The partner must maintain an API inventory and assess consumed services; OWASP also flags unsafe API consumption and improper inventory. Prefer replaceable adapters over direct coupling to one model vendor.
5. Reliability, fallback and recovery
Define behavior for model outage, high latency, invalid output, API failure, partial write and data conflict. Require circuit breaking, queues, read-only mode, manual routing, reconciliation and rollback where appropriate. A workflow that succeeds only on the happy path is a demo. Acceptance tests should inject failures and prove no material action becomes invisible or ownerless.
6. Observability, evaluation and economics
Uncorrelated logs are insufficient. Require correlation across intent, context, model call, tool, enterprise system, approval and outcome while protecting sensitive data. OpenTelemetry frames observability through traces, metrics and logs; the proposal should specify signals for quality, latency, error, cost and impact. Combine output evaluation with process metrics and total cost per result, including review, support and reprocessing.
7. Ownership, operations and transfer
Separate code, connectors, prompts, evaluations, data, models, licenses and managed services. Require repository access, documentation, infrastructure as code, runbooks, telemetry access, backlog, training, support, SLAs and exit. NIST SP 800-218A extends secure practices across the development lifecycle, which implies versions, tests and responsibilities that survive the original delivery team. The buyer must be able to operate, audit and replace components.
The finalist break test
Give every finalist the same scenario: CRM adds a required field, an API becomes slow, a credential loses permission and the model attempts to repeat a write. Ask for the architecture, event sequence, controls, telemetry, user message, reconciliation and accountable owner. The response reveals whether the company understands distributed integration or merely connects a model to an API in a demo.
Run a vertical pilot before expansion
Choose one end-to-end journey with controlled value and risk. Include common cases, exceptions, governed real data, one system write and a human fallback. Freeze criteria for quality, authorization, latency, availability, cost and reconciliation. NIST’s Generative AI Profile recommends measuring and managing capabilities, limits and impacts; the pilot should turn that guidance into evidence to stop, redesign or scale.
Compare proposals on normalized assumptions
Normalize volume, integrations, environments, data, models, availability, support and responsibilities. Compare discovery, build, licenses, consumption, observability, security, maintenance and change costs. Treat a fixed price without technical inventory—and a schedule that ignores access and system owners—as risk transfer rather than certainty. Tie payments to data contracts, validated integration, negative tests, recovery, documentation and operating acceptance.
Disqualifying signals
Reject vendors that request broad credentials, skip authorization tests, omit tool-call records, treat retries as an implementation detail, cannot explain data sent to third parties, depend on a manual environment that cannot be reproduced, or withhold export and transition. It is also a warning when a provider promises to replace legacy systems before mapping the rules and exceptions they encode.
Next step
Use this framework with MAKINAI’s workflow-automation guide at https://makinai.co/insights/en/how-to-choose-ai-workflow-automation-company, agent-company evaluation at https://makinai.co/insights/en/how-to-evaluate-ai-agent-development-company and contract/SOW guide at https://makinai.co/insights/en/what-to-include-ai-services-contract-sow. For outcome-led integration of agents, automation and enterprise systems, visit https://makinai.co/services/en/ai-agents-automation-development.