Direct answer: choose an AI agent development company based on evidence that it can constrain actions, evaluate outcomes and trajectories, protect data, operate failures, control unit economics and transfer knowledge. A polished demonstration proves only that a happy path can run. Before contracting, require a bounded use case, acceptance tests, a permission architecture, an evaluation set, an incident plan and explicit ownership terms.
AI agents are not simply chatbots with better interfaces. They retrieve context, select tools, perform multi-step work and may change systems of record. Vendor risk therefore extends beyond an inaccurate sentence. It includes an unauthorized action, an overprivileged credential, a workflow stuck in a loop and a process whose usage cost scales unpredictably. A credible partner treats these as business and engineering requirements from day one.
The MAKINAI Agent Production Readiness Matrix
Use six gates. A provider must produce minimum evidence in every gate; strong prototyping cannot compensate for a critical security or operational gap. The matrix works for early screening, a paid proof of value and production acceptance. It also makes proposals comparable without forcing every bidder onto the same model or cloud.
Gate 1 — Outcome, autonomy and action boundaries
Start with the operational outcome, not the model. Ask the provider to define the triggering event, permitted decisions, available tools, source data, stopping condition and human approval points. Establish reversible actions, spending limits, prohibited categories and behavior when information is missing. A mature design separates recommendation, preparation and execution: an agent may recommend a refund, prepare it or issue it, and each level needs different controls.
- Evidence to request: journey and tool map; role-based permission matrix; human-confirmation policy; success and stopping criteria; examples of safe failure behavior.
Gate 2 — Data, context and architecture
The partner should explain where context comes from, how it is refreshed and how outputs are tied to the right source. Ask which data enters prompts, memory, logs and external systems; how personal or confidential information is minimized; and how credentials are separated. Favor architectures where every tool has an explicit contract, input and output validation, least privilege and a timeout. The system should allow models or components to change without rebuilding the entire operation.
Request a dependency view covering the model, orchestration, retrieval, APIs, queues, databases, observability and human interfaces. This exposes lock-in, single points of failure and costs that a demo can hide. Google Cloud’s production guidance similarly treats agent architecture and deployment as a system design problem rather than a prompt-only exercise.
Gate 3 — Outcome and trajectory evaluation
Judging only the final answer is not enough. An agent can reach the right result while using the wrong tool, bypassing an approval or making ten times the necessary calls. The provider should maintain representative cases, expected outcomes, acceptable trajectories, regression tests and adversarial examples. Metrics should combine task success, factual accuracy, correct tool use, safety, latency, cost and human-escalation rate.
Require a baseline before the pilot and a reproducible report after every material change. NIST structures AI risk work around Govern, Map, Measure and Manage. That logic helps buyers connect technical tests to ownership and operating decisions instead of treating evaluation as one abstract score.
Gate 4 — Security, identity and incident response
Agents expand the attack surface because they interpret content and act through tools. OWASP now provides guidance specifically for agentic applications, while MITRE ATLAS catalogs adversary tactics and techniques against AI systems. A qualified firm should demonstrate threat modeling, defenses against malicious instructions in user input or retrieved documents, separation of data from commands, secrets management, user-bound authentication, action-level authorization, auditable records and an emergency shutdown procedure.
- Disqualifying questions: Can the agent act with a shared credential? Are high-impact actions approved? Can logs reconstruct the sequence? Is indirect prompt injection tested? Who contains and communicates an incident?
Gate 5 — Operations, observability and change
Production starts where the prototype ends. Require dashboards for availability, latency, cost, tool errors, loops, human interventions and quality by task type. Ask how prompt, tool, policy and model versions are recorded and rolled back. The company should propose gradual rollout, usage limits, review queues and fallback to a human process or deterministic automation. Otherwise, the buyer inherits an opaque system that is difficult to sustain.
The support plan should name owners, severity levels, coverage hours, response targets and escalation criteria. It must also cover changes in APIs, models and source data. A service-level agreement is useful only when it measures the end-to-end workflow, not merely model availability.
Gate 6 — Economics, ownership and internal capability
Compare total cost per successfully completed task, not only token price or developer day rate. Include inference, retrieval, storage, observability, human review, integration, support and rework. Set a budget per workflow and anomaly alerts. The proposal must clarify rights to code, prompts, evaluations, connectors, derived data and documentation, plus export and transition procedures.
Ask for enablement through runbooks, technical and operational training, acceptance criteria and an assisted handover period. The objective is not to remove the partner. It is to prevent unnecessary dependency and ensure your organization can govern the capability it owns.
How to score proposals without rewarding the best demo
Use 100 points: 20 for outcome and boundaries, 15 for data and architecture, 20 for evaluation, 20 for security, 15 for operations, and 10 for economics and ownership. Add pass/fail gates: no sensitive action without identity and authorization; no release without regression tests; no critical workflow without logs, shutdown and fallback; and no contract without ownership terms. Then run a short pilot using the same case set for finalists.
Red flags during vendor selection
- The proposal starts with a model choice but never defines the process; the demo uses curated data and hides failures; “accuracy” appears without a dataset or rubric; security is postponed until after the pilot; cost is estimated without volume and exception rates; the vendor will not deliver logs, evaluations or documentation; every change depends on one person or proprietary platform.
Next step
Before requesting prices, turn the use case into an evaluation pack with permitted actions, data, risks, test cases and metrics. MAKINAI’s RFP guide can structure the competition: https://makinai.co/insights/en/how-to-write-rfp-ai-services. For broader implementation capability, use https://makinai.co/insights/en/how-to-choose-ai-implementation-company-brazil-scorecard. If you need to design and build a first agent with production controls from the beginning, see MAKINAI’s AI agents and automation service: https://makinai.co/services/en/ai-agents-automation-development.