All insights
EN · AI Strategy & Transformation

How to evaluate an AI solution architecture proposal before hiring a provider

Compare decisions, evidence, operating risk, economics, and reversibility—not diagrams or technology name-dropping.

AI architecture proposals pass through eight evidence controls before converging into one governable system.
The AI Architecture Decision Proof-8 compares decisions, evidence, operations, economics, and reversibility—not diagram polish. · Generated with OpenAI

The strongest AI architecture proposal is not the one with the most boxes, models, or products. It explains why each decision serves the outcome, which assumptions remain unproved, how the system fails and recovers, what it costs to operate, and how components can be replaced. Before award, compare decision records and executable evidence—not diagram polish.

Architecture should be proportionate to the decision. An internal assistant with human review may need a simple integration and a managed service. An agent taking consequential actions, handling sensitive data, or operating at high volume needs stronger controls, evaluation, observability, and shutdown paths. Complexity without corresponding value or risk reduction is a liability.

The AI Architecture Decision Proof-8: 32 points

Score each dimension from zero to four: zero is absent; one is an assertion; two is a partial design; three is a justified decision with evidence and an owner; four is a decision reproduced under representative conditions, with an alternative and review trigger. As an illustrative gate, require 24 of 32, no zero, and at least three for data, security, evaluation, and operations.

  • Outcome and boundaries — user, task, autonomy, non-goals, baseline, tolerable error, and non-AI alternative are explicit.
  • Data and context — sources, rights, quality, retrieval, memory, refresh, retention, and provenance support real use.
  • Models and tools — managed services, RAG, tools, fine-tuning, or custom code are compared against evidence.
  • Integration and interfaces — systems of record, APIs, events, identity, limits, idempotency, and failure handling are defined.
  • Security and supply chain — environments, secrets, access, components, subcontractors, versions, vulnerabilities, and incidents are controllable.
  • Evaluation and oversight — sets, metrics, adverse cases, human review, approval, regression, and stop criteria are executable.
  • Reliability and operations — observability, SLOs, fallback, rollback, model change, continuity, support, and retirement have owners.
  • Economics and exit — cost per outcome, capacity, concentration, portability, formats, documentation, and transition are demonstrable.

Apply five kill criteria before scoring

  • The design depends on data, access, or an integration that nobody has verified.
  • A consequential or irreversible action lacks proportionate authorization, limits, and human intervention.
  • There is no representative evaluation set or reproducible acceptance criterion.
  • A critical third-party component lacks ownership, monitoring, contingency, or a replacement right.
  • The provider cannot export the code, configuration, data, prompts, evaluations, logs, and documentation needed for transition.

Failing a gate does not necessarily kill the use case. It changes the purchase to a capped technical discovery, integration spike, independent evaluation, or controlled pilot. Do not award full implementation to fund the late discovery of a critical architectural assumption.

Compare four patterns without chasing fashion

  • Managed service and configuration — faster delivery and less buyer-operated infrastructure; clarify limits, data use, pricing, availability, and exit.
  • RAG and enterprise integration — can ground outputs in authorized content; quality depends on ingestion, retrieval, permissions, evaluation, and refresh.
  • Agents and orchestration — can execute workflows and tools; they expand action surface, exceptions, identity, observability, and control needs.
  • Custom components or fine-tuning — may address specific domain, performance, or scale needs; require sustainable data, evaluation, MLOps, capacity, and cost.

Require seven comparable artifacts

  • Context diagram showing users, systems, data, third parties, and trust boundaries.
  • Decision records for critical choices, rejected alternatives, evidence, owner, and review trigger.
  • Threat model and component supply chain, including versions, access, and contingencies.
  • Evaluation plan tied to requirements, risk, deployment conditions, and acceptance criteria.
  • Operating model for telemetry, SLOs, incidents, changes, fallback, rollback, and retirement.
  • Cost and capacity model by scenario, including volume, latency, human review, and uncertainty.
  • Portability and transition package with formats, repositories, documentation, dependencies, and a replacement test.

Run a 90-minute adversarial architecture defense

Give finalists the same brief. Ask them to defend three decisions and one rejected alternative. Then remove the primary model, degrade retrieval, change an API schema, spike volume, and require data deletion. Demand diagnosis, isolation, fallback, cost and schedule impact, and a decision record. Score assumption clarity, component substitutability, and behavior under uncertainty—not how quickly the team draws more boxes.

Commercial and technical red flags

  • A technology appears before the problem and requirements.
  • Every choice is called a standard or best practice without comparison.
  • Integration, evaluation, observability, human review, and support sit outside the price.
  • The diagram omits identity, environments, sensitive data, third parties, or operations.
  • The solution depends on one model but has no replacement test.
  • Scale claims lack a load model, bottleneck analysis, and cost per outcome.
  • The buyer receives documentation only at the end.

U.S. buying context

For U.S. buyers, map the actual deployment and data path against applicable state, federal, sector, customer, employment, records, accessibility, export, and confidentiality obligations. Do not assume a U.S. contracting entity means U.S.-only delivery or processing. Price cloud, model, data-transfer, security, human-review, support, and internal integration costs in the same scenario model. Legal, privacy, security, enterprise architecture, and procurement should review one controlled architecture baseline.

Put architecture into the RFP and contract

Give every bidder the same scenario, constraints, and volumes. Require facts to be separated from assumptions, architectural options, buyer dependencies, acceptance criteria, and scenario pricing. Tie milestones to reproducible artifacts, reserve review rights when context or models change, and define ownership, licenses, repositories, export, transition assistance, and subcontracting limits.

Connect architecture to adjacent buying decisions

Use https://makinai.co/insights/en/buy-configure-or-build-ai-solution to set the technology boundary, https://makinai.co/insights/en/assess-data-readiness-before-hiring-ai-company to test inputs, https://makinai.co/insights/en/security-due-diligence-ai-services-company to deepen controls, and https://makinai.co/insights/en/how-to-assess-ai-vendor-lock-in-exit-plan to validate portability and exit.

When to involve MAKINAI

MAKINAI can turn architecture proposals into comparable decisions, run the adversarial defense, and convert assumptions into discovery scope, gates, acceptance criteria, and a transition plan. Explore https://makinai.co/services/en/ai-strategy-transformation-consulting.

Sources and references

  1. NIST AI RMF Core · National Institute of Standards and Technology

    Connects context, requirements, third parties, evaluation, monitoring, and retirement to documented lifecycle decisions.

    2026-09-10
  2. NIST SP 800-218 — Secure Software Development Framework · National Institute of Standards and Technology

    Structures secure development practices that buyers can embed in delivery and verify through evidence rather than assertions.

    2026-09-10
  3. NIST SP 800-161 Rev. 1 — Cybersecurity Supply Chain Risk Management · National Institute of Standards and Technology

    Guides identification, assessment, and monitoring of supplier, component, and dependency risks across the system lifecycle.

    2026-09-10
  4. UK Government — Artificial Intelligence Playbook · UK Government

    Recommends modular architecture, interoperability, security, suitable data, evaluation, and human oversight proportionate to use.

    2026-09-10
  5. UK Government — Digital, Data and Technology Playbook · UK Government Commercial Function

    Supports evaluation before purchase, open standards, interoperability, portability, and whole-life cost and risk visibility.

    2026-09-10
Making connections

Continue exploring

AI Strategy & Transformation

How to assess data readiness before hiring an AI implementation company

Read insight
AI Strategy & Transformation

Local, nearshore, or offshore AI services company: how to choose

Read insight
AI Strategy & Transformation

How to define AI provider governance and performance management before hiring

Read insight
Related capability

AI strategy & transformation

An AI transformation consultancy should answer four questions before recommending technology: where business value exists, which capabilities and data are required, how risk will be controlled, and who will operate the change. MAKINAI connects those answers in an executable plan with priorities, owners, metrics and scale decisions.

Explore this capability