All insights
EN · Websites, Platforms & Digital Products

How to choose an agency for an AI-ready enterprise website

A scorecard for evaluating strategy, UX, content, architecture, trust, measurement and operations before commissioning an AI-enabled website redesign.

Seven layers connect journey, content, interface, architecture, trust, measurement and operations in an AI-ready website.
AI-Ready Experience Proof-7 evaluates the experience as an operable system, not an isolated chatbot. · Generated with OpenAI

Direct answer: choose an agency that can prove the full experience system, not just screens or a chatbot demo. Require evidence across seven dimensions: outcome and journey, content and knowledge, interaction and human escalation, architecture and integrations, accessibility and trust, evaluation and measurement, and operations and ownership. Before commissioning a rollout, buy a paid discovery using real users, content, technical constraints and acceptance criteria. The winning proposal should explain how the site improves a valuable task, fails safely and transfers day-to-day control to your team.

An AI-ready website is not a homepage with a chat box. It may use semantic search, recommendations, assisted generation, personalization, agents or conversational interfaces, but it remains a digital service. It must help people find, understand, decide and act; respect permissions; work when a model or integration fails; and generate defensible evidence of value. A polished portfolio and familiarity with one model are therefore insufficient vendor-selection signals.

AI-Ready Experience Proof-7

Score each dimension from 0 to 4: absent, described, demonstrated, validated with users or operated in a comparable context. The maximum is 28. Do not let a high average compensate for disqualifying gaps in privacy, security, accessibility, content rights or fallback. Give every finalist the same evidence request so you compare capability rather than presentation style.

1. Outcome and journey

The agency should start with a user task and business result: finding a trustworthy answer, qualifying an opportunity, choosing a solution, resolving a problem or completing a transaction. Ask for the current journey, baseline, friction, segments and change hypothesis. Google PAIR centers AI product design on user needs and a definition of success; AI belongs only where it improves the task relative to a conventional experience.

2. Content, knowledge and data

Ask where facts, offers, policies, catalog data, evidence and instructions originate. Require an inventory, editorial authority, freshness, permissions, taxonomy, citable sources and handling for missing or conflicting content. The agency must separate public content, personal data and restricted knowledge while showing how the CMS, DAM, CRM, analytics and retrieval layer remain governable. Without that foundation, the interface merely accelerates inconsistent answers.

3. Interaction, explanation and human escalation

Evaluate how the team designs for intent, uncertainty, confirmation and control. The experience should state what it can do, confirm consequential actions, explain limits, preserve navigable alternatives and offer a clear human route. Test people who reject personalization, make ambiguous requests, switch language or correct an answer. The objective is calibrated trust, not an interface that performs humanness.

4. Architecture, performance and integrations

Request a diagram of the complete path: browser, experience layer, CMS, search or RAG, models, rules, tools, identity, CRM, commerce, observability and fallback. Require targets for speed, availability, unit cost and controlled degradation. NIST SP 800-218A adds AI-specific practices across secure development, which should translate into versioning, tests, known dependencies, separated environments and a vulnerability-response process.

5. Accessibility, privacy and trust

Accessibility belongs in discovery, design, content, components and testing, not a late audit. WCAG 2.2 supplies testable criteria, but conformance also needs informed human evaluation with assistive technologies. Request a data map, consent and preference design, retention rules, treatment of user inputs, disclosures and claims review. The proposal should name the owner and acceptance evidence for every control.

6. Evaluation and measurement

Freeze a task set and thresholds before viewing results. Combine task completion, factual quality, retrieval, accessibility, latency, cost, assisted conversion, satisfaction, escalation and avoided harm. Include common cases, edge cases, relevant languages and adversarial attempts. NIST’s Generative AI Profile treats measurement and management as continuing activities, so a vendor-selected demo cannot replace reproducible tests on representative samples.

7. Operations, ownership and transfer

Define who updates content, prompts, evaluations, integrations and policy; monitors incidents; and approves model changes. Separate buyer assets, agency components and third-party services in the contract. Require documentation, exports, log access, recurring-cost assumptions, support, training and exit. Prefer a partner that reduces dependency over time while remaining accountable for delivery.

The finalist break test

Give each finalist the same scenario: a material policy changes, the authorized source is still stale, the CRM becomes unavailable and a user requests an unauthorized action. Ask for the journey, architecture, fallback behavior, user message, recorded evidence and correction plan. The exercise reveals whether strategy, design, content and engineering operate as one system or disconnected workstreams.

Buy discovery before rollout

Commission a bounded phase to understand users, content, baseline, constraints, architecture and risk; prototype alternatives; and decide whether AI is necessary. The UK Government Service Manual frames discovery as learning about the problem and deciding whether to proceed. Deliverables should include the journey, content and integration inventory, tested prototype, architecture, backlog, acceptance criteria, risk register and revised estimate.

Disqualifying signals

Reject proposals that promise personalization without explaining data and consent; hide model costs; omit accessibility; cannot maintain a useful non-AI path; test only curated demo content; prevent source auditing; or require non-exportable assets to operate. It is also a warning when discovery is contractually guaranteed to become a rollout regardless of evidence.

Make the selection

Check non-compensable gates before adding scores. Among qualified finalists, compare hypothesis clarity, evidence quality, integrated design and engineering capability, economic transparency and transfer plan. Set acceptance by task and risk rather than feature count. Tie payments to verifiable deliverables and retain an explicit proceed, redesign or stop decision.

Next step

Use this scorecard with MAKINAI’s RFP guide at https://makinai.co/insights/en/how-to-write-rfp-ai-services, AI product partner guide at https://makinai.co/insights/en/how-to-choose-ai-product-development-company and customer-service AI guide at https://makinai.co/insights/en/how-to-choose-ai-customer-service-company. To combine brand, content, UX and technology in an AI-ready experience, visit https://makinai.co/services/en/branding-creative-digital-experience-agency.

Sources and references

  1. Google PAIR — People + AI Guidebook · Google PAIR

    Provides practical guidance for designing useful, human-centered AI products around user needs, success, mental models, trust and feedback.

    2026-08-25
  2. W3C — Web Content Accessibility Guidelines (WCAG) 2.2 · W3C

    Defines testable, technology-neutral success criteria for making web content more accessible.

    2026-08-25
  3. NIST — Secure Software Development Practices for Generative AI · NIST

    Extends secure software development practices with AI-specific tasks and considerations across the lifecycle.

    2026-08-25
  4. NIST — Generative AI Profile · NIST

    Recommends governance, measurement and management practices for generative-AI risks and impacts.

    2026-08-25
  5. UK Government — How the discovery phase works · UK Government Digital Service

    Frames discovery around understanding the problem, users, constraints and whether to proceed before committing to delivery.

    2026-08-25
Making connections

Continue exploring

Websites, Platforms & Digital Products

How to choose an AI product development company

Read insight
AI Strategy & Transformation

How to define AI provider governance and performance management before hiring

Read insight
AI Strategy & Transformation

How to evaluate an AI consulting ROI business case before hiring

Read insight
Related capability

Products, agents & automation

Building an enterprise AI agent is not just connecting a model to a chat interface. It requires product design, context, tools, integrations, identity, evaluation, guardrails, observability and human operations. MAKINAI builds the complete experience and measures whether it improves capability, quality or speed.

Explore this capability