All insights
EN · AI Strategy & Transformation

How to verify AI consulting case studies and client references before hiring

Turn case studies and testimonials into a verifiable chain of scope, baseline, team, delivery, outcome, failures and independent client evidence.

A closed portfolio is decomposed into seven evidence modules before passing through a hiring gate.
AI Delivery Evidence Chain-7 turns case studies and references into comparable proof before hiring. · Generated with OpenAI

Do not validate an AI firm by the client logo, final slide or a standalone percentage. Rebuild each case as a chain: problem and baseline; scope actually delivered; accountable team; data, models and integrations; measurement method; outcome and countermetrics; failures, recovery and current state. Then confirm critical facts with references selected for relevance by the buyer—not only the provider's happiest contact.

The best case is not necessarily the most famous. It is the closest to your decision in risk, data, integration, users, scale and operating stage. FAR 15.305 treats past performance as an indicator while requiring relevance, recency, source, context and trends. Apply the same discipline to AI services: compare contextual evidence, not logos.

The AI Delivery Evidence Chain-7

Score each domain from zero to four: zero is absent; one is a marketing assertion; two is partial evidence; three is contextual, verifiable evidence; four is confirmed through artifacts and an independent reference. Require at least 21 of 28, no zero and passage of five mandatory gates.

  • Relevance — comparable sector, decision, user, risk, data, integrations, volume, stage and constraints.
  • Baseline — prior state, period, data source, minimum quality and change hypothesis.
  • Scope and responsibility — what the firm decided, built, integrated and operated versus buyer and third-party work.
  • Delivered team — roles, seniority, allocation, subcontracting and match to the proposed team.
  • Technical and operating path — models, data, evaluation, security, adoption, observability, incidents and transfer.
  • Verifiable outcome — formula, period, denominator, countermetrics, total cost, attribution and current state.
  • Independent reference — sponsor, operator and technical owner who confirm facts, difficulty, behavior and continuity.

Five mandatory gates

  • The provider cannot separate its contribution from the buyer, platform or another partner.
  • The outcome lacks an explainable baseline, period, denominator, source or attribution method.
  • The team featured in the case is not the proposed team and substitution is not disclosed.
  • The reference confirms satisfaction but cannot validate scope, operations, failures or outcomes.
  • The provider blocks independent verification and offers only anonymous cases without alternative artifacts.

Request a two-page evidence sheet for each case

Include context, decision, baseline, scope, team, architecture, dependencies, evaluation, timeline, operating cost, outcome, countermetrics, failures, current status, available artifacts and people who can confirm. Remove confidential material while preserving definitions and causal relationships. Treat a metric without a denominator or attribution method as unverified.

CPARS evaluates quality using reliability, failure data, user comments, acceptance rates and rework. That range matters for AI: do not accept only accuracy, productivity or revenue. Ask for quality, adoption, cost per outcome, incidents, human review and operating load.

Run three references, not one institutional call

  • Executive sponsor: decision, value, governance, transparency and willingness to rehire.
  • Operating owner: adoption, quality, exceptions, support, incidents, cost and ability to operate without excessive dependency.
  • Technical or data owner: architecture, integration, security, evaluation, documentation and transfer quality.

GSA guidance recognizes tailored questionnaires, interviews and other sources for collecting past-performance information. Standardize the questions and record who answered, their relationship to the work and what they could verify. Do not request confidential information or bypass contractual restrictions.

Ten questions that outperform a testimonial

  • What was the baseline and who measured it?
  • What did this team actually deliver?
  • What had to be redesigned after work began?
  • What was the most serious failure and how did the firm respond?
  • Which outcome remained after six months?
  • Which operating costs or tasks were omitted from the case?
  • Did the promised team remain assigned?
  • How were security, privacy and model changes handled?
  • Can the buyer operate, audit and switch partners using the documentation?
  • Would you hire the same team again for a similar situation?

Normalize evidence before comparing finalists

Rate each case for relevance, evidence confidence and the provider's actual contribution. Do not automatically penalize a newer firm without famous logos: allow proof through smaller work, artifacts, the proposed team and a paid validation. World Bank guidance combines quality and value criteria with price; use references to reduce uncertainty, not preserve incumbents.

Connect cases, team, proposal and pilot

Use https://makinai.co/insights/en/how-to-evaluate-ai-consulting-proposals-scorecard to normalize proposals, https://makinai.co/insights/en/how-to-evaluate-ai-consulting-team-before-hiring to verify the delivery team and https://makinai.co/insights/en/how-to-run-paid-ai-pilot-before-hiring-partner to create evidence in your environment. Explore https://makinai.co/services/en/ai-strategy-transformation-consulting.

When to involve MAKINAI

MAKINAI can turn cases into comparable evidence sheets, facilitate structured reference interviews and convert gaps into RFP questions, finalist sessions or a paid pilot. The recommendation should record what was confirmed, what remains a hypothesis and which risk the buyer accepts.

Sources and references

  1. FAR 15.305 — Proposal Evaluation · U.S. Acquisition.gov

    Requires consideration of relevance, recency, source, context and performance trends when assessing past performance.

    2026-09-02
  2. GSA 570.306 — Evaluating Offers · U.S. General Services Administration

    Recognizes tailored questionnaires, interviews and other sources as ways to obtain past-performance information.

    2026-09-02
  3. CPARS — Quality Evaluation Area · U.S. Contractor Performance Assessment Reporting System

    Connects quality to project objectives, reliability, failure data, user comments, acceptance and rework.

    2026-09-02
  4. World Bank — Rated Criteria · World Bank

    Supports evaluating value, quality, sustainability and non-price attributes alongside price and lifecycle cost.

    2026-09-02
  5. UK Government — Sourcing Playbook · UK Government

    Structures sourcing decisions, bid evaluation, due diligence and supplier-performance management.

    2026-09-02
Making connections

Continue exploring

AI Strategy & Transformation

How to define AI provider governance and performance management before hiring

Read insight
AI Strategy & Transformation

How to evaluate an AI consulting ROI business case before hiring

Read insight
AI Strategy & Transformation

Boutique AI firm, global consultancy, or systems integrator: how to choose

Read insight
Related capability

AI strategy & transformation

An AI transformation consultancy should answer four questions before recommending technology: where business value exists, which capabilities and data are required, how risk will be controlled, and who will operate the change. MAKINAI connects those answers in an executable plan with priorities, owners, metrics and scale decisions.

Explore this capability