All insights
EN · AI Strategy & Transformation

How to choose an AI evaluation and testing company

Select an AI evaluation firm by its ability to reproduce failures, test the complete system and connect technical findings to launch and operating decisions.

Editorial AI evaluation laboratory connects blind data, failure injection, evidence trails, human review and production monitoring.
AI Evaluation Proof-8 tests the complete system and turns reproducible failures into launch and operating decisions. · Generated with OpenAI

Choose an AI evaluation company by whether it can prove three things: its tests represent real decisions and users; failures can be reproduced and explained; and findings produce a clear launch, remediation or stop decision. A long benchmark catalog, a red-team tool or a polished dashboard is not sufficient evidence of evaluation quality.

Scope must cover the complete system—model, retrieval, tools, permissions, interfaces, human review and operations—not only the base model. NIST released TEVV-Athlon on August 4, 2026 as an adaptable framework for machine learning, LLM, multimodal and agentic systems. That breadth matters because consequential failures often emerge at component boundaries.

The MAKINAI AI Evaluation Proof-8

Require evidence in the eight areas below and score each from zero to four: zero is missing; one is a generic promise; two is a described method; three is a reviewable artifact; four is evidence reproduced in your environment. Set a threshold of 24 out of 32, require no zero, and apply all mandatory gates. Lock any weights before proposals arrive.

  • Decision and boundary — decisions informed, users, prohibited uses and failure impact.
  • Representative data — coverage, provenance, relevant cohorts, rare cases, synthetic data and blind holdout.
  • Quality and utility — task metrics, baseline, uncertainty, critical errors and human judgment.
  • Robustness and security — misuse, prompt injection, leakage, distribution shifts, dependencies and tools.
  • Human factors — automation bias, comprehension, accessibility, escalation and operating burden.
  • Reproducibility and independence — versions, seeds, logs, judgments, conflicts and third-party reruns.
  • Continuous operations — telemetry, limits, alerts, regression, incidents and reevaluation triggers.
  • Transfer and exit — test data, code, reports, limitations, training and portability.

Five mandatory gates

Stop the selection if a provider cannot connect metrics to a business decision; cannot control test-data provenance and representativeness; cannot reproduce results with versions and records; tests only the model while ignoring tools, permissions and people; or omits production monitoring and reevaluation triggers. A strong average cannot offset one of these structural gaps.

Address conflicts of interest explicitly. The builder has architectural context and may remediate faster, but it can be incentivized to accept favorable evidence. An independent evaluator improves contestability but may miss operational context. For high-impact decisions, combine builder documentation with independent execution of the acceptance suite.

What a technically credible proposal should deliver

Require a versioned evaluation plan, risk map, system inventory, test-set definitions, acceptance criteria, human-review protocol, executable records, failure report and prioritized remediation backlog. Findings should distinguish confirmed defect, known limitation, unresolved uncertainty and accepted risk. An aggregate chart without traceable examples is not enough.

NIST's adversarial machine learning taxonomy can structure attacks and mitigations across the lifecycle. The UK AI Security Institute's Inspect framework demonstrates repeatable evaluations for reasoning, agentic tasks, knowledge, behavior and multimodality. A provider need not use these exact tools, but it should demonstrate equivalent capabilities and explain where automation is inadequate.

Public benchmarks, private data and human evaluation

Public benchmarks help with broad screening, but they may be contaminated by training data, reward narrow optimization and miss enterprise workflows. Private datasets increase operational relevance but require governance, sampling and access controls. Blind, sequestered holdouts, as used in NIST's AITE initiative, reduce the opportunity for test-specific tuning.

Automation increases coverage and repeatability; human evaluators capture usefulness, ambiguity, impact and escalation quality. Define rubrics before testing, train evaluators, measure agreement and retain disagreement examples. For critical outcomes, combine automated measures, structured human judgment and domain-expert review.

Run a paid finalist bake-off

Invite two finalists to the same two-to-three-week paid exercise using a controlled package: simplified architecture, synthetic data, one blind holdout and four injected failures—an unavailable source, a revoked permission, malicious content and a distribution shift. Do not request unpaid audits. The goal is to compare method, traceability and communication, not obtain a complete assurance program for free.

Have each team rerun a sample 48 hours later and explain variance. Score time to reproduce, diagnostic quality, severity judgment, business impact, recommended correction and candor about uncertainty. The best team is not the one that reports the largest raw defect count; it is the one that separates material risk from noise and supports a defensible decision.

Make the trade-offs explicit

  • Independent versus builder-integrated evaluation: stronger contestability versus faster remediation.
  • Breadth versus depth: more scenarios versus rigorous testing of critical paths.
  • Prelaunch versus continuous testing: a release gate versus detection of production change.
  • Live environment versus sandbox: higher fidelity versus stronger privacy, safety and control.
  • Executive report versus technical artifacts: decision clarity versus reproducibility and transfer.

Contracting, governance and continuity

The SOW should name systems, versions, environments, permitted data, accountable roles, severity levels, remediation timelines, retesting and artifact ownership. Define who can stop testing and who accepts residual risk. GAO's accountability framework reinforces that performance and monitoring continue after launch: changes to models, data, prompts, policies or integrations should trigger proportionate reevaluation.

To limit the initial commitment, use MAKINAI's paid-pilot guide at https://makinai.co/insights/en/how-to-run-paid-ai-pilot-before-hiring-partner. For accountability design, review https://makinai.co/insights/en/how-to-choose-ai-governance-consulting-firm and, for ongoing operations, https://makinai.co/insights/en/how-to-choose-managed-ai-services-provider.

When to involve MAKINAI

MAKINAI can help define the acceptance matrix, prepare the bake-off and connect technical evidence to an executive decision. The goal is not a generic certification; it is an evaluation process your organization can understand, repeat and use to govern the system after contracting.

Sources and references

  1. NIST — TEVV-Athlon Framework · NIST

    An August 4, 2026 framework for adaptable evaluation of ML, LLM, multimodal and agentic systems.

    2026-08-28
  2. NIST — Adversarial Machine Learning · NIST

    A taxonomy of attacks and mitigations across the AI lifecycle.

    2026-08-28
  3. UK AI Security Institute — Inspect · UK AI Security Institute

    An open framework for reproducible reasoning, agentic, knowledge, behavior and multimodal evaluations.

    2026-08-28
  4. GAO — AI Accountability Framework · U.S. Government Accountability Office

    Governance, data, performance and monitoring practices for AI systems.

    2026-08-28
  5. NIST — AI Test and Evaluation Challenge · NIST

    Blind, sequestered data can reduce contamination and improve evaluation objectivity.

    2026-08-28
Making connections

Continue exploring

AI Strategy & Transformation

How to define AI provider governance and performance management before hiring

Read insight
AI Strategy & Transformation

How to evaluate an AI consulting ROI business case before hiring

Read insight
AI Strategy & Transformation

Boutique AI firm, global consultancy, or systems integrator: how to choose

Read insight
Related capability

AI strategy & transformation

An AI transformation consultancy should answer four questions before recommending technology: where business value exists, which capabilities and data are required, how risk will be controlled, and who will operate the change. MAKINAI connects those answers in an executable plan with priorities, owners, metrics and scale decisions.

Explore this capability