Choose an AI evaluation company by whether it can prove three things: its tests represent real decisions and users; failures can be reproduced and explained; and findings produce a clear launch, remediation or stop decision. A long benchmark catalog, a red-team tool or a polished dashboard is not sufficient evidence of evaluation quality.
Scope must cover the complete system—model, retrieval, tools, permissions, interfaces, human review and operations—not only the base model. NIST released TEVV-Athlon on August 4, 2026 as an adaptable framework for machine learning, LLM, multimodal and agentic systems. That breadth matters because consequential failures often emerge at component boundaries.
The MAKINAI AI Evaluation Proof-8
Require evidence in the eight areas below and score each from zero to four: zero is missing; one is a generic promise; two is a described method; three is a reviewable artifact; four is evidence reproduced in your environment. Set a threshold of 24 out of 32, require no zero, and apply all mandatory gates. Lock any weights before proposals arrive.
- Decision and boundary — decisions informed, users, prohibited uses and failure impact.
- Representative data — coverage, provenance, relevant cohorts, rare cases, synthetic data and blind holdout.
- Quality and utility — task metrics, baseline, uncertainty, critical errors and human judgment.
- Robustness and security — misuse, prompt injection, leakage, distribution shifts, dependencies and tools.
- Human factors — automation bias, comprehension, accessibility, escalation and operating burden.
- Reproducibility and independence — versions, seeds, logs, judgments, conflicts and third-party reruns.
- Continuous operations — telemetry, limits, alerts, regression, incidents and reevaluation triggers.
- Transfer and exit — test data, code, reports, limitations, training and portability.
Five mandatory gates
Stop the selection if a provider cannot connect metrics to a business decision; cannot control test-data provenance and representativeness; cannot reproduce results with versions and records; tests only the model while ignoring tools, permissions and people; or omits production monitoring and reevaluation triggers. A strong average cannot offset one of these structural gaps.
Address conflicts of interest explicitly. The builder has architectural context and may remediate faster, but it can be incentivized to accept favorable evidence. An independent evaluator improves contestability but may miss operational context. For high-impact decisions, combine builder documentation with independent execution of the acceptance suite.
What a technically credible proposal should deliver
Require a versioned evaluation plan, risk map, system inventory, test-set definitions, acceptance criteria, human-review protocol, executable records, failure report and prioritized remediation backlog. Findings should distinguish confirmed defect, known limitation, unresolved uncertainty and accepted risk. An aggregate chart without traceable examples is not enough.
NIST's adversarial machine learning taxonomy can structure attacks and mitigations across the lifecycle. The UK AI Security Institute's Inspect framework demonstrates repeatable evaluations for reasoning, agentic tasks, knowledge, behavior and multimodality. A provider need not use these exact tools, but it should demonstrate equivalent capabilities and explain where automation is inadequate.
Public benchmarks, private data and human evaluation
Public benchmarks help with broad screening, but they may be contaminated by training data, reward narrow optimization and miss enterprise workflows. Private datasets increase operational relevance but require governance, sampling and access controls. Blind, sequestered holdouts, as used in NIST's AITE initiative, reduce the opportunity for test-specific tuning.
Automation increases coverage and repeatability; human evaluators capture usefulness, ambiguity, impact and escalation quality. Define rubrics before testing, train evaluators, measure agreement and retain disagreement examples. For critical outcomes, combine automated measures, structured human judgment and domain-expert review.
Run a paid finalist bake-off
Invite two finalists to the same two-to-three-week paid exercise using a controlled package: simplified architecture, synthetic data, one blind holdout and four injected failures—an unavailable source, a revoked permission, malicious content and a distribution shift. Do not request unpaid audits. The goal is to compare method, traceability and communication, not obtain a complete assurance program for free.
Have each team rerun a sample 48 hours later and explain variance. Score time to reproduce, diagnostic quality, severity judgment, business impact, recommended correction and candor about uncertainty. The best team is not the one that reports the largest raw defect count; it is the one that separates material risk from noise and supports a defensible decision.
Make the trade-offs explicit
- Independent versus builder-integrated evaluation: stronger contestability versus faster remediation.
- Breadth versus depth: more scenarios versus rigorous testing of critical paths.
- Prelaunch versus continuous testing: a release gate versus detection of production change.
- Live environment versus sandbox: higher fidelity versus stronger privacy, safety and control.
- Executive report versus technical artifacts: decision clarity versus reproducibility and transfer.
Contracting, governance and continuity
The SOW should name systems, versions, environments, permitted data, accountable roles, severity levels, remediation timelines, retesting and artifact ownership. Define who can stop testing and who accepts residual risk. GAO's accountability framework reinforces that performance and monitoring continue after launch: changes to models, data, prompts, policies or integrations should trigger proportionate reevaluation.
To limit the initial commitment, use MAKINAI's paid-pilot guide at https://makinai.co/insights/en/how-to-run-paid-ai-pilot-before-hiring-partner. For accountability design, review https://makinai.co/insights/en/how-to-choose-ai-governance-consulting-firm and, for ongoing operations, https://makinai.co/insights/en/how-to-choose-managed-ai-services-provider.
When to involve MAKINAI
MAKINAI can help define the acceptance matrix, prepare the bake-off and connect technical evidence to an executive decision. The goal is not a generic certification; it is an evaluation process your organization can understand, repeat and use to govern the system after contracting.