Your data does not need to be perfect before you hire an AI company. It must be understood, reachable, authorized, and testable enough for bidders to estimate delivery, risk, and outcomes without pricing assumptions as change orders. Before award, prove the path from use case to a representative sample, with ownership, rights, quality, lineage, controls, and a refresh mechanism. If that path cannot be proved, buy a capped data discovery—not full implementation.
Do not ask whether the organization has enough data. Ask whether it has the right evidence for this decision under expected operating conditions. A large lake may be unusable because rights, coverage, or freshness are weak. A smaller collection may support a controlled pilot if it represents real work and enables repeatable evaluation.
The AI Data Readiness Evidence-8: 32 points
Score each proof from zero to four: zero is unknown; one is an assertion; two is a partial sample or process; three is consistent evidence with an owner and gap plan; four is evidence reproduced with a finalist. As an illustrative gate, require 24 of 32, no zero, and at least three for rights, security, and representativeness before full implementation. Weight the dimensions to the use-case risk.
- Outcome and decision — task, user, action, baseline, tolerable error, and measurable value are defined.
- Sources and ownership — systems of record, owners, contracts, third parties, and responsibilities are known.
- Access and permissions — purpose, authority, environments, retention, deletion, and use restrictions are approved.
- Quality and profile — completeness, validity, duplicates, timeliness, consistency, and outliers are measured by segment.
- Coverage and representativeness — periods, regions, channels, groups, exceptions, and adverse conditions match intended use.
- Lineage and provenance — origin, transformations, versions, labels, synthetic content, and use constraints are traceable.
- Security and privacy — classification, minimization, segregation, access, logging, de-identification, and incident response work.
- Pipeline and operations — ingestion, refresh, monitoring, drift, correction, rollback, cost, and production ownership are executable.
Six kill criteria for a full implementation award
- No accountable owner can authorize the primary source.
- Usage rights, purpose, or contractual restrictions remain unknown.
- The proposal requires sensitive data in an unapproved environment.
- No sample represents the workflow and its material exceptions.
- Labels, ground truth, or correction criteria cannot be reproduced.
- Production depends on a manual refresh with no owner, frequency, or cost.
A kill criterion does not mean abandoning AI. It changes the purchase. Buy inventory, profiling, access remediation, evaluation design, or a controlled pilot. The commercial mistake is funding an implementation team to discover late that the decisive input does not exist or cannot be used.
Require eight comparable artifacts
- Decision brief connecting data, user, action, risk, and metric.
- Source inventory with owner, purpose, volume, frequency, and constraints.
- Access matrix by person, system, environment, and phase.
- Profiling report with rules, segments, exceptions, and measurement date.
- Lineage card with origin, transformations, versions, and dependencies.
- Development, evaluation, and holdout-set plan with leakage controls.
- Remediation backlog with impact, owner, cost, and critical path.
- Operating plan for refresh, observability, incidents, and retirement.
Run a 90-minute data-defense test
Give every finalist the same masked sample and decision brief. Allow 30 minutes for a quality profile and risk hypotheses; 20 for an evaluation-split proposal; then inject a changed definition, an underrepresented segment, and a late source. In the final 20 minutes, require a revised plan, cost and schedule impact, and an implement, remediate, or stop recommendation. Score causal reasoning, questions, reproducibility, and data protection—not dashboard polish.
Choose implementation, discovery, or stop
- Implement — critical proofs score three or four, gaps have owners, and the production pipeline is feasible.
- Capped discovery — value is plausible but access, profile, coverage, or evaluation still needs proof; buy the eight artifacts with a time and price ceiling.
- Stop or redesign — rights, representativeness, security, or operating capacity has no acceptable path, or the data cannot support the proposed decision.
Make the decision per use case, not per enterprise. The same knowledge corpus may be ready for an internal search assistant with human review and unfit for an automated high-impact decision. Readiness changes when purpose, population, source, model, autonomy, or operating environment changes.
U.S. buying context
Bring privacy, security, legal, records, accessibility, and sector specialists into the inventory according to actual exposure. Map state, federal, contractual, customer, employment, health, financial, export, and confidentiality constraints before bidders receive data. A technical quality score does not establish a right to use data, and a legally permitted source may still be unrepresentative or operationally fragile.
Proposal red flags
- The bidder promises to clean data later without measuring the gap.
- Pricing excludes integration, labeling, evaluation, governance, or operation.
- Quality is one average with no segments or adverse cases.
- Evaluation data can leak into development or prompts.
- Provenance and synthetic content are not identified.
- The solution relies on manual extracts that nobody will own in production.
- The provider requests a broad copy before defining minimization and controls.
Carry readiness into the RFP and contract
Give competitors the same decision brief, minimum data dictionary, and aggregated profile without exposing personal data or trade secrets beyond what is necessary. Require each proposal to separate facts, assumptions, client dependencies, and remediation; price scenarios; name decision owners; and tie the first payment gate to reproducing the baseline and evaluation set. Add change control when a source, purpose, population, or refresh mechanism changes.
Connect data readiness to the buying system
Use https://makinai.co/insights/en/should-you-pay-for-ai-discovery-before-implementation to choose the first engagement, https://makinai.co/insights/en/how-to-evaluate-realistic-ai-project-timeline to expose dependencies, https://makinai.co/insights/en/how-to-choose-ai-evaluation-testing-company to design evaluation, and https://makinai.co/insights/en/security-due-diligence-ai-services-company to deepen access and third-party controls.
When to involve MAKINAI
MAKINAI can run a focused readiness assessment, normalize vendor evidence, and turn data gaps into scope, gates, economics, and a defensible sourcing decision before implementation. Explore https://makinai.co/services/en/ai-strategy-transformation-consulting.