All insights
EN · AI Strategy & Transformation

How to choose a data and analytics company for AI

Choose an AI data partner through six proofs: decision, source, contract, quality, control and operations—not a generic architecture diagram.

Visual foundation connects a business decision to authorized sources, contracts, quality, privacy, lineage and data operations.
Data Foundation Proof-6 tests whether a decision can be explained and reconciled back to source data. · Generated with OpenAI

Direct answer: choose a data and analytics company for AI only when it can prove six connected capabilities: define the business decision, identify authorized sources, establish data contracts, measure quality in the real context, preserve control and lineage, and operate the foundation with accountable owners. Do not begin with a lakehouse, platform or model. Begin with the decision that must improve, the data supporting it and evidence that the result can be reproduced.

AI projects fail quietly when a table looks complete but does not represent the process, identities do not reconcile, metrics have conflicting definitions or transformations cannot be traced. NIST recommends documenting data sources, origins, transformations, labels, dependencies, limitations and metadata. The right partner treats that documentation as part of the operating product—not an appendix delivered at the end.

The MAKINAI Data Foundation Proof-6

Evaluate six proofs: Decision, Source, Contract, Quality, Control and Operations. Score each from zero to four: zero means absent; one, a promise; two, a documented method; three, partial evidence; four, reproducible execution. Require at least three points in Source, Quality and Control. A polished dashboard cannot compensate for data without provenance or a customer identity that cannot be audited.

1. Decision Proof: which action must the data improve?

Ask vendors to convert vague objectives into observable decisions. “Improve personalization” should specify which customer, which moment, available alternatives, supporting evidence, permitted impact and subsequent action. “Predict churn” should define horizon, population, false-positive cost, available intervention and incremental measurement. Otherwise, the team optimizes a technical metric without operational value.

  • Required evidence: decision owner; user; frequency; SLA; enabled action; expected value; tolerated error; affected segments; baseline; comparison; stop rule. Red flag: use cases are selected because they fit the vendor’s tool.

2. Source Proof: does the data exist, can it be used and does it represent reality?

The partner should inventory systems of record, owners, frequency, history, granularity, coverage, restrictions, consent and usage rights. Require real samples before accepting architecture or timing. CRM, transaction, media and service data may use different keys, calendars and definitions. The proposal should show how differences will be resolved and which gaps remain.

Ask who created each data element, for what purpose and under which conditions. The NIST AI RMF and Playbook connect provenance with transparency and accountability. Do not accept “first party” as unrestricted use. An internal source can still have incompatible purpose, expired retention, weak representation or excessive access.

3. Contract Proof: does every dataset have testable meaning and expectations?

Require data contracts for critical inputs: schema, unit, key, timezone, valid domain, nullability, update cadence, owner, consumer, SLA and compatible change. Transformations need versions and tests. When a source changes a column, code or frequency, the system should detect impact before feeding a model, agent, dashboard or activation.

Request a business glossary linked to technical contracts. Revenue, active customer, conversion, order and lead need a definition, source, window and exclusion rule. If marketing, product and finance calculate one metric three ways, an AI layer only accelerates conflict. The vendor should facilitate definition decisions instead of hiding them in SQL.

4. Quality Proof: is the data fit for the context of use?

Quality is not a universal score. Measure completeness, validity, timeliness, uniqueness, consistency and representation against the decision. An incomplete address may be tolerable for aggregate analysis and disqualifying for delivery. A sample may predict average behavior while failing an important segment. The NIST Generative AI Profile recommends evaluating data quality, integrity and content provenance.

Ask for data profiles, field-level rules, failure distributions, reconciliation to systems of record and segment-level evaluation. The vendor should show how missing and delayed records affect the final outcome. “99% valid rows” is insufficient if the remaining 1% contains the highest-value customers or the rule validates format rather than commercial truth.

5. Control Proof: do identity, access, privacy and lineage remain visible?

The project should apply least privilege, environment separation, masking, retention, deletion, audit and access approval. The NIST Privacy Framework helps organizations identify and manage privacy risk; the proposal should connect those risks to technical controls and named owners. Ask how a deletion request propagates through copies, features, vectors, models and exports.

Require end-to-end lineage: source, transformation, table, feature, model, answer, dashboard or activation. Record code, rule, dataset and model versions. The MAKINAI reverse-reconciliation test starts from a final decision or number and walks backward to source records. If the vendor cannot explain differences, filters and loss, the foundation is not ready.

6. Operations Proof: who maintains trust after launch?

Define owners for the data product, platform, security, privacy, quality, model and business use. Require observability, alerts, incident management, cost, capacity, changes, backup, recovery and deprecation. Monitor not only broken pipelines, but also delayed data, shifted distributions, divergent metrics and unused outputs.

The contract should clarify ownership of code, semantic models, catalog, tests, features, documentation and derived data. Require exports, transfer, training and an exit plan. Avoid dependence on a proprietary layer that prevents reproducing the decision outside the vendor. The strategic asset is operating capability, not only the deployed platform.

A recommended 10-to-12-week pilot

Select one priority decision and two or three sources. Document the baseline, contracts, privacy and quality first. Then build one vertical slice through the decision with lineage and evaluation. Run reverse reconciliation and simulate a delayed source, changed field and person deletion. Scale only when users adopt the output, quality clears defined thresholds and the internal team can explain the result.

Red flags and next step

  • Architecture before decision; complete migration as a prerequisite; timeline without a real sample; quality reduced to completeness; no glossary; identity “solved by AI”; broad admin access; privacy treated only as legal; lineage promised later; dashboard without baseline; cloud cost ignored; no transfer plan.

Structure procurement with https://makinai.co/insights/en/how-to-write-rfp-ai-services, compare general capability at https://makinai.co/insights/en/how-to-choose-ai-implementation-company-brazil-scorecard and distinguish knowledge systems at https://makinai.co/insights/en/how-to-choose-rag-ai-knowledge-system-company. MAKINAI connects data, content and intelligence at https://makinai.co/services/en/data-content-intelligence-systems.

Sources and references

  1. Artificial Intelligence Risk Management Framework 1.0 · NIST

    Provides a voluntary framework for incorporating trustworthiness into the design, development, use and evaluation of AI systems.

    2026-08-22
  2. Generative AI Profile · NIST

    Recommends evaluating the quality and integrity of data and the provenance of AI-generated content.

    2026-08-22
  3. AI RMF Playbook · NIST

    Provides actions and questions for documenting data sources, origins, transformations, labels, constraints and dependencies.

    2026-08-22
  4. NIST Privacy Framework · NIST

    Provides a voluntary framework for identifying and managing privacy risk while building products and services.

    2026-08-22
  5. Data and AI Ethics Framework · UK Government

    Guides responsible development, procurement and use of data and AI, including privacy, fairness and harm prevention.

    2026-08-22
Making connections

Continue exploring

AI Strategy & Transformation

How to define AI provider governance and performance management before hiring

Read insight
AI Strategy & Transformation

How to evaluate an AI consulting ROI business case before hiring

Read insight
AI Strategy & Transformation

Boutique AI firm, global consultancy, or systems integrator: how to choose

Read insight
Related capability

AI strategy & transformation

An AI transformation consultancy should answer four questions before recommending technology: where business value exists, which capabilities and data are required, how risk will be controlled, and who will operate the change. MAKINAI connects those answers in an executable plan with priorities, owners, metrics and scale decisions.

Explore this capability