Run the selection through seven gates: market map, shortlist, comparable brief, written proposal, evidence session, paid proof, and final offer. Start with enough alternatives to expose different delivery models, then narrow the field before expensive stages. For a typical enterprise engagement, a useful planning range is to map eight to twelve firms, qualify five or six, invite three or four proposals, and take one or two teams into a paid proof. These are not universal rules; adjust them to market depth, risk, complexity, and buyer capacity.
Competition improves a decision only when every firm receives the same problem, constraints, permitted data, criteria, timeline, and communication rules. A parade of generic capability decks creates volume, not evidence. Preselecting a favorite and using other bids only as price leverage damages trust. The goal is to identify which partner can reduce uncertainty about value, execution, and risk.
The AI Partner Selection Gate-7: 28 points before signature
Score each gate from zero to four: zero means absent; one is an informal intention; two is a partial artifact; three is an evidence-backed gate; four is a documented, comparable, auditable decision. Require at least 21 of 28, no zero, and a minimum of three for the brief, evidence session, and paid proof before a material implementation award.
- Market map — problem, delivery models, available capability, constraints, and signs of healthy competition.
- Shortlist — objective screens for eligibility, conflicts, critical capability, geography, and team availability.
- Comparable brief — outcomes, users, baseline, data, integrations, risks, deliverables, operations, budget, and process rules.
- Written proposal — solution, assumptions, named team, plan, total economics, dependencies, evidence, and exceptions.
- Evidence session — the actual delivery team works a common scenario under equal questions and controlled pressure.
- Paid proof — a small representative slice with authorized data, injected failure, acceptance criteria, and reusable deliverables.
- Final offer — normalized scope, resolved risks, verified references, essential terms, decision, and rationale.
Gate 1: research the market without designing for one vendor
Hold exploratory conversations before the RFP to test whether the problem is buyable, what capabilities exist, what must remain in-house, and which assumptions drive price or schedule. Ask consistent questions and record learning without promising advantage. The DDaT Playbook connects early engagement to solution feasibility, testing, evaluation, intellectual property, and exit while emphasizing equal treatment. Private buyers can use that integrity principle even when public procurement rules do not apply.
Gate 2: use a small number of true knockout screens
Separate eligibility from scoring. Eliminate only for genuinely mandatory conditions: an unmanageable conflict, missing critical capability, incompatible data or residency restriction, unavailable key staff, refusal to disclose material subcontractors, or inability to accept an essential term. Certifications and company size do not substitute for problem-specific competence. Too many screens can exclude specialists and smaller firms before they demonstrate value.
Gate 3: issue a brief that makes bids comparable
- Business decision, users, current process, and known baseline.
- Included and excluded scope; data, integrations, environments, and volumes.
- Unacceptable risks, controls, human review, and prohibited decisions.
- Phase deliverables, acceptance criteria, operations, transfer, and exit.
- Pricing format, assumptions, third-party costs, and consumption scenarios.
- Criteria, weights, required evidence, calendar, contact, and question rules.
Publish criteria before offers arrive and score only what was requested. FAR Part 15 is a useful discipline reference, not a rule for every private company: factors should enable meaningful discrimination, and the relative importance of price and non-price factors should be clear. World Bank guidance adds that a limited set of tailored criteria keeps the evaluation focused on genuine differentiators.
Gate 4: constrain narrative and expose assumptions
Require the same structure and page limits. Ask each bidder for its architecture, value path, named team, first-weeks plan, client dependencies, initial risk register, evaluation approach, scenario costs, and contract exceptions. Separating fact, assumption, and pending decision is more informative than confident prose. Normalize optional items and model, cloud, license, and data costs before comparing totals.
Gate 5: replace the pitch with an evidence session
Give finalists the same case and run a 90- to 120-minute working session with the people who would deliver. Introduce a data change, integration failure, risk constraint, and budget reduction. Observe who asks, who decides, how the team records uncertainty, which trade-offs it protects, and when it recommends stopping. Score behaviors against a rubric written before the meeting; do not reward charisma or slide polish.
- The presenters match the people named for delivery.
- Questions uncover value, workflow, data, operations, and risk.
- Assumptions become explicit and receive a validation plan.
- Controls survive schedule or budget pressure.
- Decisions, owners, and next tests are recorded.
- Solution limits are explained without evasive language.
Gate 6: pay for the proof; do not demand free implementation
Use a paid proof only when material uncertainty cannot be resolved through proposals, references, or the workshop. Give finalists comparable conditions, compensate the work, restrict data use, and predefine ownership, confidentiality, security, deliverables, and disposal. A good proof runs one end-to-end slice, measures quality and cost, injects at least one failure, and returns reusable assets. Do not turn an ideas contest into unpaid speculative delivery.
Gate 7: request final offers after uncertainties are reduced
Before the final round, send each firm its observed gaps, preserve consistent treatment, and allow a controlled proposal revision. Compare the same scope, volume, staffing, risk, and commercial baseline. Document why technical or operating benefits justify any price premium, which risks remain, who accepted them, and what evidence supports the trade-off. Do not quietly reopen criteria already used to remove competitors.
A decision matrix that keeps unlike factors separate
- Outcome and method, 20% — problem understanding, value, discovery, and learning plan.
- Team and collaboration, 20% — named people, relevant seniority, availability, and decision behavior.
- Engineering and integration, 15% — architecture, data, security, evaluation, observability, and quality.
- Delivery and operations, 15% — plan, dependencies, change, support, transfer, and continuity.
- Risk and governance, 15% — controls, supply chain, rights, auditability, and exit.
- Total economics, 15% — phase price, consumption, third parties, operating cost, and scenarios.
Those weights are an illustrative starting point. Tailor them before release and never after seeing names or prices. Use behavioral anchors: zero absent; one assertion; two partial approach; three sufficient evidence; four superior, tested evidence. Require a short rationale for every score and have evaluators score independently before calibration. Disagreement is a prompt to investigate, not a reason to average quickly.
Avoid four errors that contaminate competition
- Inviting too many firms into a high-effort stage.
- Giving bidders different material information or privileged access.
- Changing weights after recognizing a preferred brand or solution.
- Treating the lowest rate as a proxy for lowest total cost and risk.
Name a process lead; business, technology, data, risk, and operations evaluators; and a final decision owner who cannot rewrite the evaluation alone. Declare conflicts, maintain one question channel, share material answers with all bidders, and preserve versions of briefs, proposals, scores, and decisions. Documentation should be proportionate but sufficient to explain the award months later.
U.S. operating context
Private companies should align the process with internal procurement policy, applicable privacy and sector rules, and delegated authority. In regulated or sensitive-data work, involve security, privacy, legal, compliance, and records stakeholders before the final shortlist. Federal, state, and public-sector acquisitions follow their own rules; this framework is a management tool for enterprise sourcing, not a substitute for applicable procurement law.
Connect each gate to a ready decision asset
Use https://makinai.co/insights/en/how-to-write-rfp-ai-services for the brief, https://makinai.co/insights/en/how-to-evaluate-ai-consulting-proposals-scorecard for the matrix, https://makinai.co/insights/en/verify-ai-consulting-case-studies-client-references for due diligence, and https://makinai.co/insights/en/how-to-run-paid-ai-pilot-before-hiring-partner for the paid proof.
When to involve MAKINAI
MAKINAI can help map the market, turn the decision into a comparable brief, prepare the evidence session, and structure a paid proof before final selection. Explore https://makinai.co/services/en/ai-strategy-transformation-consulting. The process should end with a defensible decision, an executable team, and a first delivery gate—not merely a vendor name.