A proposal needs AI when variability, volume, or ambiguity makes explicit rules and conventional software insufficient—and representative data shows material incremental value after error, oversight, security, and lifecycle cost are included. The supplier should prove that value against a simpler alternative. Without a baseline, comparison, and fallback, AI describes the sale rather than the need.
The right question is not ‘where can we add AI?’ It is ‘what is the least complex solution that meets the outcome within our risk tolerance?’ Require each finalist to compare conventional software, deterministic automation, machine learning, and generative AI or agents against the same case, data, and metrics.
Four paths every proposal should compare
- Conventional software — stable workflows, known logic, high predictability, and little benefit from inference.
- Deterministic automation — repeatable tasks, auditable rules, predictable integrations, and bounded exceptions.
- Machine learning — classification, prediction, or recommendation where data patterns outperform handcrafted rules.
- Generative AI or agents — language, content, or assisted decisions in variable environments, with evaluations, action limits, and proportionate oversight.
Hybrid designs often win: rules control permissions and transactions; ML prioritizes; generative AI interprets content; people approve higher-impact decisions. The goal is not category purity. It is preventing a probabilistic component from owning work that requires deterministic behavior.
AI Necessity & Fit Proof-8: a 32-point scorecard
Score each dimension from zero to four: zero is absent; one is an assertion; two is partial evidence; three is a current test with representative data; four adds independent comparison, limits, and a reassessment trigger. As an illustrative gate, proceed at 24 of 32, with no zero and at least three for baseline, incremental lift, risk, and fallback.
- Problem shape — variability, volume, and response time justify inference.
- Output fit — uncertainty tolerance, explainability, and review match the technique.
- Data signal — lawful, representative, accessible data contains information useful to the outcome.
- Simple baseline — rules, manual workflow, or existing software is measured with the same metrics.
- Incremental lift — added quality, speed, or scale is material and attributable to AI.
- Operating controls — evaluations, monitoring, oversight, action limits, and incident handling are defined.
- Lifecycle economics — integration, usage, human review, operations, error, and change enter cost.
- Reversibility — fallback, shutdown, portability, and a non-AI path work when needed.
Apply six kill criteria
- The supplier does not measure the current process or define a baseline.
- No deterministic alternative is designed and tested.
- The demo uses supplier-selected data rather than a representative buyer sample.
- The business case assigns all improvement to AI without separating process redesign, integration, and automation.
- There is no failure threshold, human review, fallback, or degraded mode.
- Value disappears when consumption, operations, evaluation, error, and provider change are included.
Run a 90-minute counterfactual test
Give finalists one real workflow, twenty typical examples, and five adversarial cases. Request four sketches: no change, software or rules, ML, and generative AI or an agent. Each must show metrics, cost, time, risk, dependencies, and abandonment conditions. Then remove half the data or require human review; see whether the recommendation changes coherently.
Require five artifacts before award
- Problem and user map, including exceptions and impact of error.
- Reproducible baseline for the current process and simplest alternative.
- Four-option comparison with assumptions and evidence.
- Pilot plan with sample, metrics, thresholds, and stop criteria.
- Fallback architecture with ownership, cost, and reassessment triggers.
Put the proof into the RFP, pilot, and payments
Do not prescribe a model or platform before the problem is clear. Ask for the recommended and minimum-sufficient options, unbundled costs, and evidence that would make the team choose less AI. Pay the pilot for comparable evidence. Gate implementation on incremental lift and risk approval—not a standalone demo that merely works.
United States context
Map applicable federal, state, sector, privacy, accessibility, employment, and consumer obligations to the actual use. For consequential decisions, legal, security, privacy, and process owners should validate purpose, data, human authority, notice, contestability, and monitoring. Counsel should tailor requirements; official procurement frameworks are rigorous references, not universal private-sector law.
Connect this decision to adjacent diligence
Use https://makinai.co/insights/en/buy-configure-or-build-ai-solution to define paths, https://makinai.co/insights/en/practical-framework-prioritize-ai-use-cases to prioritize the case, https://makinai.co/insights/en/evaluate-ai-consulting-roi-business-case-before-hiring to test value, https://makinai.co/insights/en/how-to-run-paid-ai-pilot-before-hiring-partner to produce evidence, and https://makinai.co/insights/en/evaluate-ai-solution-architecture-proposal-before-hiring to validate the design.
When to involve MAKINAI
MAKINAI can structure the baseline, compare AI and non-AI solutions, run the counterfactual test, and translate the recommendation into verifiable investment and delivery gates. Explore https://makinai.co/services/en/ai-strategy-transformation-consulting.