Procuring AI is not like procuring traditional software. A conventional enterprise application behaves deterministically: given the same input, it produces the same output, every time. An AI system is probabilistic, opaque, and prone to performance degradation over time. These fundamental differences demand a procurement process that goes well beyond functional requirements and pricing negotiations.
Why Traditional Procurement Fails for AI
When an organisation purchases a customer relationship management platform or an enterprise resource planning system, the evaluation criteria are relatively straightforward: does it meet our functional requirements, integrate with our existing stack, and fit our budget? The software's behaviour is predictable, and the vendor's track record with similar implementations provides reasonable confidence in the outcome.
AI procurement introduces several dimensions that traditional software evaluation does not address:
Opacity. Many AI systems — particularly those built on large language models or deep neural networks — operate as functional black boxes. The vendor may not be able to explain precisely why the system produced a specific output, making it difficult to assess reliability or diagnose failures.
Performance variability. An AI system's accuracy in the vendor's demonstration environment may not transfer to your specific data distribution, domain vocabulary, or operational context. The model may perform well on the vendor's benchmark datasets while failing on the edge cases that matter most to your business.
Temporal degradation. Unlike traditional software that continues to function as designed until explicitly updated, AI models degrade as the real-world data they encounter diverges from their training data. A model that performs well at deployment may become unreliable within months if the underlying patterns shift.
Data dependencies. AI systems are inseparable from their training data. Understanding how the model was trained — on what data, with what biases, under what assumptions — is essential to assessing its suitability and risk profile.
Regulatory exposure. The EU AI Act, effective August 2025, imposes specific obligations on both providers and deployers of AI systems. Organisations that deploy a high-risk AI system without adequate due diligence may find themselves liable for violations they could have anticipated during procurement.
A Five-Dimension Evaluation Framework
Effective AI vendor due diligence requires assessment across five interconnected risk dimensions:
1. Technical Risk
The technical evaluation must go beyond accuracy metrics to assess the system's reliability, explainability, and fitness for your specific use case.
Model documentation. Request Model Cards and Datasheets for Datasets — standardised documentation formats that describe the model's intended use cases, known limitations, training data characteristics, and performance across different demographic groups or data segments. If the vendor cannot or will not provide this documentation, treat that as a significant risk indicator.
Performance validation. Insist on evaluating the system against your own data — not the vendor's curated demonstration set. Pay particular attention to performance on edge cases, underrepresented data segments, and adversarial inputs. A model that achieves 95 percent accuracy on average but fails systematically on a critical minority of cases may be worse than no model at all.
Explainability. For high-stakes applications, the ability to understand why the system reached a particular decision is not optional. Assess whether the vendor provides meaningful explanations (not just confidence scores) and whether those explanations are actionable for your domain experts.
2. Operational Risk
Operational assessment focuses on the system's reliability and the vendor's ability to support it over time.
Drift management. How does the vendor monitor and address model drift? Do they provide automated drift detection? What is their retraining cadence? Who is responsible for performance monitoring — the vendor or the deployer?
Service level agreements. Standard uptime SLAs are necessary but insufficient. AI-specific SLAs should cover prediction accuracy thresholds, response latency, and incident response procedures for model failures — not just server availability.
Human-in-the-loop capabilities. For high-risk applications, the system must support meaningful human oversight. This means providing confidence indicators, escalation mechanisms for uncertain predictions, and audit trails that allow reviewers to understand and challenge the system's reasoning.
3. Financial Risk
The AI vendor landscape is volatile. Rigorous financial due diligence protects against service disruption.
Vendor viability. Assess the vendor's funding runway, revenue model, and market position. A technically excellent AI system becomes a liability if the vendor cannot sustain operations. Request financial disclosures appropriate to the contract size and dependency level.
Total cost of ownership. AI costs extend beyond licence fees. Factor in integration, customisation, ongoing monitoring, retraining, and the internal expertise required to operate the system effectively. Hidden costs — such as data labelling, compute for fine-tuning, and compliance auditing — often exceed the initial licence cost.
Exit strategy. Before signing, define the off-boarding process. Can you export your data? In what format? What happens to models fine-tuned on your data? Is there a transition period? Vendor lock-in in AI is particularly acute because switching often requires not just technical migration but retraining from scratch.
4. Legal and Regulatory Risk
AI procurement contracts must address risks that do not arise with traditional software.
Regulatory compliance. Under the EU AI Act, deployers of high-risk AI systems bear specific obligations — including conducting fundamental rights impact assessments and ensuring human oversight. Verify that the vendor's system is designed to support these obligations, and that the contract clearly allocates responsibility for compliance.
Data processing. Understand precisely how the vendor processes your data. Is it used to improve the vendor's base model? Is it stored, and where? Who has access? A clear data processing agreement that explicitly prohibits using your proprietary data for model training without consent is essential.
Intellectual property. If the vendor fine-tunes a model on your data, who owns the resulting weights? If the model generates content that infringes third-party IP, who is liable? These questions must be answered in the contract, not discovered during litigation.
Audit rights. The contract should grant you the right to audit the vendor's AI systems, bias documentation, security logs, and compliance processes. Without audit rights, your due diligence is limited to the vendor's self-reported claims.
5. Ethical and Reputational Risk
AI failures are headline news. Ethical due diligence protects the organisation's reputation and social licence to operate.
Bias assessment. Request the vendor's bias testing methodology and results. Has the system been evaluated for disparate impact across protected characteristics? What mitigation measures are in place? If the vendor has not conducted bias testing, the risk falls entirely on the deployer.
Labour practices. AI systems often rely on human data labelling, content moderation, and red-teaming. Understanding the vendor's labour practices for these functions is both an ethical obligation and a reputational risk factor.
Societal impact. For high-stakes applications — hiring, lending, healthcare, law enforcement — the broader societal implications of deploying the system deserve explicit assessment. This is not merely altruistic; regulators and the public increasingly hold organisations accountable for the downstream effects of their AI deployments.
Standardised Assessment Tools
Several industry frameworks can structure the evaluation process:
CSA AI Controls Matrix (AICM) and its companion questionnaire provide a comprehensive, standardised assessment framework specifically designed for AI vendor evaluation. Using a recognised standard reduces the risk of overlooking critical assessment areas.
Evidence-based verification has replaced policy-based trust. Rather than accepting vendor claims at face value, request evidence packs that include third-party audit reports, signed runtime receipts, penetration test summaries targeting AI-specific attack vectors, and documented bias assessment results.
Quantitative risk modelling frameworks help translate technical AI risks into financial loss-exposure metrics that procurement and legal teams can use to make informed decisions about acceptable risk levels.
Making Due Diligence Operational
AI vendor due diligence is not a one-time procurement exercise but an ongoing governance process. After contract signing, organisations should establish a regular cadence for reviewing performance metrics, reassessing bias, verifying security posture, and confirming continued regulatory compliance. The vendor relationship should include defined triggers for escalation — performance thresholds that, if breached, require the vendor to provide a remediation plan within specified timeframes.
The goal is not to create an impossibly high bar that prevents AI adoption, but to ensure that the organisation's enthusiasm for AI capabilities is matched by a clear-eyed understanding of the risks those capabilities bring. Thorough due diligence is not a barrier to innovation — it is the foundation on which sustainable AI adoption is built.
