ctaio.dev Ask AI Subscribe free

Guides / Buying AI / AI Due Diligence

Guides · Buying AI · Updated 2026-09-14

AI Due Diligence: What to Ask Before You Sign.

An RFP collects answers. Diligence establishes which ones survive evidence.

AI vendor due diligence is the verification step between a vendor's claims and a signature: reproduce the performance claim on data the vendor has never seen, trace the model supply chain to whoever owns the weights, establish who can read a prompt, read the SOC 2 report for scope rather than for the logo, and confirm the vendor will still be operating at renewal. This page is the procedure, starting with a bench test you actually run and ending with the one clause diligence adds to the contract.

AI due diligence: what to ask before you sign

30-SECOND POV

  • A benchmark number is a claim about a test set, not about your work. Documented contamination in public coding benchmarks has moved a headline resolution rate by a factor of three. Reproduce the number on data the vendor has never seen, or discount it.
  • Check the vendor's model-change notice against their supplier's published retirement date. Recent Anthropic deprecations ran 61 and 62 days from announcement to retirement. Whatever your vendor promises you above their supplier's floor, they are paying for out of their own margin, and you should find out whether they know that.
  • SOC 2 says nothing about the model. It is an examination of controls at a service organization against criteria the service organization helped scope. There is no criterion for hallucination rate, prompt-injection resistance, or training-data provenance.
  • The thin-wrapper question has a numeric answer. Run their product and the base API side by side on the same held-out set, then compare the delta against the price difference. Most of the time the arithmetic decides the deal.

What is this page for?

An RFP collects answers. Due diligence establishes which of those answers survive contact with evidence. The two are usually run by the same people in the same three weeks, which is how a benchmark screenshot ends up in a board pack as a performance commitment.

This page is the verification half. It assumes the vendor has already told you what their system does, through something like the RFP, and asks a narrower question: what can you check yourself, before signature, with the access a serious vendor will grant a serious buyer during evaluation.

THE BENCH TEST

What do you actually run?

Everything else on this page is a question. This is the part you run. Budget one engineer for a week and a second for two days of labelling. If a vendor will not support this during evaluation, the evaluation is over. The numbers below are reasoned defaults, not findings: nobody has measured that 40 items beats 25. Adjust them to your own risk. The order is not adjustable, because steps 2 and 4 stop working out of sequence.

01

Build a holdout set the vendor has never seen

Between 30 and 60 items drawn from your own production data, chosen to match the distribution you actually face rather than the interesting cases. Two of your people label the expected output independently, a third resolves disagreements. Do not send the set to the vendor. Do not describe it to the vendor.

02

Freeze the grader before you see any output

Write the scoring rubric, commit it, and have someone who is not running the test hold the reference file. This is the step teams skip, and skipping it is how a 61% becomes "basically what they claimed" in a summary deck.

03

Run the vendor's product on the holdout, with your people driving

Vendor engineers observe. They do not touch the keyboard and they do not tune prompts mid-run. Record every input and output, wall-clock latency per item, and per-item cost at your projected volume rather than the trial tier.

04

Run the same holdout against the base model you could call yourself

Same items, same rubric, a competent but unoptimised prompt your team writes in an afternoon. This is your floor. Everything you are paying the vendor for lives in the gap between this number and theirs.

05

Re-run the vendor's product seven days later, unannounced

A silent model swap, a prompt change, or a capacity-driven quality drop shows up here and nowhere else in a procurement cycle. Ask afterwards whether anything changed, and compare their answer to your numbers.

06

Write down the two deltas

Vendor minus base model, week one minus week two. Those two numbers are the technical case for the contract. If nobody on the deal team can state them from memory, the diligence did not happen.

How do you test a claimed eval result?

Asking for the eval suite is the RFP's job. This starts one step later, when the artifacts arrive and somebody has to decide what they are worth.

Recompute the headline from the item-level file, not the aggregate in the deck. That catches arithmetic, and more usefully it catches selective exclusion: a dropped category is visible per item and invisible in a mean. If they will not produce one, stop treating the number as evidence.

Then check whether the test set is public, and discount accordingly. An analysis of SWE-bench found that 32.67% of successful patches involved solutions provided directly in the issue report or comments, and a further 31.08% of passing patches were suspicious because the tests were too weak to verify correctness. After filtering those issues out, the resolution rate for SWE-Agent with GPT-4 fell from 12.47% to 3.97% (arXiv:2410.06992). A separate study reported that state-of-the-art models identified buggy file paths from issue descriptions alone with up to 76% accuracy on SWE-Bench-Verified, dropping to at most 53% on repositories not included in the benchmark (arXiv:2506.12286).

The mechanism matters more than the two papers. A benchmark stops measuring capability the moment it becomes a target, and public benchmarks became targets some time ago. When a vendor's differentiation rests on a leaderboard position, you are being sold a number whose relationship to your workload is unestablished. Not dishonest. Just unable to support the weight being put on it.

Then have them run their harness, unmodified, against the holdout from the bench test. A vendor whose harness holds up on data it has never seen has earned the number. One who declines, or who needs a week of prompt work first, has told you what the published figure was measuring.

The version pin is the other half. A result produced on a model identifier since retired is a historical artifact, not a product claim. Ask them to re-run it on what they will actually serve you.

Who actually owns the model you depend on?

Four questions, in order of how much trouble they save.

Who owns the weights?

Three answers, three risks. They serve a third-party API, so their reliability is a subset of the upstream provider's. They host open weights, so ask where the weights came from, what licence governs them, and whether they can serve you if that source disappears. Or they trained the model, so ask for the compute provider, the data provenance, and what happens to the weights in an insolvency.

Does their notice survive a check against their supplier's?

The published notice tiers are the RFP's territory. Diligence adds the check, and it takes ten minutes. Take the model identifier the vendor named, find its retirement date on the provider's own deprecation page, and hold that next to the term length in the draft contract.

Do the arithmetic on intervals, not on promises. Anthropic deprecated claude-opus-4-1-20250805 on 5 June 2026 and retired it 5 August 2026: 61 days. Sonnet 4 and Opus 4 were deprecated 14 April 2026 and retired 15 June 2026: 62 days. OpenAI's Assistants API shut down 26 August 2026, and gpt-3.5-turbo and gpt-4 are scheduled for 23 October 2026. Two months is the window a real migration has actually had. Plan against that, not against the reassurance in the meeting.

A vendor offering twelve months on top of it is either absorbing a migration you cannot see or has not read their supplier's policy. Both are answers. Ask which, and ask who pays for revalidation.

Which platform serves you?

Anthropic's dates apply to Anthropic-operated platforms; Amazon Bedrock and Google Cloud set their own schedules, so status and dates can differ. If your vendor resells through a hyperscaler, the date in their contract and the date in the provider's documentation are two different facts. Get both in writing.

What breaks that is not a retirement?

Lifecycle is the visible failure mode, API contract changes the quiet one. Anthropic now returns a 400 when temperature, top_p, or top_k are set to a non-default value on Claude Opus 4.7 and later, and the Python SDK from v1.0 removes them outright. No model retired. Code that worked stopped working. Ask how they learned about the last upstream breaking change, and how long their customers were affected.

Who can read a prompt, and in which jurisdiction?

Residency is the question everyone asks and the weaker of the two. "Data stays in the EU" constrains geography without constraining who can read a prompt.

The terms that govern the chain (retention in days, training use, subprocessor notice, deletion on termination) belong in the RFP, which asks for all four as numbers rather than assurances.

Diligence adds one move, and it is the one that gets skipped. There are two subprocessor lists, not one. Your vendor has theirs. Every model provider they route to has another, and the second is the one nobody pulls. Get both, read the current published version rather than the screenshot in the deck (trust.anthropic.com/subprocessors, openai.com/policies/sub-processor-list), and diff them against what the vendor told you in writing. A discrepancy between a vendor's claim and its supplier's public documentation is the highest-signal finding this process produces, and it costs nothing.

Two follow-ups. Which legal entity holds the contract and where is it incorporated, because that decides which regulator can reach them. And what happens on support escalation, the most common route by which a prompt you believed was excluded from training reaches a human being.

EU AI Act dates and your provider-versus-deployer role sit on the AI compliance page. The diligence question is narrower: a vendor whose compliance roadmap contains no dates has not started, and one email establishes it.

What does SOC 2 tell you, and what does it not?

SOC 2 is an attestation examination performed against the Trust Services Criteria, established by the AICPA's Assurance Services Executive Committee and published as TSP Section 100, currently the 2017 criteria with revised points of focus issued in 2022. The five categories are security, availability, processing integrity, confidentiality, and privacy (AICPA).

Read that sentence for what it does not contain. No criterion for model output quality, hallucination rate, prompt-injection resistance, eval methodology, training-data provenance, or agentic action safety. A vendor can hold a clean report and have no evaluation practice at all. The report is evidence that a control environment was examined, not evidence that the AI system works.

Three things to do with the report you are handed.

  • Ask for the whole report, not the certificate graphic or the cover letter. The value is in Section 4, where the auditor describes the tests performed and the exceptions found. Zero exceptions across a broad scope is itself a question about scope.
  • Read the scope section first. The AICPA's own framing is controls relevant to security, availability, processing integrity, confidentiality, or privacy. That "or" is elective. Find out which categories and which systems were in scope, and specifically whether the AI inference path was examined or only the corporate SaaS environment around it.
  • Check the period and the subservice organizations. A Type 2 report covers a historical window, so ask when it closed and what has changed since, including any model provider swap. Then ask which subservice organizations were excluded, because controls at an excluded provider are controls nobody tested.

Where do you find the incident history?

Vendors disclose incidents accurately in writing and vaguely on a call. Ask in writing, for something specific: every incident in the trailing 24 months that triggered a customer notification, with dates, duration, affected customers, and remediation. A false statement in a signed representation is a different legal category from one in a sales meeting, which is the whole reason for putting it in the contract.

Independent of what they tell you:

  • The vendor's status page history. Read the postmortems, not the uptime percentage. You are assessing whether they can describe a failure honestly, and whether the remediation named a mechanism or a promise to be more careful.
  • The AI Incident Database at incidentdatabase.ai, run by the Responsible AI Collaborative, which indexes real-world harms from deployed AI systems and publishes a complete database download.
  • CVE and NVD searches on the vendor and any open-source component they name, plus the CISA Known Exploited Vulnerabilities catalog if their stack touches anything internet-facing.
  • Court and regulator records where they are incorporated, and their engineering blog read backwards. Sudden silence for a quarter is information.

Then the question that separates a mature vendor from a young one: what is your customer notification threshold and window, in hours, and is it in the contract? Vendors who have run a real incident answer immediately.

Will the vendor still exist at renewal?

For an early-stage vendor this is a technical risk, not a finance-team formality, because insolvency means your production dependency stops taking traffic.

Ask for audited financials or the auditor's name, runway in months as of a stated date, the date and lead of the last round, gross margin on a workload shaped like yours, and the revenue share held by their largest customer. The last one is the most informative and the least often asked. A vendor with one customer at 40% of revenue has a roadmap that belongs to someone else.

Verify independently where you can. SEC EDGAR full-text search surfaces a Form D for a US private placement, Companies House carries filed accounts for UK entities, and the lead investor's portfolio page confirms the round happened. That last check is a low bar, and a surprising number of claimed raises fail it.

Then convert the risk into terms rather than a spreadsheet: escrow covering weights or model access, prompts, configuration and eval sets, with release triggers including insolvency; a perpetual licence to fine-tuned artifacts derived from your data; and payment terms that do not prepay twelve months to a vendor whose runway you were just told is nine.

Escrow that has never been tested is a document. Ask for a release rehearsal.

Is the vendor a thin wrapper over an API you could call yourself?

Some wrappers are worth paying for. A wrapper has value when the vendor holds something you do not: a labelled domain corpus, an eval suite built over years of production failures, an integration you would spend six months building. The failure mode is the vendor whose product is a system prompt, a UI, and a markup on tokens you could buy directly.

The test is arithmetic, and the bench test already gave you most of the inputs.

  • Compare the deltas. Vendor performance minus base-model performance on your holdout, against vendor annual cost minus your own token cost at the provider's published rate for your volume. Run the number. It is frequently decisive and frequently uncomfortable.
  • Ask to see the system prompt under NDA. A vendor with real engineering behind the product will usually show it, because the prompt is not where the value is. A refusal on a product that is mostly a prompt is the answer.
  • Ask what fraction of engineering headcount works on retrieval, evaluation and data versus UI and integrations. Then ask to speak to one of the evaluation engineers. Whether that person exists tells you more than the org chart.
  • Ask which upstream model, at which version pin. A vendor who cannot name it, or who says "we route across providers" without being able to say which one served your last request, is passing traffic through.
  • Look at latency. A wrapper adds a network hop and some orchestration. If their p50 on your holdout is indistinguishable from the raw API on the same items, little is happening between the two. Streaming and caching confound this one, so treat it as a signal rather than proof.
  • Ask what they do that you could not do in a quarter. Then cost that quarter honestly, including the maintenance you would own forever. Sometimes the answer favours buying. The point of the test is knowing which.

Which clause does diligence add to the contract?

The clauses that turn any of this into an obligation (model change notice, export on demand, eval access, no training) belong with the AI vendor RFP, which sets out all four and how to weight them.

Diligence adds exactly one, and it exists only because you ran the tests.

Benchmark and performance claims restated as warranties, measured against the test procedure you have now actually executed, with a remedy when the number is missed in production. Every other clause protects you from a change. This one protects you from the original claim. A vendor unwilling to warrant a figure they put in a deck has told you what that figure was worth, and finding out costs one sentence in a redline instead of two quarters of production disappointment.

Three neighbouring pages carry the parts this one deliberately leaves out. AI vendor capture risk covers what happens after you have signed and the switching cost has compounded. The AI audit checklist covers auditing your own systems rather than a supplier's. AI compliance covers the EU AI Act, ISO 42001, and NIST AI RMF obligations in full, which this page touches only where they set a date you have to plan against.

Due diligence that ends in a summary memo has produced nothing. Due diligence that ends in one warranted number and two measured deltas has produced the only things here that survive the vendor's next reorganisation.

AI Due Diligence: FAQ

What is AI vendor due diligence?
AI vendor due diligence is the verification step between a vendor's claims and a signature: reproducing performance claims on data the vendor has not seen, tracing the model supply chain to whoever owns the weights, establishing who can read a prompt, reading the security attestation for scope rather than for the logo, and confirming the vendor will still be operating at renewal. An RFP collects claims. Due diligence tests them.
How do I verify an AI vendor's benchmark results?
Recompute the headline from the item-level results file rather than accepting the aggregate, which catches selective exclusion as well as arithmetic. Then have the vendor run their own unmodified harness against a held-out set built from your production data that they have never seen. Discount any figure from a public benchmark, because contamination is documented: filtering flawed issues out of SWE-bench dropped one agent's resolution rate from 12.47% to 3.97% (arXiv:2410.06992), and models identified buggy file paths with up to 76% accuracy on SWE-Bench-Verified against at most 53% on repositories outside it (arXiv:2506.12286).
Does SOC 2 cover AI risk?
No. SOC 2 is an examination against the AICPA Trust Services Criteria: security, availability, processing integrity, confidentiality, privacy. None of those addresses model output quality, hallucination rate, prompt-injection resistance, evaluation methodology, or training-data provenance. A vendor can hold a clean Type 2 report and have no AI evaluation practice at all. Read the scope section to find out which systems and categories were actually examined.
What happens if my AI vendor's model provider deprecates the model?
You migrate, and the question is who pays and how much notice you actually get. Take the model identifier your vendor named, find its retirement date on the provider's deprecation page, and compare that to the notice in your contract. Recent intervals have been short: Anthropic deprecated claude-opus-4-1-20250805 on 5 June 2026 and retired it 5 August 2026, 61 days. A vendor cannot offer more notice than their supplier gives them without absorbing the migration, so ask who revalidates and who pays. Models served through Amazon Bedrock or Google Cloud follow those platforms' own schedules.
How do I know if an AI vendor is just a wrapper around an API?
Run their product and the base model side by side on the same holdout with the same rubric, then compare the performance gap to the price gap at your real volume. Ask to see the system prompt under NDA, ask which upstream model and version pin serves your requests, and compare their p50 latency to the raw API on the same items. A wrapper worth paying for earns its margin through evaluation infrastructure, domain data, or workflow. Ask them to show you that artifact.
Where can I find an AI vendor's incident history?
Ask in writing for every incident in the trailing 24 months that triggered customer notification, with dates and remediation, and get it into the contract as a representation. Independently: read the vendor's status-page postmortems rather than its uptime figure, search the AI Incident Database at incidentdatabase.ai (run by the Responsible AI Collaborative, with a full database download available), run CVE and NVD searches on the vendor and its named components, and check court and regulator records where it is incorporated.

Sources

All figures above were retrieved from these primary pages on 27 August 2026 and re-verified on 14 September 2026.

  1. arXiv:2410.06992: SWE-bench contamination figures (32.67% solution leakage, 31.08% weak tests, 12.47% to 3.97%)
  2. arXiv:2506.12286: 76% vs 53% file-path identification on and off SWE-Bench-Verified
  3. OpenAI deprecations: Assistants API shutdown 26 August 2026, gpt-3.5-turbo and gpt-4 shutdown 23 October 2026
  4. Anthropic model deprecations: deprecation and retirement date pairs; Bedrock and Vertex set their own schedules; temperature/top_p/top_k 400 on Opus 4.7 and later and the Python SDK v1.0 removal
  5. AICPA Trust Services Criteria: five categories, ASEC, TSP Section 100, 2017 criteria with 2022 revised points of focus
  6. AI Incident Database: operated by the Responsible AI Collaborative, complete database download

Note on what is not here: the contents of the OpenAI and Anthropic subprocessor lists are not asserted, only pointed at. No vendor markup multiple appears above because I could not source one.

One warranted number, two measured deltas

That is what survives the vendor's next reorganisation.

·
Thomas Prommer
Thomas Prommer Technology Executive — CTO/CIO/CTAIO

These salary reports are built on firsthand hiring experience across 20+ years of engineering leadership (adidas, $9B platform, 500+ engineers) and a proprietary network of 200+ executive recruiters and headhunters who share placement data with us directly. As a top-1% expert on institutional investor networks, I've conducted 200+ technical due diligence consultations for PE/VC firms including Blackstone, Bain Capital, and Berenberg — work that requires current, accurate compensation benchmarks across every seniority level. Our team cross-references recruiter data with BLS statistics, job board salary disclosures, and executive compensation surveys to produce ranges you can actually negotiate with.