Guides · Buying AI · Updated 2026-09-14
The AI Vendor RFP That Isn't Theatre.
Exit cost is the deciding variable. Weight it that way.
An AI vendor RFP that works asks what happens when you want out. Feature coverage converges within eighteen months. Switching cost compounds from day one, through the model changing underneath you without consent and the artifacts you built on it being unreadable anywhere else. This page is a reusable question set across seven sections, each question with a note on what a non-answer looks like, a scoring split that weights portability at 30, and the four clauses that turn diligence into an obligation.
30-SECOND POV
- Grade on the artifact, not the prose. If a vendor answers a question in a paragraph where a document would do, that is the finding.
- Check every notice promise against the supplier's floor. A vendor promising 90 days on a preview model is promising more than its own supplier owes it. Ask which of them is paying for the gap.
- Zero retention and the best model can be mutually exclusive. Anthropic's Covered Models require 30-day retention. Find that out in the RFP, not in the migration.
- Question 5.8 is the highest-yield question in the document. "Name a customer who has migrated off you." The tell is not the answer. It is the pause.
Why do most AI RFPs select the wrong vendor?
Most AI RFPs are a feature bake-off with a compliance appendix. You send forty questions, three vendors answer yes to thirty-eight of them, and procurement scores the difference. Every question that gets asked is a question the vendor's sales engineer has answered two hundred times, and the answers converge because the questions were written to be answerable.
The questions that separate vendors are the ones about what happens when you want out. Not "do you support SSO" but "what is the file format of my fine-tune artifacts, and can anything except you read it." The second question has a bad answer, which is why it rarely gets asked.
Exit cost is the deciding variable in AI procurement. Feature coverage converges within eighteen months. Switching cost compounds from day one, and the two components that compound fastest are the ones nobody writes into the contract: the model underneath you changing without your consent, and the artifacts you built on top of it being unreadable anywhere else.
That is my position as a practitioner, and it rests on less than I would like. I went looking for a documented migration off an AI application vendor with a real duration or cost attached, and I did not find one I would put my name to. What circulates is a vendor-sponsored survey with no published methodology and a set of dollar figures that appear only on content farms with no primary source behind them. Every number below this paragraph traces to a provider's own documentation. The claim in this paragraph does not, and you should read it as judgement rather than evidence.
The three failures of a standard AI RFP
It treats the model as a fixed thing. It is not. It has a retirement date, and in some cases that date was set programmatically the day it launched. Microsoft Foundry sets a general-availability model's retirement 18 months out at launch (12 months for partner models such as Anthropic, DeepSeek and Mistral) and exposes it through the Models API, and Microsoft's answer to whether that date can be extended is a flat no: "Retirement dates aren't extendable." Your RFP is scoring a system whose core component has a documented expiry you did not choose.
It asks about capability, not about change. A vendor demoing on GPT-5 today may serve you a different model in nine months. On Foundry, Standard-tier deployments are auto-upgraded when a version retires and Provisioned deployments are not, which means the same vendor running the same product can hand two customers materially different behaviour depending on a deployment property neither customer saw.
It confuses a security questionnaire with a security posture. SOC 2 tells you the vendor has controls. It tells you nothing about whether anyone has tried to break the model, who tried, what they found, or whether you get to see the report.
What follows is the question set. Lift it. Each section has questions, and each question has a note on what a non-answer looks like. Grade on the artifact, not the prose.
Section 1: Which model serves you, and what is its change policy?
- 1.1Name every model that serves production traffic in this product today: exact model identifier including version or snapshot suffix, the provider, and the hosting platform (first-party API, Azure/Foundry, Bedrock, Vertex, self-hosted).
- 1.2For each model in 1.1, state its published retirement or deprecation date and the source URL for that date.
- 1.3What notice do you give us before the model serving our traffic changes? State it in days, as a contractual minimum, not as an intention.
- 1.4When your upstream provider retires a model, what is your default behaviour: auto-upgrade, pinned version until forced migration, or customer election?
- 1.5Do we have the right to remain on a pinned model version, and for how long past your migration date?
- 1.6Describe your last three model migrations: the dates, the eval deltas you measured before and after, and what regressed.
- 1.7Do you route requests across multiple models or providers (cost routing, fallback, cascading)? If yes, can we see the routing policy, and can we disable it?
- 1.8Under what circumstances can you change the model with no notice at all?
The upstream numbers are public, so 1.3 is checkable against them. Anthropic commits to "at least 60 days' notice before model retirement for publicly released models," and its own history sits right on that floor: Claude Opus 4.1 was deprecated on 5 June 2026 and retired on 5 August 2026, a 61-day window. OpenAI's stated minimums are longer and tiered: "at least 6 months" for generally available models, "at least 3 months" for specialized variants, and preview models that "may be retired with much shorter notice, such as 2 weeks." Microsoft commits to at least 60 days for GA models and at least 30 for preview, and reserves an emergency retirement with shortened notice if a model has compliance or security issues.
Read what that means for question 1.3. If a vendor promises you 90 days of notice while sitting on a preview model, the promise exceeds what its own supplier owes it. That vendor is either self-insuring the gap or has not read the deprecation page. Ask which.
Question 1.4 has a specific technical answer on Azure, and asking for it by name flushes out whether the vendor knows its own stack. The versionUpgradeOption deployment property takes three values: OnceNewDefaultVersionAvailable, OnceCurrentVersionExpired, and NoAutoUpgrade. A vendor that cannot tell you which one their deployments carry does not control the model you are buying.
Section 2: What happens to your data, and who can train on it?
- 2.1Is our input or output data used to train, fine-tune, evaluate, or otherwise improve any model, yours or a third party's? Answer separately for each.
- 2.2Where is that commitment written? Give the clause reference in the contract we will sign, not the marketing page.
- 2.3State the retention period in days for: prompts, completions, embeddings, uploaded files, tool-call arguments and results, and abuse or trust-and-safety logs. Six numbers.
- 2.4Which of your product features are incompatible with zero-retention operation, and what breaks if we turn them off?
- 2.5Which models are unavailable to us if we require zero retention?
- 2.6Under what conditions is data retained beyond the stated period, and for how long?
- 2.7Which subprocessors see our data, in which jurisdictions, and what is your notice period for adding one?
- 2.8On termination, what is deleted, on what timeline, and what evidence of deletion do we receive?
Question 2.3 asks for six numbers because a single number is always the flattering one. Anthropic's own documentation is the model of how granular the real answer is: code execution container data is retained up to 30 days and is not zero-retention eligible, the Activity Feed retains for 6 years, local session transcripts from Claude Code default to 6 years, and content flagged by automated trust-and-safety systems may be retained up to 2 years even under a zero-retention arrangement. One vendor, one product family, four different clocks.
Question 2.5 is the one that surprises people, and it is not hypothetical. Anthropic designates Claude Fable 5, Fable 5.1, Mythos 5 and Mythos 5.1 as Covered Models that require 30-day data retention, so zero retention is not available for any of them unless Anthropic expressly authorises it. A request from a non-conforming organisation returns a 400 invalid_request_error reading "In order to access this model, your organization or workspace must have data retention enabled." Your strictest data posture and your best model can be mutually exclusive, and you want to find that out during the RFP rather than during the migration.
The same shape appears on OpenAI. Zero Data Retention excludes customer content from abuse monitoring logs and forces the store parameter to false on /v1/responses and /v1/chat/completions. It does not cover a long list of stateful endpoints including /v1/assistants, /v1/threads, /v1/vector_stores, /v1/files, /v1/fine_tuning/jobs, /v1/evals, and /v1/batches. A vendor whose product is built on vector stores and file uploads cannot give you zero retention on the parts that hold your documents, whatever the front page says.
On 2.1, the good answers exist and are quotable. Anthropic's commercial terms say "Anthropic may not train models on Customer Content from Services" and assign output rights to the customer. OpenAI states that "As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)." If your vendor sits on top of those APIs, its own commitment should be at least as strong. If it is weaker, the gap is the vendor's own doing and should be priced.
Section 3: Who runs the evals, and on whose data?
- 3.1Provide the eval suite you used to produce every accuracy or quality number in your proposal: the dataset, its size, how it was constructed, and who labelled it.
- 3.2Can we run your eval suite ourselves, against your production endpoint, on our data? If not, why not?
- 3.3What is your regression policy? Give the metric, the threshold, and the action taken when a model change crosses it.
- 3.4Do you re-run evals after every model change, including upstream provider changes you did not initiate?
- 3.5Give us the eval results for the last model change you made, including anything that got worse.
- 3.6Who inside your company can veto a model rollout on eval grounds, and has anyone ever done it?
- 3.7How do you measure quality on inputs that look nothing like your eval set?
A vendor benchmark is a vendor artifact. The point of 3.2 is to move the eval to your side of the boundary, because the only eval that predicts your outcome runs on your distribution. Microsoft's own migration guidance says the quiet part: evaluate candidates "against your own application, prompts, and representative data. Compare quality, latency, and cost together rather than relying on public benchmarks alone."
Question 3.6 is a culture question wearing a process costume. Every vendor has a regression policy on paper. Very few have ever held a launch because of one. The answer to "has anyone ever done it" is the answer to whether the policy is real. How to test the eval artifacts once they arrive is the subject of the due diligence page.
Section 4: What does the availability commitment actually cover?
- 4.1State the availability commitment as a monthly uptime percentage, the measurement window, and who measures it.
- 4.2What counts as downtime? Specifically: do elevated error rates, rate-limit rejections, capacity exhaustion, and degraded-but-serving latency count?
- 4.3State latency commitments at p50, p95 and p99, separately for time-to-first-token and total completion, at a named token count.
- 4.4What is the remedy when you miss? Give the service-credit table.
- 4.5What are the exclusions? List every one.
- 4.6Does the SLA cover features in preview or beta? Which of the features in our scope are in preview today?
- 4.7What happens to our latency when your other customers spike, and what isolation do we get?
- 4.8Is there a capacity reservation available, what does it cost, and does it change the SLA?
Question 4.2 is where most availability commitments die. A 99.9% uptime number attached to a definition of downtime that excludes rate-limiting is a number about the vendor's servers, not about whether your feature worked. Rate-limit rejections are the most common way an LLM feature fails in production, and they are also the most commonly excluded.
Question 4.6 matters more in AI procurement than elsewhere because preview lifecycles are genuinely different. On Foundry, preview models launch with a "not sooner than" retirement date typically 90 days out, and when the decision comes, deployments are force-upgraded or terminated with 30 days' notice. There is no option to stay. Microsoft's own guidance is that preview models "aren't recommended for production workloads." If a vendor's differentiating feature depends on a preview model, that feature has a 30-day fuse and no SLA.
You will notice no uptime percentages or service-credit tiers for the major providers appear on this page. I could not retrieve them from the primary SLA pages in a form I would quote, and I am not going to paraphrase a contract. That is the reason question 4.4 exists. Make the vendor paste its own table into the response, under signature.
Section 5: What can you take with you when you leave?
This is the section that decides the procurement. Weight it accordingly.
It is the pre-purchase half of a pair. The recurring half is the quarterly independence audit in AI vendor capture risk, whose data sovereignty line requires that embeddings, fine-tuned weights and training artifacts stay exportable. That checklist states the requirement and audits your own estate against it. The questions below are what you put to a vendor before the requirement is yours to enforce.
- 5.1List every artifact we will create inside your product: conversation history, prompts and prompt versions, tool and function definitions, retrieval corpora, chunking configuration, embeddings, fine-tune training sets, fine-tune weights or adapters, eval sets, eval results, guardrail policies, routing rules, and audit logs.
- 5.2For each artifact in 5.1: is it exportable, through what interface, in what file format, and how long does a full export take at our data volume?
- 5.3For embeddings: which model produced them, what dimensionality, and are they exportable as raw vectors with their source-document identifiers intact?
- 5.4For fine-tunes: do we receive weights or adapter files? In what format? If not, state plainly that we do not.
- 5.5If we cannot take the weights, we can take the training data. Confirm we can export every training and evaluation set in the exact form used, including any transformations you applied.
- 5.6Are prompts, system instructions and tool schemas exportable as text or structured config, or are they only editable inside your interface?
- 5.7Give us the runbook for a customer who left. Not a policy statement. The actual sequence, with durations.
- 5.8Name a customer who has migrated off you and tell us how long it took.
Question 5.4 usually has a bad answer, and the bad answer is normal rather than scandalous. OpenAI's optimization documentation describes no path to download or export fine-tuned weights, and states that "All fine-tuned models will remain available for inference until their base models are deprecated." Read that as a lifespan clause. Your fine-tune expires when its base model does, on a schedule you do not set. Microsoft publishes the equivalent dates outright: a gpt-4o 2024-08-06 fine-tune has a deployment retirement date of 1 October 2027, after which inference and deployment return errors.
So 5.5 is the real ask. If the weights are not portable, the training data is the asset, and it needs to leave in the form you can retrain from.
For what a portable model artifact actually looks like, Amazon Bedrock Custom Model Import is the reference standard to point the vendor at. It accepts a .safetensors weights file, config.json, tokenizer_config.json, tokenizer.json and tokenizer.model in Hugging Face format, across named architectures including Llama, Mistral, Mixtral, Qwen and GPT-OSS, with imported weights under 200GB for text models and 100GB for multimodal, and maximum context under 128K. That is a concrete, checkable definition of "portable." Ask the vendor whether what they will hand you meets it. Any answer other than a file list is a no.
Question 5.8 is the highest-yield question in the entire document, and the tell is not the answer. It is the pause.
Section 6: How does the bill behave under load?
- 6.1Give the unit price and the unit. If tokens, state input, output, cached-read and cached-write prices separately.
- 6.2What is the price sensitivity of our workload? Show your own estimate of our monthly cost with the assumptions written out, including average input and output token counts.
- 6.3How does an agentic or multi-step workflow bill? One task can be dozens of model calls. Which of them do we pay for, including retries, tool-call round trips, and reasoning tokens.
- 6.4Do we pay for your retries? For failed or filtered generations? For requests that hit a guardrail?
- 6.5What discount mechanics exist and are they passed through: batch processing, prompt caching, provisioned or reserved capacity, committed spend?
- 6.6What are your price-change rights mid-term? State the notice period and any cap.
- 6.7If your upstream provider raises prices, who absorbs it?
- 6.8What happens at 10x our forecast volume: does unit price fall, does a rate limit bind, or do we get throttled?
Question 6.5 is worth real money and is easy to check upstream. Anthropic's Message Batches API cuts costs by 50% for asynchronous work, with most batches finishing in under an hour. Prompt caching prices cache reads at 0.1x the base input token price on most models (0.025x on Fable 5.1 and Mythos 5.1), with cache writes at 1.25x for the 5-minute cache and 2x for the 1-hour cache. A vendor whose product does heavy repeated-context work and does not use caching is billing you 10x the marginal rate on the cacheable portion. A vendor that uses caching and does not pass it through is keeping the difference. Both are legitimate business models. Neither should be a surprise in month four.
Question 6.3 is where agentic products get expensive in ways the demo does not show. Reasoning tokens, tool-call round trips and retry loops are all billable and all invisible in a scripted walkthrough. Ask for a token trace of a single representative task, end to end. Pair this section with the LLM cost calculator to model the numbers the vendor gives you, and with what frontier-model access should cost for the upstream price floor a vendor's markup sits on.
Section 7: Has anyone tried to break it?
- 7.1Has this system been red-teamed against prompt injection, indirect injection through retrieved content, and tool-use abuse? By whom, when, and do we get the report?
- 7.2What is the trust boundary between model output and any tool, API or database it can reach? Describe it as an architecture, not as a principle.
- 7.3What can the model do without a human in the loop? Enumerate every action, including writes.
- 7.4How is retrieved third-party content isolated from instruction-following?
- 7.5What is your vulnerability disclosure process, and your notification SLA to us for an incident affecting our data?
- 7.6Which certifications do you hold, with the audit date and the scope statement, not the badge.
- 7.7Under the EU AI Act, what is your role for our deployment: provider, deployer, or distributor? Which of our obligations do you take on contractually?
- 7.8If our use case is high-risk under Annex III, what documentation do you supply and by when?
Questions 7.7 and 7.8 have a calendar attached. The EU AI Act entered into force on 1 August 2024. Prohibitions and AI literacy requirements applied from 2 February 2025. General-purpose AI model obligations began on 2 August 2025, with models placed on the market before that date having until 2 August 2027 to comply. Most high-risk obligations, and the Article 50 transparency requirements, apply from 2 August 2026. High-risk systems operated by public authorities have until 2 August 2030. The detail behind 7.6 to 7.8 (EU AI Act, ISO 42001, NIST AI RMF) is on the AI compliance page; the practice behind 7.1 is AI red teaming.
If you are buying into a high-risk use case, the obligations are live now, and the contract is where the split of responsibility gets decided. A vendor that cannot name its role under the regulation has not done the work, and you will be doing it.
Question 7.3 is the one to read aloud in the room. Enumerating every unsupervised write action a model can perform tends to produce a shorter list than the product page implies, and occasionally produces a longer one than the vendor's own engineers expected.
How should you score the responses?
Weight exit and portability at least as heavily as capability. My split for a system that will hold production data:
| Area | Weight |
|---|---|
| Exit and portability (Section 5) | 30 |
| Model provenance and change policy (Section 1) | 20 |
| Data handling and training rights (Section 2) | 20 |
| Pricing mechanics under load (Section 6) | 15 |
| Security and red-teaming posture (Section 7) | 10 |
| Latency and availability (Section 4) | 5 |
| Capability (Section 3 and the demo) | Gate, not weight |
Capability gets scored separately, as a gate rather than a weight, because a vendor that fails the capability threshold is not in the comparison at all. That looks lopsided until you notice what the weights buy you. Capability is the thing you can re-test in a year, and re-testing is cheap. Portability is the thing that determines whether re-testing changes anything.
Which four clauses belong in the contract?
Everything above is diligence. These four clauses are what turn diligence into something a vendor has to honour.
- Model change notice. A stated minimum in days, with the right to remain on a pinned version through that window, and a termination right without penalty if you cannot get there.
- Export on demand. Not on termination. On demand, at any time, in the formats named in your Section 5 answers, with a stated maximum turnaround. An export right you can only invoke while leaving is a right you will invoke exactly once, badly, under time pressure.
- Eval access. The right to run your own evals against production, with a defined regression threshold and a remedy when it is crossed.
- No training, in the contract. Not in the trust centre, not in the FAQ. In the clause you sign, with the subprocessor list attached and a notice period for changes to it.
Section 7 of the AI policy template ("Vendor and third-party AI") sets the internal requirements these clauses verify, and the reversibility framework is the one-way-door test that Section 5 applies to a purchase.
What does this RFP not solve?
Two things this RFP will not save you from.
An excellent RFP response is still a description of intent. The vendor's answer to 1.4 describes the behaviour they believe their platform has. Half of the failures I have seen came from vendors who answered honestly about a default they had never verified. Ask for the configuration value, not the description of it, whenever a configuration value exists.
The second is that exit cost is not only technical. Once a workflow has been rebuilt around a product, the migration cost is mostly the retraining of people, and no export format touches that. Portability buys you the option. Whether you can afford to exercise it is a separate budget, and worth naming out loud before you sign, because it is the number that will actually decide.
AI Vendor RFP: FAQ
How long should an AI vendor RFP be?
Do vendors actually answer the exit questions?
Should we ask an early-stage vendor the same questions?
Is an AI vendor RFP an alternative to a security questionnaire?
What if the vendor is a thin wrapper over a frontier model?
Sources
All figures above were retrieved from these primary pages on 27 August 2026 and re-verified on 14 September 2026.
- Microsoft Foundry Models lifecycle and support policy: 18-month GA lifecycle (12 months for partner models), 60/30-day notice,
versionUpgradeOptionvalues, provisioned deployments not auto-upgraded, no retirement extensions, emergency retirement, preview lifecycle, fine-tune retirement dates, migration guidance - Anthropic model deprecations: 60-day notice commitment, Opus 4.1 deprecation and retirement dates
- OpenAI deprecations: 6-month / 3-month / 2-week notice tiers
- Anthropic API and data retention: Covered Models (Fable 5, Fable 5.1, Mythos 5, Mythos 5.1) 30-day requirement and 400 error, ZDR scope and exclusions, 2-year flagged-content retention, 6-year Activity Feed and session transcripts, 30-day code-execution container retention
- OpenAI data controls: no training since 1 March 2023 absent opt-in, ZDR behaviour and the non-eligible endpoint list
- Anthropic commercial terms: no training on Customer Content, customer owns Outputs
- Anthropic batch processing: 50% cost reduction, sub-1-hour typical completion
- Anthropic prompt caching: 1.25x / 2x write multipliers, 0.1x read multiplier, 5-minute and 1-hour TTLs
- OpenAI model optimization: fine-tuned models available until base model deprecation; no documented weight export
- Amazon Bedrock Custom Model Import: required artifacts, supported architectures, size and context limits
- EU AI Act implementation timeline: application dates by article
The answers arrived. Now check them.
An RFP collects claims. Due diligence tests them.