Last fully reviewed: August 19, 2026
Evidence standard: buyer-run verification, negotiated commitments, and current primary sources—not vendor badges or sales assurances.
Choosing an AI inference provider is not a model leaderboard exercise. It is a combined architecture, security, privacy, reliability, capacity, commercial, and exit-risk decision. The strongest provider for a public summarization tool may be unacceptable for regulated records. The cheapest token price may become the most expensive option after retries, long outputs, support delays, unused commitments, and migration work.
The practical answer is to run procurement in this order:
- Define the workload and its non-negotiable constraints.
- Eliminate providers that fail a hard gate.
- Score only the survivors with workload-specific weights.
- Test finalist configurations under production-like load and failure.
- Convert the winning evidence into enforceable contract language.
- Keep a tested exit path before the first production request.
That sequence matters. A polished weighted score can hide a fatal privacy, residency, capacity, or contractual gap. Hard gates prevent that mistake.
Kingy verdict: There is no universal “best AI inference provider.” The defensible winner is the provider-service-model-region configuration that clears every mandatory gate, performs best on the buyer’s real tasks, and accepts enforceable operating and exit terms.
Download the complete procurement package
- Download the 41-page enterprise RFP and procurement guide (PDF)
- Download the editable RFP, scoring, POC, contract, and exit toolkit (Excel)
The downloadable package contains 32 RFP domains, a 16-gate knockout checklist, five editable weighting profiles, 13 POC metrics, contract and exit worksheets, and a source ledger. The article below explains how to use it.
Start with the decision, not the vendor list
An inference service is not one indivisible product. The relevant unit of evaluation is the exact combination of:
- contracting entity and reseller or marketplace path;
- service and purchasable tier;
- model family and pinned model version;
- API endpoint and enabled features;
- processing, storage, logging, and backup regions;
- shared, priority, provisioned, or dedicated capacity mode;
- support plan and incident path;
- data-use, retention, and human-review configuration.
A provider-level answer such as “supports zero retention” or “available in Europe” is too broad. One feature may create application state while another is stateless. One deployment type may remain inside a selected geography while a global route may not. A public rate limit may be a ceiling, not reserved capacity. A service-level agreement may exclude quota exhaustion, previews, specific models, or failures outside the provider’s defined control.
Write the intended configuration at the top of the RFP and require every answer to apply to that scope. If a vendor cannot say which entity, service, model, feature, tier, and region its answer covers, treat the answer as unresolved.
Use an evidence ladder
Procurement teams often collect a large volume of evidence without ranking its strength. A certification logo, a sales email, a documentation page, a signed amendment, and a controlled deletion test are not equivalent.
Scroll or swipe to compare →
| Rank | Evidence | How to use it |
|---|---|---|
| 1 | Verified technical behavior | Repeatable buyer-run test, configuration export, API response, telemetry, logs, or controlled failure exercise for the exact scope. |
| 2 | Written contractual commitment | Executed order form, DPA or BAA, SLA, security schedule, or negotiated term that clearly covers the exact configuration. |
| 3 | Current official documentation | Dated vendor documentation or policy. Useful, but changeable unless incorporated into the agreement. |
| 4 | Vendor assertion | Questionnaire response, sales email, roadmap statement, certification logo, or unsupported attestation. |
| 5 | Assumption or inference | Architectural guess or extrapolation from another service. Never treat this as a pass. |
The evidence ladder prevents two common failures. First, it stops procurement from treating “the vendor said so” as equivalent to a tested behavior or binding term. Second, it exposes where a promising product still depends on assumptions.
For every material answer, record the source URL or document, publication or execution date, retrieval date, exact product scope, evidence rank, owner, and refresh date. Screenshots without a stable source or timestamp are weak evidence. A link without a quoted proposition is difficult to audit. “SOC 2” without the report period, entity, service scope, bridge coverage, and exceptions is not a complete control answer.
Classify the workload before setting thresholds
The same provider can be acceptable for one workload and a knockout failure for another. Build a short workload dossier before sending the RFP.
Data and regulatory profile
Identify the highest data class that may reach the service: public, internal, confidential, personal, regulated, privileged, health, payment, export-controlled, or customer-restricted. Include prompts, outputs, embeddings, files, fine-tuning data, tool calls, metadata, logs, cached context, abuse-monitoring records, support tickets, and backups. Hidden data paths are still data paths.
State the applicable jurisdictions and contractual obligations. HIPAA, financial-services rules, public-sector requirements, employment law, data-localization commitments, and customer agreements may create different hard gates. For health data, for example, HHS explains that a cloud provider maintaining electronic protected health information can be a business associate even when it cannot view encrypted content, and the required relationship must be covered by an appropriate BAA. See HHS guidance on HIPAA and cloud computing and business associate contracts.
This guide is an operational procurement aid, not legal advice. Qualified counsel should review applicable law, transfer mechanisms, negotiated terms, sector obligations, and the final agreement.
Availability and recovery profile
Define business impact by outage duration. Record the minimum acceptable service, target availability, recovery time objective, recovery point objective, regional-failure posture, and whether degraded model quality is acceptable during failover. If the workload is safety-critical or materially affects customers, “we can switch providers” is not a recovery plan until the alternate has been tested for quality, capacity, data flow, and operational readiness.
Load and latency profile
Measure input and output tokens separately. Record steady and peak requests per minute, tokens per minute, concurrent streams, context-length distribution, output-length distribution, time-to-first-token, time-to-last-token, batch needs, tool use, and seasonal or launch spikes. Include expected growth and a headroom factor.
Do not confuse a documented rate limit with guaranteed capacity. A limit describes what the service may allow; it does not necessarily promise that capacity will be available when many customers are competing for it. Reserved or provisioned offerings must be evaluated for model and regional coverage, activation lead time, overage behavior, commitment duration, scaling mechanics, and remedies.
Quality and safety profile
Define representative tasks, failure classes, evaluation datasets, human-review rules, prohibited outputs, fairness or safety requirements, and the cost of a wrong answer. Benchmark averages are not a substitute for task-level acceptance. A model that wins on general reasoning can still fail document extraction, code transformation, structured output, grounded support, or a specific language mix.
The 16 hard gates
Hard gates are binary. A provider either supplies acceptable evidence for the exact proposed configuration or it does not advance. “Partially meets” belongs in remediation only when the buyer has explicitly accepted a time-bound condition before award.
Scroll or swipe to compare →
| Gate | Pass condition | Knockout example |
|---|---|---|
| Required service scope | Exact model version and every required endpoint are generally available in the required region and tier. | A critical feature is preview-only, unavailable in-region, or outside the SLA. |
| Minimum capacity and latency | Demonstrated or contracted capacity meets peak load plus headroom, and p95/p99 latency passes under load. | A public maximum rate limit is offered as proof of guaranteed capacity. |
| Data processing and retention | Content, metadata, state, logs, caches, support paths, and backups meet the workload’s retention rule. | An undisclosed or incompatible retention exception remains. |
| Training and product improvement | Buyer data is excluded from training, tuning, evaluation, improvement, and human review except for explicit, narrow instructions. | The restriction is only a mutable dashboard setting or excludes a required feature. |
| Residency and transfer | Every processing and storage path satisfies the approved geography and transfer mechanism. | A global route or support process can move covered data outside the approved boundary. |
| Security baseline | Identity, least privilege, key management, encryption, network controls, logging, vulnerability management, and independent assurance cover the exact service. | The report or certificate covers a different entity or omits a material service. |
| Incident response | Notification, severity, containment, evidence preservation, RCA, and regulator/customer support meet the buyer’s timetable. | The vendor only promises notice under a vague “without undue delay” standard. |
| Availability commitment | The covered service, regions, measurements, exclusions, credits, and claim mechanics meet the business requirement. | Credits are the sole remedy but do not address chronic failure or termination. |
| Business continuity | Regional and control-plane failure plans meet RTO/RPO and can be exercised. | The architecture has a single undocumented dependency or no recovery evidence. |
| Model lifecycle | Version pinning, change notice, deprecation period, rollback, and migration support match the workload. | The provider may silently update behavior or retire a model without workable notice. |
| Output rights and IP | The agreement provides necessary rights, allocates infringement risk, and contains workable indemnity conditions. | Essential output use is constrained or exclusions make protection illusory. |
| Subprocessors and supply chain | Current subprocessors, locations, functions, notice, objection rights, and flow-down duties are acceptable. | A critical processor or location is undisclosed or change rights are ineffective. |
| Audit and evidence rights | Reports, testing summaries, bridge evidence, remediation, and audit escalation are sufficient. | The buyer must accept a stale report with no bridge or remediation visibility. |
| Commercial transparency | Rates, meters, rounding, caches, batches, tools, commitments, taxes, support, and change controls are clear. | Material charge categories or unilateral price-change rights remain open. |
| Support and escalation | 24×7 severity path, response/update/restore targets, named ownership, and RCA timing meet need. | Marketplace and provider each disclaim responsibility for the same incident. |
| Exit feasibility | Data export/deletion, model and prompt portability, assistance, commitment relief, and tested alternatives are workable. | The buyer cannot retrieve needed artifacts or exit a failing commitment. |
The gate register should include the requirement, proposed threshold, vendor answer, evidence, evidence rank, reviewer, result, exception owner, expiry date, and award condition. A waiver without an owner and expiry date tends to become permanent risk.
Ask RFP questions that force testable answers
Broad prompts produce broad answers. “Describe your security” invites marketing copy. A better question specifies the service scope, requests evidence, and asks what is excluded.
Use these question patterns across the full questionnaire:
- Scope: Identify the legal entity, service, model versions, endpoints, features, tiers, regions, and partner platforms covered by every answer. List exclusions.
- Architecture: Provide a data-flow diagram showing processing, storage, logs, safety review, application state, caches, backups, support access, subprocessors, and cross-region paths.
- Data use: State whether prompts, outputs, files, embeddings, metadata, or feedback are used for training, evaluation, product improvement, safety, or human review. Separate defaults from configurable and contractual controls.
- Retention: Give the duration and deletion behavior for each data type and feature. Explain legal holds, backup deletion, support records, flagged-content exceptions, and certificate availability.
- Residency: Identify every processing and storage country for normal operation, failover, support, telemetry, abuse review, and backups. Explain how the buyer can verify routing.
- Identity and access: Describe SSO, SCIM, MFA, RBAC, service identities, key rotation, IP or network restrictions, privileged access, separation of duties, and audit logs.
- Assurance: Provide current SOC or ISO evidence, report periods, scope, exceptions, bridge coverage, penetration-test summary, vulnerability disclosure, and remediation status.
- Capacity: State available and contractible RPM, input TPM, output TPM, concurrency, context, batch, and streaming limits. Explain reservation, burst, ramp, throttling, scaling lead time, and priority treatment.
- Reliability: Define uptime calculation, time windows, error classes, latency treatment, regional aggregation, exclusions, maintenance, credits, claim windows, and chronic-failure rights.
- Lifecycle: Explain model pinning, version changes, default aliases, notice periods, retirement, rollback, evaluation support, and partner-platform differences.
- Commercials: Provide unit rates and meters for input, cached input, output, reasoning or hidden tokens, tools, storage, batches, fine-tuning, provisioned capacity, support, egress, taxes, and minimum commitments.
- Exit: Describe data return and deletion, configuration export, logs, fine-tuned artifact portability, prompt and evaluation export, transition help, continued service, commitment relief, and post-termination verification.
Require a response format of answer, exact scope, evidence, exception, proposed contractual commitment, and owner. If the vendor marks a question confidential, use a controlled data room rather than accepting no evidence.
Score only after the gates pass
Weighted scoring is useful for choosing among viable finalists. It is dangerous when used to average away a knockout issue.
A defensible scoring method has four parts:
- A 0–5 score with anchored definitions.
- A workload-specific weight for each domain.
- An evidence-confidence multiplier.
- Separate treatment of price and residual risk.
One practical formula is:
Adjusted domain score = weight × normalized score × evidence confidence
For example, a score of 4 supported only by a vendor assertion should not equal a score of 4 supported by a test and negotiated commitment. A simple confidence scale might assign 1.00 to buyer-verified evidence or a strong signed obligation, 0.85 to current official documentation, 0.65 to a detailed vendor assertion, and zero to an unverified assumption for a mandatory requirement.
Use anchored scoring:
- 0 — Unacceptable: missing, contradicted, or fails the requirement.
- 1 — Material deficiency: major gap with no credible near-term remedy.
- 2 — Below requirement: partial capability or weak evidence; substantial remediation needed.
- 3 — Meets requirement: acceptable capability and evidence for the proposed scope.
- 4 — Exceeds requirement: stronger control, performance, or terms with solid evidence.
- 5 — Materially superior: differentiated result demonstrated on the buyer’s workload or locked into unusually strong terms.
Weights should change by workload. A regulated profile may emphasize privacy, residency, security, and auditability. A latency-critical interactive product may emphasize quality at latency, predictable capacity, streaming behavior, and regional performance. A bursty public product may favor elasticity and cost efficiency. A continuity-critical system should increase reliability, lifecycle, support, and exit weights. Never reuse one “enterprise” scorecard without explaining what it optimizes.
Calculate effective cost, not list price
Token prices are inputs to a model, not the procurement conclusion. Estimate cost for representative task distributions and include:
- input and output tokens;
- cached-input eligibility and hit rate;
- reasoning or otherwise billable hidden tokens where applicable;
- tool calls, search, code execution, image or audio processing;
- batch discounts and latency trade-offs;
- failed requests, retries, fallbacks, and duplicate work;
- prompt expansion and output verbosity;
- provisioned-capacity commitments and unused capacity;
- support plans, egress, storage, taxes, and marketplace margin;
- engineering, observability, evaluation, and migration effort.
Normalize quality before comparing cost. If one configuration needs more retries, human review, larger prompts, or a second-pass model to achieve the required outcome, its apparent token discount may disappear.
Run sensitivity cases for steady load, expected load, peak load, growth, lower cache hit rate, longer output, degradation, and commitment under-utilization. Report cost per successful business task, not only cost per million tokens.
Run a production-like proof of capability
A good POC is an acceptance exercise, not a demo. Two to four weeks is usually enough to test finalists if the dataset, harness, and thresholds are prepared in advance.
Freeze the test plan
Before the first run, record:
- exact model identifiers and versions;
- API parameters, system prompts, tools, and structured-output schemas;
- dataset version, sampling method, and prohibited test leakage;
- regions, capacity modes, accounts, and network path;
- runtime, SDK, retry, timeout, and fallback configuration;
- evaluators, scoring rubrics, and human-review protocol;
- load shapes and failure injections;
- acceptance thresholds and decision rules.
Use the same eligible prompts across providers, but do not force identical settings when platforms expose materially different controls. Document every difference.
Measure the outcomes that matter
At minimum, collect:
- task success and severity-weighted error rate;
- groundedness, citation correctness, or extraction accuracy where applicable;
- structured-output validity and tool-call success;
- safety/refusal correctness for permitted and prohibited requests;
- time to first token and time to complete response at p50, p95, and p99;
- input/output throughput and concurrency behavior;
- 429s, capacity 503s, other 5xx errors, timeouts, truncations, and retry exhaustion;
- effective cost per successful task;
- regional routing, storage, logging, retention, and deletion conformance;
- failover recovery time and degraded-service quality;
- model-update, rollback, and version-pin behavior;
- support response, escalation quality, and RCA performance during a planned exercise.
Do not combine latency samples from materially different output lengths without normalization. Separate provider errors from client timeouts and quota rejections. Preserve raw request IDs and timestamps without retaining sensitive content unnecessarily. Use synthetic or properly authorized test data.
Test failure, not only success
Inject rate spikes, region loss, endpoint errors, malformed outputs, long contexts, tool failure, provider timeout, credential rotation, model retirement simulation, and deletion requests. Observe retry storms and cascading cost. Confirm that the fallback provider has sufficient capacity and acceptable quality before calling it resilient.
No statement in this article represents a Kingy hands-on benchmark of a specific provider. The POC framework tells buyers how to generate their own controlled evidence.
Turn POC evidence into contract language
The commercial negotiation should not start from a blank page after technical selection. Translate every material acceptance criterion into the order form, SLA, DPA, BAA, security schedule, support plan, or amendment with the correct order of precedence.
Prioritize these positions:
- Scope and precedence: name the exact services, features, regions, model versions, commitments, and documents. Make negotiated terms prevail over conflicting click-through or online terms.
- Capacity: state minimum available RPM, input/output TPM, concurrency, or reserved units; burst and ramp behavior; scaling lead time; effective date; and remedy for capacity failure.
- SLA measurement: define clock, denominator, errors, regions, retries, exclusions, data source, reporting cadence, audit/recalculation rights, and claims automation.
- Chronic failure: add termination, commitment relief, and transition support after repeated failures; credits alone are rarely sufficient.
- Data use and retention: impose purpose limitation, training restrictions, feature-level retention, deletion requirements, backup treatment, human-review boundaries, legal-hold notice, and evidence of completion.
- Residency and transfers: identify allowed locations and transfer mechanisms, cover support and failover, require notice of changes, and preserve meaningful objection or termination rights.
- Security and incidents: specify controls, subprocessor flow-down, notification deadlines, cooperation, evidence preservation, RCA, remediation, and responsibility allocation.
- Model lifecycle: require version pinning where feasible, material-change notice, deprecation windows, rollback, migration support, and relief when a replacement does not meet acceptance criteria.
- IP and outputs: clarify input/output rights, confidential information, infringement handling, indemnity scope, exclusions, defense control, and caps.
- Pricing: lock meters, rates, tiers, rounding, discounts, commitment treatment, audit rights, change notice, and renewal mechanics.
- Exit: require data and configuration export, deletion, continued service during transition, assistance rates, commitment relief, and survival of confidentiality, IP, audit, and assistance duties.
Legal review remains essential. Published provider terms are useful evidence of the default offer, but the signed document set controls the deal.
Read provider documentation at feature level
Current official documentation shows why generic provider labels are unreliable.
- OpenAI documents endpoint-level data controls, including abuse-monitoring retention, Zero Data Retention and Modified Abuse Monitoring scope, application-state behavior, and regional processing considerations. Review OpenAI’s data controls, Services Agreement, and subprocessor list against the exact features in use.
- Anthropic separately documents commercial retention, Zero Data Retention scope and exclusions, rate-limit semantics, service tiers, and model deprecations. See its commercial data-retention guidance, ZDR scope, rate limits, and model deprecations.
- AWS distinguishes service protection, SLA scope, cross-region inference, and Provisioned Throughput. Review Amazon Bedrock data protection, the Bedrock SLA, cross-region inference, and Provisioned Throughput.
- Microsoft describes data and safety-monitoring paths plus global, data-zone, regional, standard, provisioned, and batch deployment types. Review data, privacy, and security for Azure Direct Models, deployment types, and the model retirement schedule.
- Google Cloud documents feature-specific retention considerations, dynamic shared quota, Provisioned Throughput, SLA scope, service-specific terms, and its data processing addendum. Review Vertex AI and zero data retention, throughput quota, Provisioned Throughput purchasing, the Vertex AI SLA, and the Cloud Data Processing Addendum.
These links do not establish that one provider is safer or more reliable. They demonstrate that retention, routing, capacity, lifecycle, and contract answers depend on a specific product configuration. Refresh all material sources immediately before award and again before renewal.
The NIST AI Risk Management Framework and NIST Generative AI Profile provide useful governance and risk-management structure. They do not replace workload testing, legal analysis, or negotiated obligations.
Design the exit before signing
Exit planning is an architecture requirement. It should not begin after a deprecation notice, security event, chronic capacity failure, or price increase.
Maintain these minimum capabilities:
- an abstraction boundary around provider-specific APIs, schemas, tools, and authentication;
- versioned prompts, policies, evaluation sets, routing logic, and configuration;
- portable logs and telemetry that do not depend on one vendor’s console;
- documented model-specific behaviors and fidelity gaps;
- a tested alternate provider or model for critical tasks;
- data inventory, export, deletion, and legal-hold procedures;
- emergency degraded-service and planned-migration runbooks;
- capacity lead times and commercial prerequisites for the alternate;
- trigger thresholds for SLA failure, adverse policy or subprocessor change, price increase, deprecation, security event, financial distress, or strategic need;
- a maximum migration time for emergency and planned paths.
Test at least one component of the exit plan quarterly. A paper architecture that has never replayed prompts, schemas, tools, and safety policies on an alternate provider is not a proven exit route.
Keep the decision current
Inference services change faster than ordinary enterprise contracts. Model aliases move, features gain state, data controls change scope, subprocessors change, new regions open, preview capabilities become generally available, and rate or lifecycle terms change.
Use an operating cadence:
- Monthly: incidents, status history, capacity headroom, spend and commitment utilization, invoice reconciliation, unresolved vulnerabilities, support performance, and fallback readiness.
- Quarterly: task evaluations against an alternate, source-ledger refresh, subprocessor and policy changes, regional feature matrix, access review, and an exit-drill component.
- After material change: targeted gate review and regression testing for new model versions, endpoints, regions, tools, safety controls, data paths, or terms.
- Annually and before renewal: full gates and scorecard, current assurance reports, bridge evidence, DPA/BAA/SLA/rate card, business continuity, market check, and negotiation plan.
Assign one accountable owner for each unresolved issue. Record accepted residual risk separately from vendor score. A high score does not approve a risk; an authorized business owner does.
The procurement decision record
The final recommendation should be short enough for an executive to understand and detailed enough for an auditor to reproduce. Include:
- Proposed provider, legal entity, service, model version, tier, regions, capacity mode, support plan, and contract path.
- Workloads approved and explicitly excluded.
- Hard-gate results and any time-bound conditions.
- Weighted scores, evidence-confidence adjustments, and sensitivity cases.
- POC results with versioned test configuration.
- Three-year effective-cost range and commitment exposure.
- Negotiated protections and remaining deviations.
- Residual risks, owners, expiry dates, and monitoring.
- Exit plan, alternate configuration, and migration time objective.
- Approval signatures and next review date.
The goal is not to predict which AI platform will dominate. It is to make a bounded, evidence-backed decision that remains reversible when technology, terms, or business needs change.
Frequently asked questions
Is an AI inference RFP the same as a model benchmark?
No. A benchmark measures selected model behavior under a particular test. An enterprise RFP also evaluates the serving layer, data handling, identity controls, assurance scope, regional architecture, capacity, support, pricing mechanics, contracts, lifecycle, and exit. Benchmark evidence belongs inside the quality portion of the decision, not in place of the broader procurement process.
Does zero data retention mean no data is ever stored?
Not necessarily. Scope can vary by service, feature, endpoint, abuse-monitoring arrangement, application-state function, file or cache behavior, support path, and contract. Ask for a feature-level retention matrix and test or inspect the exact configured service. Separate request-content retention from account, billing, security, and operational metadata.
Should an enterprise choose one provider or multiple providers?
It depends on business impact and operating maturity. Multi-provider capability can reduce concentration and lifecycle risk, but it adds evaluation, routing, observability, security, and commercial complexity. A useful minimum is a primary configuration plus a tested alternative for critical tasks. Active-active routing is not automatically better than a well-tested, capacity-ready recovery path.
How many providers should reach the POC?
Usually two or three. Hard gates and documented capability should narrow the field before expensive testing. Sending six weakly screened vendors into a POC tends to reduce test depth and delay contracting. Keep a reserve candidate if a finalist fails the evidence or negotiation stage.
How often should the RFP be repeated?
Run a full reassessment before renewal and after a material change in workload, law, geography, provider architecture, model lifecycle, data controls, or risk classification. Between full reviews, use monthly operational monitoring and quarterly source, alternate-provider, access, and exit checks. The right cadence follows risk and change—not a generic annual calendar alone.
Editorial disclosure: The featured image is an AI-generated conceptual illustration created for Kingy.ai. It is not a provider interface, measured benchmark, certification, or product photograph. No provider paid for inclusion in this guide. Provider documentation was reviewed as primary evidence of published positions; buyers should verify current scope and negotiate their own terms.
