Buying access to a language model now works like buying compute did a decade ago: the question is never just which model answers best, it is whose hardware it runs on, under whose contract, and what happens to a prompt after it is sent. This ranking is aimed at the buyer who signs the actual contract, not the developer picking a model for a weekend project, and it judges each platform on where inference physically runs, whether prompts and completions train the vendor’s next model, what the token bill becomes once caching, provisioned capacity and fine-tuning are added, and how hard it would be to leave.
Visibility in this ranking can be paid for. Payment moves a vendor's position within the
shortlist; it never adds a vendor, and it never changes a word of the review. The largest vendors in generative AI platform
cannot hold places 1 to 3. How it works: placement disclosure ·
editorial process.
How we ranked these
How you actually get in: API key, cloud console, or your own hardware
Three access models sit on this list, and they are not interchangeable. A direct API vendor, OpenAI, Anthropic, Mistral AI, Cohere, Together AI and Heabsy among them, hands over an API key in minutes and bills per token from the first request. A hyperscaler platform, Google Vertex AI, Amazon Bedrock or Microsoft Azure AI Foundry, wraps that same kind of access inside a wider cloud console, which is faster to adopt for a team already billing through that cloud and slower for anyone starting from zero. A governed or on-premises option, IBM watsonx.ai or Databricks Mosaic AI, plugs model serving into infrastructure and permissions a company already runs, which is the right fit only when that infrastructure already exists.
What happens to a prompt after it is sent
This is the single fact worth checking before price, because it varies more than buyers assume. Anthropic and OpenAI both state that API inputs and outputs are not used to train models by default, with zero-retention terms available on request from Anthropic. Heabsy goes further on its own flagship model specifically: prompts and completions are never logged at all, a stronger guarantee than 'not used for training.' Platforms that route requests to third-party model providers, Amazon Bedrock, Azure AI Foundry, Together AI’s open-model catalogue and the non-flagship part of Heabsy’s own catalogue, inherit whatever policy that underlying provider sets, which is a materially different guarantee than a platform that owns the hardware end to end.
The real price: tokens, caching, and the capacity nobody demos
Every vendor here publishes a per-token rate card, but the number on that page rarely survives contact with a production workload. Prompt caching cuts the effective input price sharply on Anthropic, OpenAI and Heabsy for any workload that resends the same context repeatedly, which agentic coding tools do constantly. Provisioned or dedicated capacity, Amazon Bedrock’s provisioned throughput, Azure’s provisioned throughput units, Together AI’s dedicated GPU clusters, buys predictable latency but is billed hourly whether or not it is used, which can cost more than pay-as-you-go at low or spiky volume. Databricks Mosaic AI bills in consumption units that do not translate cleanly into a per-token figure at all, which is the detail a buyer comparing it against a plain API misses most often.
Model portability: what happens if you leave
Switching cost varies enormously across this list. An open-weight model licensed by Mistral AI or served by Together AI can, in principle, be redeployed on different infrastructure entirely, since the weights themselves are portable. A closed frontier model, GPT, Claude or Gemini, is not: leaving the vendor means re-prompting and re-evaluating against a different model, not moving a file. Routing the same closed model through a hyperscaler, Claude via Bedrock instead of Anthropic directly, changes the contract and the region without changing that underlying portability problem. A buyer who has not tested a second model against the same evaluation set has not actually checked how locked in they are, regardless of what the API looks like.
Who it’s for: three shapes of buyer, not one
A startup or product team optimizing for capability and ecosystem defaults correctly to OpenAI or Anthropic directly, where the newest model ships first and the most third-party tooling already exists. An enterprise already standardized on one cloud gets a faster procurement path and existing identity and billing integration through that cloud’s own platform, Vertex AI, Bedrock or Azure AI Foundry, even if the underlying models are the same ones available directly. A regulated buyer or one with a hard EU data-residency requirement needs to look past the marketing page to the actual processing region: Mistral AI, Heabsy and, for eligible workloads, the EU regions on the three hyperscalers are the options that hold up under that specific question, and a platform with no independent EU region, Anthropic and OpenAI sold directly, is not a fit no matter how strong the model is.
The 11 tools, reviewed
#1 Mistral AI
French frontier models you can rent by token or self-host · France · mistral.ai
EU processing Open weights Published token price
The only vendor on this list that publishes token prices, serves them from inside the EU by contract, and will licence weights a customer can run on its own hardware. That combination is what makes an exit possible, which is rare in this category. The models trail the largest American labs on the hardest reasoning benchmarks, the enterprise console is younger than the competition’s, and support outside France is thin.
Where it falls short
Models trail the largest US labs on the hardest reasoning benchmarks, the enterprise console is younger than OpenAI’s or Google’s, and support outside France is thin.
Wrong for
A buyer chasing the single strongest model on every benchmark regardless of where it runs; Mistral trades some frontier performance for EU processing and an open-weight exit.
Pricing: Per-token pricing published on La Plateforme, with batch (50% discount) and cached-token pricing; self-hosted enterprise deployment quoted separately (rate-limited free tier on la plateforme for testing; no perpetual free production tier)
Visit Mistral AI →
#2 Together AI
Inference and fine-tuning cloud for open-weight models, priced per token · United States · together.ai
Open models Fine-tuning Published price
An independent American inference cloud with one of the widest catalogues of open-weight models, per-token prices published on the website, and fine-tuning plus dedicated endpoints when shared capacity is not enough. Moving from a closed model to an open one here is mostly a change of endpoint. It is a US company without the EU processing commitments of Mistral or Scaleway, and it sells no frontier closed model of its own.
Where it falls short
A US company without the EU processing commitments of Mistral AI, and it sells no closed frontier model of its own, so it depends entirely on the pace of the open-weight ecosystem.
Wrong for
A buyer who specifically needs a closed frontier model like GPT or Claude; Together AI’s catalogue is open-weight models only.
Pricing: Serverless inference published per model, roughly $0.0015 to $4.50 per million tokens depending on the model; dedicated GPU clusters from about $3.99 to $9.99 per GPU-hour; fine-tuning priced per token (no published free production tier; billing is pay-as-you-go from the first request)
Visit Together AI →
#3 Cohere
Enterprise models you can deploy inside your own network · Canada · cohere.com
Private deployment Retrieval focus Non-US vendor
Built for the buyer who will not send text to a public endpoint: the same models run in a customer’s own cloud account or on its own servers, and the retrieval and reranking parts are what customers actually keep using. Behind the frontier labs on general reasoning, a smaller developer community, and private deployment is a quoted project rather than a self-serve price.
Where it falls short
Behind the frontier labs on general reasoning benchmarks, a smaller developer community than OpenAI’s or Anthropic’s, and private deployment is a quoted project rather than a self-serve price.
Wrong for
A small team that just wants the cheapest public API call; Cohere’s pricing and packaging are built around enterprise deployment, not hobbyist usage.
Pricing: Per-token pricing published for the hosted API; private and on-premises deployment quoted per project (rate-limited trial api keys for evaluation; not licensed for production use)
Visit Cohere →
#4 OpenAI Platform
The widest model range and the ecosystem everything targets first · United States · platform.openai.com
EU residency option Large ecosystem Published price
The default choice, and the one every library supports before it supports anyone else. API data is not used for training by default, and European data residency is offered for eligible endpoints on business plans. Against that, model deprecations arrive on OpenAI’s schedule rather than the buyer’s, and pricing and packaging have changed often enough to make a two-year forecast guesswork.
Where it falls short
Model deprecations and pricing or packaging changes arrive on OpenAI’s own schedule, which makes a multi-year cost forecast difficult, and it publishes no EU-only processing region of its own.
Wrong for
A buyer whose compliance team requires EU-only processing with no exceptions; OpenAI’s residency options cover only eligible endpoints on qualifying plans.
Pricing: Per-token pricing published per model family; enterprise agreements quoted for volume commitments (no standing free tier for the api; the separate chatgpt consumer product has its own free tier)
Visit OpenAI Platform →
#5 Anthropic Claude API
Frontier models with no training on API traffic by default · United States · anthropic.com
No training on inputs Published price Prompt caching
Strong on long documents and on code, with published prices and a default that API inputs and outputs are not used to train models; zero-retention terms can be arranged. Anthropic publishes no EU-only processing region of its own, so a European residency requirement means routing the same models through Bedrock or Vertex and paying that cloud’s margin on top.
Where it falls short
Publishes no EU-only processing region of its own, so a hard European residency requirement means routing the same models through Bedrock or Vertex and paying that cloud’s margin on top.
Wrong for
A buyer with a hard EU-only data residency requirement and no appetite to route through a hyperscaler; Anthropic has no independent EU region.
Pricing: Per-token pricing published per model tier (Haiku, Sonnet, Opus); batch and prompt-caching discounts available (limited free trial credit for new api console accounts; no perpetual free production tier)
Visit Anthropic Claude API →
#6 Google Vertex AI
Gemini and third-party models inside a European cloud region · United States · cloud.google.com
EU regions Model garden BigQuery link
The obvious option when the data already sits in BigQuery, and Google will contract for processing in European regions with sovereignty controls layered on top. The console is a maze, the same capability appears under two or three product names depending on which page a buyer lands on, and quota limits on new projects turn up without warning.
Where it falls short
The console reads as a maze, with the same capability sometimes appearing under two or three product names, and quota limits on new projects can appear without much warning.
Wrong for
A buyer outside the Google Cloud ecosystem who wants a lean, single-purpose API rather than a full cloud console.
Pricing: Per-token or per-hour pricing published per model; committed-use discounts for sustained volume (new google cloud accounts receive limited trial credit applicable to vertex ai; no perpetual free tier for production inference)
Visit Google Vertex AI →
#7 Amazon Bedrock
One API in front of a dozen different model vendors · United States · aws.amazon.com
Model choice EU regions AWS contract
Buys model portability inside one contract: Anthropic, Mistral, Meta and Amazon’s own models behind a single API, in Frankfurt, Ireland or Paris, with the European Sovereign Cloud as the stricter tier. The catch is that not every model reaches every European region on release, and provisioned throughput, the only route to predictable latency, is billed hourly whether or not it is used.
Where it falls short
Not every model reaches every European region on release, and provisioned throughput, the only route to predictable latency, is billed hourly whether or not it is used.
Wrong for
A buyer who wants the very newest model on day one in every region; Bedrock’s rollout of new models to non-US regions regularly lags the vendor’s own direct API.
Pricing: Per-token pricing published per model; provisioned throughput billed hourly for guaranteed capacity (no dedicated free tier for bedrock inference itself; standard aws trial credits may apply to new accounts)
Visit Amazon Bedrock →
#8 Microsoft Azure AI Foundry
OpenAI models under a Microsoft agreement and the EU Data Boundary · United States · azure.microsoft.com
EU Data Boundary Provisioned capacity Microsoft contract
For an organization already buying Microsoft this is the path of least procurement resistance: OpenAI models on an existing agreement, inside the EU Data Boundary, with Entra identity already attached. Capacity for the newest models is rationed by region, provisioned throughput units are expensive and sold in blocks, and the product has been renamed more than once in recent years.
Where it falls short
Capacity for the newest models is rationed by region, provisioned throughput units are sold in expensive blocks, and the product has been renamed more than once in recent years.
Wrong for
A buyer with no existing Azure or Microsoft footprint, who gains none of the Entra ID and billing integration that justifies the product’s added complexity.
Pricing: Per-token or provisioned-throughput-unit pricing published; enterprise agreements available (new azure accounts receive limited trial credit; no perpetual free tier for production inference)
Visit Microsoft Azure AI Foundry →
#9 Databricks Mosaic AI
Model serving that sits next to your already governed data · United States · databricks.com
Catalogue governance Fine-tuning Consumption billing
Makes sense when the training data and its governance already live in a lakehouse, because the serving endpoint inherits the same catalogue permissions. Fine-tuning and evaluation are part of the product rather than bolted on. It is a poor way to buy plain inference: consumption units are hard to translate into a per-token figure, and a buyer is adopting Databricks first.
Where it falls short
A poor way to buy plain inference on its own: consumption units are hard to translate into a per-token figure, and adopting Mosaic AI effectively means buying the wider Databricks platform first.
Wrong for
A team that just wants a model API and has no other Databricks footprint; the consumption pricing and governance model exist to serve lakehouse customers first.
Pricing: Consumption-based DBU (Databricks Unit) pricing on a published rate card; committed-spend discounts available (14-day free trial of the databricks platform; no perpetual free tier for production model serving)
Visit Databricks Mosaic AI →
#10 IBM watsonx.ai
Governed model platform for organizations that must show their work · United States · ibm.com
On-premise option Model governance Regulated sectors
The governance tooling, model documentation, drift monitoring and an audit trail an examiner will accept, is ahead of the field, and the whole platform can run on a buyer’s own hardware through Cloud Pak for Data. IBM’s own models are modest, third-party models arrive late, and almost nothing here gets bought without a services engagement attached.
Where it falls short
IBM’s own Granite models are modest next to frontier labs' offerings, third-party models typically arrive on the platform later than on their native APIs, and almost nothing here gets bought without a services engagement attached.
Wrong for
A team that wants to move fast with the newest model the day it ships; IBM’s governance-first approach and services-led sales motion favor caution over speed.
Pricing: Resource-unit (per-token) pricing published; software licence available for on-premises deployment via Cloud Pak for Data (ibm cloud lite includes a limited free allocation of watsonx.ai resource units)
Visit IBM watsonx.ai →
#11 Heabsy
Slovak inference API for open models, on the company’s own EEA hardware with zero data retention · Slovakia · heabsy.com
EEA-only flagship model OpenAI/Anthropic-compatible No minimum spend
Heabsy sells access to models rather than a model of its own: an OpenAI- and Anthropic-compatible API built so a team already using the OpenAI SDK, or running Claude Code, Cursor or Cline, can point its client at a new base URL and change nothing else. Its own flagship model, Qwen3.8 27B, runs on hardware the company owns inside the EEA with a documented zero-retention policy, and it publishes real throughput and latency numbers rather than marketing claims. It is a young platform, with no published founding date or team page, and only that one model carries the full EEA-only, zero-retention guarantee; the rest of the catalogue depends on third-party providers whose compute may sit outside the EEA.
Where it falls short
Only one model carries the full EEA-only, zero-retention guarantee; the rest of the catalogue depends on third-party providers whose compute may sit outside the EEA. Young platform with no published founding date or team page, and the flagship model’s 262,144-token context is smaller than some competitors on this list offer.
Wrong for
A buyer that wants an entire model catalogue under the same zero-retention guarantee, or one that needs a multi-million-token context window; only Heabsy’s own flagship model carries the strict guarantee, and its context tops out well below some rivals.
Pricing: Pay-as-you-go; flagship model (Qwen3.8 27B) at $0.04 per million input tokens, $0.04 per million cached input tokens and $0.30 per million output tokens; routed catalogue models priced separately from $0.04 per million input tokens; no minimum spend (no free tier; no credit card required to obtain a spending-limited api key)
Visit Heabsy →
What the data says about this market
Ownership here concentrates even more tightly than in most software categories on this site. Of the eleven platforms ranked, eight are US-headquartered, including all three hyperscaler platforms (Google, Amazon, Microsoft) and both leading closed-model labs (OpenAI, Anthropic), alongside Together AI, Databricks and IBM. Three sit outside the US: Mistral AI in France, Cohere in Canada, and Heabsy in Slovakia. That leaves a buyer who specifically needs a non-US-owned vendor, rather than merely an EU processing region from a US-owned one, with a genuinely short list.
Pricing has converged on per-token billing at the sticker-price level, but the real economics increasingly hinge on what sits around that headline rate. Prompt caching, which several vendors here now price at a fraction of the fresh-token rate, rewards exactly the repeated-context pattern that agentic coding and long-running assistants produce, and a platform’s cache-hit rate on a real workload matters more to the final bill than the rate card’s base price. At the other end of the market, provisioned and dedicated capacity, sold hourly rather than per token, exists specifically for buyers who have outgrown pay-as-you-go latency variance and are willing to pay for a reserved slice of hardware instead.
Enterprise adoption of generative AI moved from experimentation to mainstream use over 2024: 78% of organizations reported using AI in some form that year, up from 55% the year before, and global private investment in generative AI specifically reached $33.9 billion, an 18.7% increase on 2023 (Stanford HAI, AI Index Report 2025). That growth is what has pulled cloud providers, data platforms and now specialist inference vendors like Heabsy into a category that, three years ago, was effectively three companies.
The 11 ranked vendors, counted
- Headquarters by region: North America 9, Europe 2
- By country: United States 8, Canada 1, France 1, Slovakia 1
- Pricing model: Consumption units (DBUs) 1, Pay-as-you-go per token; no subscription 1, Per token (API) or quoted (private deployment) 1, Per token (API) or quoted (self-hosted licence) 1, Per token or per hour; committed-use discounts 1, Per token or provisioned units; enterprise agreement 1, Per token; batch and caching discounts 1, Per token; enterprise agreements available 1, Per token; hourly provisioned throughput 1, Resource units (per token) or on-premises licence 1, Usage-based (serverless) or hourly (dedicated GPU clusters) 1
- Free option: 14-day free trial of the Databricks platform; no perpetual free tier for production model serving 1, IBM Cloud Lite includes a limited free allocation of watsonx.ai resource units 1, Limited free trial credit for new API console accounts; no perpetual free production tier 1, New Azure accounts receive limited trial credit; no perpetual free tier for production inference 1, New Google Cloud accounts receive limited trial credit applicable to Vertex AI; no perpetual free tier for production inference 1, No dedicated free tier for Bedrock inference itself; standard AWS trial credits may apply to new accounts 1, No free tier; no credit card required to obtain a spending-limited API key 1, No published free production tier; billing is pay-as-you-go from the first request 1, No standing free tier for the API; the separate ChatGPT consumer product has its own free tier 1, Rate-limited free tier on La Plateforme for testing; no perpetual free production tier 1, Rate-limited trial API keys for evaluation; not licensed for production use 1
Counted from the 11 vendors on this page. More in our market data.
For the wider market behind generative AI platform, read our report The Global Shift to ICT Services,
or browse all industry reports.
Questions and answers
What is the best generative AI platform in 2026?
There is no single best platform because the category splits by what a buyer actually needs. OpenAI and Anthropic lead on raw model capability and ecosystem depth for a team building a new product. A buyer already standardized on AWS, Google Cloud or Azure gets a faster procurement path through that cloud’s own AI platform, on the same models plus a few of its own. A buyer with a hard EU processing requirement should start with Mistral AI or Heabsy rather than a US lab sold directly.
Is it better to use a model directly (OpenAI, Anthropic) or through a hyperscaler (Bedrock, Vertex, Azure AI Foundry)?
Direct usually means the newest model on day one and the simplest bill; the hyperscaler route usually means slower access to the newest release but existing identity, billing and data-residency controls already wired into infrastructure a company runs. Neither is more expensive by definition, since hyperscaler pricing generally mirrors the model vendor’s own rate card, but availability and region coverage genuinely lag on the hyperscaler side for the newest releases.
Which of these platforms are free to use?
None offers an unlimited free production tier, since running inference costs real GPU time. Most give a rate-limited or time-limited trial: Mistral AI’s La Plateforme has a rate-limited free tier for testing, Cohere issues free trial API keys not licensed for production, Databricks offers a 14-day platform trial, and IBM Cloud Lite includes a small free allocation of watsonx.ai resource units. Heabsy has no free tier at all but requires no credit card to obtain a spending-limited key, and the rest price from the first token.
Does using these APIs mean the vendor trains its next model on my data?
Not by default on most of this list. Anthropic and OpenAI both state that API inputs and outputs are not used for training by default, with stricter zero-retention terms available on request. Heabsy’s flagship model goes further: prompts and completions are never logged at all. Where a platform routes a request to a third-party model provider, most of Together AI’s catalogue, part of Amazon Bedrock’s and Azure AI Foundry’s, and the non-flagship part of Heabsy’s catalogue, the data-use policy is whatever that underlying provider sets, which is worth checking per model rather than assumed from the platform’s own headline policy.
What is Heabsy, and why is it on this list next to OpenAI and Amazon?
Heabsy is a Slovak inference platform: an OpenAI- and Anthropic-compatible API that lets a team point an existing SDK or agent tool at a new base URL. Its flagship model, Qwen3.8 27B, runs on hardware the company owns inside the EEA with zero data retention, while a wider catalogue of open models is routed through third-party providers under one contract. It is a much smaller and younger company than the hyperscalers and frontier labs on this list, which is reflected in its ranking, but it is a genuine, verifiable option for a buyer specifically weighing European data residency and pay-as-you-go pricing against the majors.
Which of these vendors are not US companies?
Mistral AI is French, Cohere is Canadian, and Heabsy is Slovak. OpenAI, Anthropic, Google (Vertex AI), Amazon (Bedrock), Microsoft (Azure AI Foundry), Together AI, Databricks and IBM are all US-incorporated, which puts customer data within reach of US legal process on a valid order regardless of which region physically hosts it. Whether that distinction matters depends on a buyer’s own compliance requirements, not on the software itself.
What is the difference between Amazon Bedrock, Azure AI Foundry and Google Vertex AI?
All three put several model vendors' catalogues behind one console and one cloud bill, but the catalogue and the integration differ. Bedrock’s roster spans Anthropic, Mistral, Meta and Amazon’s own Nova models. Azure AI Foundry is built specifically around OpenAI’s models under a Microsoft enterprise agreement, with Entra ID identity attached natively. Vertex AI centers on Google’s own Gemini models plus a model garden of open and partner models, with the deepest native link to BigQuery. The right one is usually whichever cloud a buyer’s data and identity already live in.
Why is Databricks Mosaic AI on a generative AI platform list rather than a data platform list?
Because it sells model serving, not just data infrastructure: Foundation Model APIs, fine-tuning and a AI Gateway sit directly on top of a Databricks lakehouse, and the same Unity Catalog permissions that govern a company’s data also govern who can call its models. It is included here specifically for buyers whose training and reference data already live in Databricks, for whom it is a materially different proposition than a standalone API.
Can I avoid vendor lock-in by choosing an open-weight model provider?
Partially. Mistral AI licenses open weights a buyer can self-host, and Together AI serves a wide catalogue of open-weight models over infrastructure it operates, so switching between models within that ecosystem, or moving to self-hosted infrastructure later, is genuinely easier than with a closed model. It does not eliminate lock-in entirely: prompts, evaluation harnesses and fine-tuned checkpoints are still tuned to a specific model’s behavior, and re-validating that work against a different model is real effort regardless of whether the weights are open.
How often is this ranking updated?
Whenever a fact underneath it changes: a pricing update, a new data-residency region, a model deprecation, or a platform changing its default training policy. The published and last-reviewed dates at the top of this guide are real, and a review means someone checked the vendor’s current documentation, not that a date was moved forward without a re-check.