Sooner or later a product team stops pasting text into a chat window and starts calling a model from code: a support tool that drafts replies, an invoice parser, a search box that answers in sentences. At that point the buyer is choosing an API, not an assistant, and the questions change. Which models can one key reach? What does a million tokens cost on the day traffic doubles? Does the prompt stay in memory or land in a log, and in which country does it run? This ranking covers twelve hosted providers a small or mid-sized company can sign up for today: three model makers, one cloud platform, two European infrastructure companies, and a set of inference and routing specialists. We read each vendor's own pricing, documentation, terms and privacy pages, and ordered them by what a buyer can confirm in writing before the first invoice.
Visibility in this ranking can be paid for. Payment moves a vendor's position within the
shortlist; it never adds a vendor, and it never changes a word of the review. The largest vendors in LLM API providers
cannot hold places 1 to 3. How it works: placement disclosure ·
editorial process.
How we ranked these
Setup: an API key takes minutes, a production quota takes longer
Every provider here hands out a key quickly, and most accept the OpenAI client library with a changed base URL: OpenRouter, Together AI, Fireworks AI, Groq, Scaleway, OVHcloud, Cohere and Google all document that route, Google still labeling it beta and Groq calling itself mostly compatible. Anthropic and Mistral document their own SDKs and request formats, though Anthropic also exposes an OpenAI-style endpoint. What slows a launch is everything after the key. OVHcloud caps anonymous calls at 2 requests per minute and authenticated ones at 400 per project and model. Cohere's trial key is free but limited. Bedrock adds an AWS account, IAM roles and a per-Region retention setting before the first call is clean.
The real price: tokens are cheap, the multipliers are not
All twelve bill input and output tokens separately, so the list price per million tokens tells only half the story. The rest sits in modifiers. Anthropic charges 1.1 times the standard rate when inference is pinned to the United States. Google's paid tier adds context caching and a 50 percent batch discount, while its free tier lets Google use submitted content to improve products. OpenRouter adds a platform fee of 5.5 percent on its standard plan and 8 percent on Business. Mistral prices Large at $0.5 per million input tokens and $1.5 output, with batch jobs at half. Run a month of real traffic through each meter before comparing, because caching and batching change rankings.
Getting your data out: what stays on the provider's side after the call
Portability is the easy half here: a prompt is text, and an OpenAI-style request body moves between vendors with little rewriting. The harder half is what the provider keeps. Scaleway, Fireworks and Groq state that inference content is not retained by default, with narrow exceptions: Scaleway may store a failing request for up to two weeks, and Groq may log inputs and outputs for up to 30 days when it detects misuse, with an opt-out. OpenAI keeps abuse-monitoring logs for up to 30 days unless a zero-retention arrangement applies. Files for batch jobs and fine-tuning datasets are the other holdings: Groq keeps batch files 30 days, and fine-tuned weights stay until the customer deletes them.
Independence from the vendor: one model maker or many
Lock-in here is mostly about models, not code. A provider that makes its own models, such as OpenAI, Anthropic or Cohere, can retire a version on its own schedule, and your prompts were tuned to that version. Providers that serve several open-weight families, including Together AI, Fireworks AI, Groq, Scaleway and OVHcloud, let a team move between models without changing the contract. OpenRouter goes furthest, with 500 plus models from more than 80 providers behind one key. Bedrock sits in between, carrying third-party models inside one AWS bill. A thin abstraction in your own code, with the base URL and model name in configuration, is the cheapest insurance.
Who it's for: match the provider to the constraint, not the benchmark
Start from the constraint that cannot move. A team that needs the strongest closed model picks the maker directly, OpenAI, Anthropic or Google, and accepts a single supplier. A team that wants to compare models weekly, or keep a fallback when one provider has an outage, is the audience for OpenRouter. A company bound by European data-protection duties should look at Scaleway, Mistral and OVHcloud first, since each contracts through a French entity. Latency-sensitive features such as voice or autocomplete suit Groq and Fireworks. Companies already standardized on AWS can keep procurement, billing and identity in one place with Bedrock, even if its model catalog moves slower than the makers' own.
The 12 tools, reviewed
#1 OpenRouter
A single API in front of 500+ models from 80+ providers · United States · openrouter.ai
Many models, one key Opt-in prompt logging EU in-region routing
OpenRouter takes first place because it removes the commitment that hurts most in this category: being tied to one model owner. The pricing page lists more than 500 models behind one key, and the standard plan adds routing preferences, budgets and caching. Its data page says prompts and responses are not stored unless the customer enables logging, and that retention is never on by default. Business and Enterprise customers can send requests to an EU or US host so that processing stays in that region. For a company that expects to swap models every few months, it is the cheapest way to keep that option open.
Where it falls short
The fee on top of model cost grows with spend, 5.5% on standard and 8% on Business. Each request passes through a middleman, so data handling also depends on whichever upstream provider is chosen, and OpenRouter says those providers differ on retention and training. EU in-region routing is not on the free or standard plans, and the free plan is capped at 50 requests a day.
Wrong for
A company that must contract directly with its model maker, or one sending regulated personal data without paying for in-region routing. Scaleway or a direct contract with the model owner is a cleaner fit.
Pricing: Pay as you go with a platform fee of 5.5% on the standard plan and 8% on Business; a free plan offers 25+ free models at 50 requests a day. (free plan)
Visit OpenRouter →
#2 Together AI
Open-weight model hosting with fine-tuning and dedicated capacity · United States · together.ai
Open-weight catalog Zero-retention switch Fine-tuning included
Together sits second because it is the plainest option for a team that wants open-weight models and a clean exit. Its docs describe an API compatible with the OpenAI REST API and SDKs across chat, completions, vision, image generation, speech and embeddings, so switching in takes a changed base URL. The terms put data control in a visible setting: answer No to storing prompts and allowing training and the account runs under zero data retention. The same company also trains and fine-tunes models, which helps a team that starts with prompts and later wants a tuned model.
Where it falls short
The pages we read did not state where Together's inference runs, so a buyer with residency duties must ask. Zero retention is a setting that applies only from the moment it is enabled and does not reach earlier data. The pricing page is a long list of per-model prices in several units, and we found no standing free tier in the pages read.
Wrong for
A European company that needs a stated hosting country, or a non-technical team that wants a chat interface and no code. Scaleway fits the first case; a chat product fits the second.
Pricing: Billed per million tokens with published per-model prices, cached-token discounts on some models, and dedicated capacity sold separately. (none)
Visit Together AI →
#3 Scaleway Generative APIs
Serverless open-weight model endpoints from a Paris data center · France · scaleway.com
Zero retention by default Paris-hosted Free tier
Scaleway is the strongest answer here for a buyer whose first question is where the prompt goes. Its data-privacy page says a zero-retention policy applies by default, that the content of inputs and outputs is not read or reused, and that data is not used to train base models. The FAQ says all serverless models are hosted in a data center in Paris. The contracting entity is a French simplified joint-stock company with a registered office in Paris. The API follows the OpenAI format, so existing client code works, and a free tier lets a team test before paying.
Where it falls short
The catalog is a curated set of open-weight models, so a team that needs the newest closed models from OpenAI or Anthropic will not find them. The FAQ says serverless models currently sit in one Paris site, and for single-region processing the vendor recommends the dedicated option. Failing requests may be stored for up to two weeks for debugging.
Wrong for
A team that must use a particular closed frontier model, or one that wants a gateway across many providers. A direct maker contract or OpenRouter is the better look.
Pricing: Billed per token on serverless endpoints with a free tier; dedicated deployments are available when a team wants isolation. (free plan)
Visit Scaleway Generative APIs →
#4 OpenAI Platform
The maker's own API for GPT models, audio and images · United States · platform.openai.com
Reference API EU data residency Zero-retention eligible
As the vendor whose API format became the standard, OpenAI is a safe default for a team that wants its own models and wide tooling support, but a large incumbent cannot take a top-three place. Its data-controls page says API data is not used to train models unless the customer opts in. Abuse-monitoring logs are kept for up to 30 days unless law requires longer, and a long list of endpoints, including chat completions, responses, embeddings and realtime, qualifies for zero data retention. Storage and processing are available in the US and in Europe (the EEA and Switzerland).
Where it falls short
Everything depends on one maker: model versions retire on its schedule, and a prompt tuned to one model rarely transfers unchanged. Retention defaults to a 30-day log, and zero retention and regional processing are arrangements rather than the starting point. The pages we read did not give a street address for the contracting company.
Wrong for
A company whose legal team requires an EU-only supplier, or one that wants to change models freely. Mistral, Scaleway or a gateway such as OpenRouter fits those cases better.
Pricing: Billed per token with separate input and output rates for each model; enterprise terms add residency and retention options. (none)
Visit OpenAI Platform →
#5 Anthropic Claude API
The maker's own API for Claude models · United States · anthropic.com
Own models US-only inference option Zero-retention arrangement
Anthropic sits at fifth because its terms are short and direct: they say Anthropic may not train models on customer content from the services. The data-residency page lets a request be pinned with inference_geo to the US or left global, and workspaces can restrict which geographies are allowed. A zero data retention arrangement is described in the docs, with a feature table showing what it covers. The Messages API is documented across several languages, and an OpenAI-compatible endpoint helps teams that want to test it from existing code.
Where it falls short
The only geographies documented are global and US, and the only workspace geography is the US, so a buyer who needs processing inside the EU cannot get it from the first-party API on these pages; the docs add that on Bedrock and Google Cloud the region follows the endpoint or profile chosen. Pinning to the US raises the price by 10 percent, and the pinning setting does not apply through the OpenAI-compatibility endpoint.
Wrong for
A European company that must keep inference inside the EU, or a team wanting open-weight models it could later self-host. Mistral or Scaleway are the places to start.
Pricing: Billed per token; US-only inference is priced at 1.1 times the standard rate on Claude 4.6 and later models. (none)
Visit Anthropic Claude API →
#6 Fireworks AI
Fast open-model inference with prepaid credits · United States · fireworks.ai
Zero retention for open models Prepaid credits OpenAI-format calls
Fireworks earns a mid-table place on one strong sentence: its docs say it does not log or store prompt or generation data for any open model without explicit opt-in, and that data exists only in volatile memory for the length of the request. The privacy notice repeats that inputs are not used to train models without opt-in. The base URL follows the OpenAI format. Billing is prepaid: a team buys credits, sets auto reload or a monthly limit, and cannot be surprised by an overage, which suits small companies and side projects with a firm budget.
Where it falls short
The zero-retention default covers open models only; the Responses API stores conversation data when store is true, which is its default. The pages we read did not state where inference runs, and we could not confirm the company's street address from its own pages, so only the city is given. No standing free tier appears on the pricing page.
Wrong for
A buyer who needs a closed frontier model or a stated hosting region in Europe. A maker's own API or Scaleway answers those needs more directly.
Pricing: Prepaid credits deducted per token, with Standard, Priority and Fast serving paths and a monthly spend limit option. (none)
Visit Fireworks AI →
#7 Groq
Very fast inference on custom hardware with a free plan · United States · groq.com
Fast inference Free plan No default retention
Groq's pitch is speed, and the docs back the privacy side: by default it does not retain customer data for inference requests, keeping only usage metadata. It lists text generation, transcription, speech output and image recognition on one platform, with prompt caching and structured outputs. The base URL follows the OpenAI format, which the docs describe as mostly compatible, with a short list of unsupported features. A free plan with published limits lets a team prototype first. For voice agents and autocomplete, where waiting a second is too long, it deserves a test run.
Where it falls short
The model list is smaller than the gateways and cloud platforms, and some OpenAI features are unsupported. Inputs and outputs may be logged for up to 30 days when misuse triggers it, unless the customer opts out. Batch files stay for 30 days and fine-tuning data until deleted. The legal address on its terms is a post office box in California.
Wrong for
A buyer who needs the newest closed models or EU-based contracting. OpenAI or Anthropic suits the first need; Mistral or Scaleway fits the second.
Pricing: A free plan with rate limits and pay-as-you-go billing per token on paid usage, with batch jobs available. (free plan)
Visit Groq →
#8 Mistral AI
French model maker with EU-hosted API and open-weight models · France · mistral.ai
EU-hosted by default Own open-weight models Batch discount
Mistral is a rare case of a model maker and a European contracting party in one. The privacy policy names a French company registered in Paris. Its help center states that data is hosted in the European Union by default, with a US API endpoint for those who choose it. The commercial terms say customer data is not used for training except in listed cases, such as opted-in features, feedback and preview models. The pricing page publishes per-token rates, a 50 percent batch discount and cached-input savings. A free plan and monthly API credits make a pilot inexpensive.
Where it falls short
Abuse monitoring runs on the API unless zero data retention is activated, and the help center says data can be temporarily transferred outside the EU depending on the feature, to the sub-processors in its Trust Center. Preview models can be used for training. It is one family of models, so a buyer wanting several makers behind one key needs a gateway.
Wrong for
A team that wants the broadest model choice behind one key, or one that insists on a specific closed model. OpenRouter or the maker's own API fits better.
Pricing: A free plan, paid plans with monthly API credits, and pay-as-you-go credits beyond them; Mistral Large is listed at $0.5 per million input tokens and $1.5 output. (free plan)
Visit Mistral AI →
#9 Google Gemini API
Google's own Gemini models via a free and paid API · United States · ai.google.dev
Free tier Context caching Batch discount
Google's API is the easiest place to start for free: the pricing page lists a free tier with free input and output tokens, and a paid tier with higher limits, context caching and batch pricing at half cost. The terms draw a hard line between tiers: on paid services Google says it does not use prompts or responses to improve its products, while the unpaid tier does. An OpenAI-compatible endpoint exists and the docs label it beta. Billing for paid services runs through Google Cloud, a plus for teams already there.
Where it falls short
On the free tier, submitted content is used to improve Google products, and the terms limit use for audiences in the EEA, Switzerland and the UK to paid services only. Paid services log prompts for a limited period to detect abuse, and grounding features keep data 30 days. We found no EU-only processing option on the pages read.
Wrong for
A European company sending personal data on a free key, or one needing a vendor other than a big platform. Move to paid terms, or look at Mistral or Scaleway.
Pricing: A free tier of limited models, and a paid tier billed per million tokens with context caching and a 50 percent batch discount. (free plan)
Visit Google Gemini API →
#10 OVHcloud AI Endpoints
Serverless open-model API from a French cloud provider · France · ovhcloud.com
French cloud No stored user data OpenAI specification
OVHcloud suits a company that already runs infrastructure with it and wants models inside the same account. The docs describe AI Endpoints as serverless and say user data is not stored, and that its LLM APIs follow the OpenAI specification, so existing client code works. The catalog lists open models such as gpt-oss, Llama and Qwen alongside embedding, speech and image models. The company is a French simplified joint-stock company registered in Roubaix. Anonymous requests work without an account, which allows a quick test before a Public Cloud project is created.
Where it falls short
Authenticated limits are 400 requests per minute per project and model, and keys from projects in Discovery mode without a payment method cannot use the service. The playground caps output at 1,024 tokens. The pages we read did not state the processing location of each model, and the model list is smaller than a gateway's.
Wrong for
A team that needs a wide closed-model catalog or large rate limits on day one. A gateway or the model makers fit better.
Pricing: Billed on a Public Cloud project with a payment method; anonymous use is limited to 2 requests per minute per IP and model. (free plan)
Visit OVHcloud AI Endpoints →
#11 Cohere
Enterprise language, embedding and rerank models · Canada · cohere.com
Embeddings and rerank Trial key Canadian vendor
Cohere is the retrieval specialist of the list: chat models alongside embeddings, which fits a company building search over its own documents. A trial key is free but limited, and production keys bill per token, with input and output priced apart. A compatibility API at api.cohere.ai/compatibility/v1 covers chat completions, embeddings and audio transcription for the OpenAI client libraries. The vendor is a Toronto company and its terms are governed by Ontario law. Enterprise customers can negotiate a data processing addendum and control how data is used for training.
Where it falls short
The privacy policy says inputs and outputs on the platform are generally retained 30 days for enterprise users, and that trial and research inputs may be used for improvement after de-identification. We found no zero-retention option in that policy. The compatibility API lists three endpoint groups only, and the model range is narrower than the gateways.
Wrong for
A team that needs zero retention without an enterprise contract, or one that wants the widest open-model choice. Scaleway, Fireworks or OpenRouter answer better.
Pricing: A free trial key with limits, and production keys billed per token with different input and output rates. (free trial)
Visit Cohere →
#12 Amazon Bedrock
Managed access to many makers' models inside AWS · United States · aws.amazon.com
AWS account Per-Region retention Third-party models
Bedrock suits a company whose procurement, identity and logging already live in AWS. Its documentation says model providers have no access to the accounts where models run, and therefore not to customer prompts or completions. The retention page lists modes from none, which writes nothing to durable storage, to aws_review, and lets an organization enforce a mode with IAM or service control policies. The same account can reach several makers through one bill. For a security team that wants guardrails in policy code rather than a vendor promise, that control is its main strength.
Where it falls short
Some models require retention for human review within the AWS boundary, up to 30 days, and zero retention for them needs approval through the AWS account manager. The setting is per Region and does not carry over. Setup takes an AWS account, roles and Region choices, which is heavy for a small team that wants a key and a curl command.
Wrong for
A small team without AWS experience, or one that wants the maker's newest model on release day. A direct maker API or OpenRouter is simpler.
Pricing: Billed through AWS per token by model and Region; the retention mode is set per Region at account or project level. (none)
Visit Amazon Bedrock →
What the data says about this market
Of the twelve providers ranked, eight are headquartered in the United States (OpenRouter, Together AI, Fireworks AI, Groq, OpenAI, Anthropic, Google and Amazon), three in France (Scaleway, OVHcloud and Mistral AI) and one in Canada (Cohere). No provider here is headquartered in Germany, the Netherlands or the Nordics, which means a buyer in those countries who wants a domestic supplier will not find one in this list. All twelve meter usage per token, and the vendors differ mainly in what they add on top: caching discounts, batch discounts, regional surcharges or a gateway fee.
Seven of the twelve let a developer make calls without a payment method or with a standing free allowance: OpenRouter (25 or more free models at 50 requests a day), Groq, Google, Mistral, Scaleway, OVHcloud (anonymous calls at a low rate) and Cohere (a trial key). The free tiers are not equal. Google states that on its free tier submitted content is used to improve its products, and that only paid services may be used when serving users in the European Economic Area, Switzerland or the United Kingdom. A prototype built on a free key can therefore be unusable in production for a European audience until the account is moved to a paid tier.
Where requests run is the clearest split. Scaleway serves its serverless models from a data center in Paris, Mistral hosts in the European Union by default and offers a separate US endpoint, OpenAI offers storage and processing in Europe for eligible customers, and OpenRouter offers EU in-region routing on its Business and Enterprise plans. Anthropic's documented geography options are global or US-only, and its workspace geography is US only. The category also sits inside a wider shift: ICT services rose from 9.1% of world service exports in 2013 to 14.48% in 2023 (World Bank), and a metered model call is a small unit of that trade.
The 12 ranked vendors, counted
- Headquarters by region: North America 9, Europe 3
- By country: United States 8, France 3, Canada 1
- Pricing model: Usage-based 9, Usage-based (prepaid) 1, Usage-based + plans 1, Usage-based + platform fee 1
- Free option: Free plan 6, None 5, Free trial 1
Counted from the 12 vendors on this page. More in our market data.
For the wider market behind LLM API providers, read our report The Global Shift to ICT Services,
or browse all industry reports.
Questions and answers
What is the best LLM API provider in 2026?
OpenRouter leads this ranking because one key reaches more than 500 models from over 80 providers, prompts are not stored unless the customer opts in, and a failed provider can be routed around. Together AI is the better pick for a team that wants open-weight models and a documented zero-retention setting. Scaleway is the better pick when the prompt must stay in the EU, since it applies zero retention by default and hosts serverless models in Paris. A team committed to one maker's models should contract with that maker directly.
Is there a European alternative to the OpenAI API?
Yes, three in this list contract through French companies: Scaleway, Mistral AI and OVHcloud. Scaleway hosts its serverless models in a Paris data center and applies zero retention by default. Mistral's help center says data is hosted in the EU by default, with a US endpoint as an option. OVHcloud says its AI Endpoints do not store user data. All three accept requests close to the OpenAI format, though Mistral documents its own SDK. None offers OpenAI's own models, so prompts may need retesting.
Can I switch providers without rewriting my code?
Often, for the basic chat call. OpenRouter, Together AI, Fireworks AI, Groq, Scaleway, OVHcloud, Cohere and Google all document an OpenAI-style endpoint where you change the base URL, key and model name. Groq calls its version mostly compatible and Google labels its beta. Features beyond plain chat, such as tool use, structured outputs, batch jobs and caching, differ in names and limits, so run your own test set against each. Keep model names in configuration to make the switch a small change.
Do these providers train on my prompts?
The pages we read say no by default for most, with exceptions worth knowing. OpenAI says API data is not used for training unless you opt in. Anthropic's terms say it may not train on customer content from the services. Scaleway, Fireworks and Groq describe no retention of inference content by default. Google's free tier does use submitted content to improve its products, and Cohere says trial and research inputs may be used after de-identification. Mistral excludes preview models and feedback from its no-training promise.
How long do providers keep my requests?
It ranges from nothing to 30 days. Scaleway and Groq keep no inference content by default, with narrow misuse exceptions of up to two weeks and 30 days. OpenAI keeps abuse-monitoring logs up to 30 days unless a zero-retention arrangement applies. Cohere says enterprise inputs and outputs are generally retained 30 days. Bedrock lets the customer set a retention mode per Region, though models that require human review keep content up to 30 days. Ask for the retention clause in writing and check which endpoints it covers.
How is an LLM API priced?
By the token, with input and output counted and priced separately. Providers then add modifiers: Google and Mistral discount cached input and batch jobs by half, Anthropic charges 1.1 times the standard rate for US-only inference, and OpenRouter adds a fee of 5.5 or 8 percent on top of model cost. Fireworks uses prepaid credits. Estimate from your own traffic: average prompt length, answer length, how often a prefix repeats and how many calls can wait for a batch. List prices alone mislead.
Which providers have a free tier?
OpenRouter offers more than 25 free models at 50 requests a day. Groq, Google, Mistral and Scaleway each describe a free plan or free tier, Cohere gives a limited trial key, and OVHcloud allows anonymous calls at 2 requests per minute. Free tiers carry data terms: Google uses free-tier content to improve its products and limits audiences in the EEA, Switzerland and the UK to paid services. Use a free key for prototypes and move to paid terms before real customer text goes through it.
When does a gateway like OpenRouter make sense?
A gateway fits when you compare models often, want an automatic fallback during a provider outage, or prefer one invoice for many vendors. The cost is a platform fee of 5.5 percent on the standard plan and 8 percent on Business, plus an extra hop in the data path. OpenRouter says retention depends on each upstream provider and gives controls to avoid providers that train. For a single stable model and large volume, a direct contract is cheaper.
Can I keep processing inside the EU?
Some can. Scaleway serves serverless models from Paris, Mistral hosts in the EU by default, OpenAI offers European storage and processing to eligible customers, and OpenRouter offers EU in-region routing on Business and Enterprise plans. Anthropic documents only global and US-only inference, with US the sole workspace geography. Several vendors did not state a processing location on the pages we read, including Together AI, Fireworks and OVHcloud for individual models. Get the location in the contract, not a marketing page.
Should I use the model maker or a cloud platform like Bedrock?
The maker gets new models first and sells the clearest terms for its own models. A cloud platform makes sense when your company already buys, authenticates and logs everything in that cloud. Bedrock keeps billing and IAM policy in AWS and says model providers cannot see prompts or completions, but some models require retained data for review, and setup assumes AWS skills. A small team without that background will ship sooner with a maker's API or a gateway.
What breaks when a provider retires a model?
The model name in your code starts returning errors, or a replacement answers differently, so prompts tuned for the old model degrade. Makers such as OpenAI, Anthropic and Cohere retire their own versions on their own schedules. Providers that host open-weight models can keep older families longer but may drop them too. Defend against it by keeping a fixed evaluation set of prompts, re-running it each quarter, and holding a second provider ready so a retirement notice becomes a planned migration.