Skip to content
AI software · 12 vendors ranked

Best LLM API Providers in 2026

Sooner or later a product team stops pasting text into a chat window and starts calling a model from code: a support tool that drafts replies, an invoice parser, a search box that answers in sentences. At that point the buyer is choosing an API, not an assistant, and the questions change. Which models can one key reach? What does a million tokens cost on the day traffic doubles? Does the prompt stay in memory or land in a log, and in which country does it run? This ranking covers twelve hosted providers a small or mid-sized company can sign up for today: three model makers, one cloud platform, two European infrastructure companies, and a set of inference and routing specialists. We read each vendor's own pricing, documentation, terms and privacy pages, and ordered them by what a buyer can confirm in writing before the first invoice.

What it is: An LLM API provider hosts large language models and lets a developer call them over HTTP: send a prompt, receive generated text, pay per token. Some providers train their own models, others run open-weight models on their own GPUs, and gateways route one request to many upstream providers. The product is the endpoint plus its keys, rate limits, billing meter, logging settings and the contract that says where inputs are processed and how long they are kept.

Visibility in this ranking can be paid for. Payment moves a vendor's position within the shortlist; it never adds a vendor, and it never changes a word of the review. The largest vendors in LLM API providers cannot hold places 1 to 3. How it works: placement disclosure · editorial process.

The top three

  1. #1

    OpenRouter

    Teams comparing models or needing failover

    One key and one endpoint reach hundreds of models, prompts are not stored unless the customer opts in, and a fallback exists when a single provider goes down.

  2. #2

    Together AI

    Developers on open-weight models

    Hosts a wide catalog of open-weight models with OpenAI-compatible calls, a documented zero-retention switch, and fine-tuning on the same account.

  3. #3

    Scaleway Generative APIs

    EU companies with data-residency duties

    A French provider that applies zero data retention by default, serves models from a Paris data center, accepts OpenAI-style calls and offers a free tier.

How we ranked these

Setup: an API key takes minutes, a production quota takes longer

Every provider here hands out a key quickly, and most accept the OpenAI client library with a changed base URL: OpenRouter, Together AI, Fireworks AI, Groq, Scaleway, OVHcloud, Cohere and Google all document that route, Google still labeling it beta and Groq calling itself mostly compatible. Anthropic and Mistral document their own SDKs and request formats, though Anthropic also exposes an OpenAI-style endpoint. What slows a launch is everything after the key. OVHcloud caps anonymous calls at 2 requests per minute and authenticated ones at 400 per project and model. Cohere's trial key is free but limited. Bedrock adds an AWS account, IAM roles and a per-Region retention setting before the first call is clean.

The real price: tokens are cheap, the multipliers are not

All twelve bill input and output tokens separately, so the list price per million tokens tells only half the story. The rest sits in modifiers. Anthropic charges 1.1 times the standard rate when inference is pinned to the United States. Google's paid tier adds context caching and a 50 percent batch discount, while its free tier lets Google use submitted content to improve products. OpenRouter adds a platform fee of 5.5 percent on its standard plan and 8 percent on Business. Mistral prices Large at $0.5 per million input tokens and $1.5 output, with batch jobs at half. Run a month of real traffic through each meter before comparing, because caching and batching change rankings.

Getting your data out: what stays on the provider's side after the call

Portability is the easy half here: a prompt is text, and an OpenAI-style request body moves between vendors with little rewriting. The harder half is what the provider keeps. Scaleway, Fireworks and Groq state that inference content is not retained by default, with narrow exceptions: Scaleway may store a failing request for up to two weeks, and Groq may log inputs and outputs for up to 30 days when it detects misuse, with an opt-out. OpenAI keeps abuse-monitoring logs for up to 30 days unless a zero-retention arrangement applies. Files for batch jobs and fine-tuning datasets are the other holdings: Groq keeps batch files 30 days, and fine-tuned weights stay until the customer deletes them.

Independence from the vendor: one model maker or many

Lock-in here is mostly about models, not code. A provider that makes its own models, such as OpenAI, Anthropic or Cohere, can retire a version on its own schedule, and your prompts were tuned to that version. Providers that serve several open-weight families, including Together AI, Fireworks AI, Groq, Scaleway and OVHcloud, let a team move between models without changing the contract. OpenRouter goes furthest, with 500 plus models from more than 80 providers behind one key. Bedrock sits in between, carrying third-party models inside one AWS bill. A thin abstraction in your own code, with the base URL and model name in configuration, is the cheapest insurance.

Who it's for: match the provider to the constraint, not the benchmark

Start from the constraint that cannot move. A team that needs the strongest closed model picks the maker directly, OpenAI, Anthropic or Google, and accepts a single supplier. A team that wants to compare models weekly, or keep a fallback when one provider has an outage, is the audience for OpenRouter. A company bound by European data-protection duties should look at Scaleway, Mistral and OVHcloud first, since each contracts through a French entity. Latency-sensitive features such as voice or autocomplete suit Groq and Fireworks. Companies already standardized on AWS can keep procurement, billing and identity in one place with Bedrock, even if its model catalog moves slower than the makers' own.

Compared at a glance

#ToolBest forPricing modelFree optionHeadquarters
1OpenRouter Teams comparing models or needing failoverUsage-based + platform feeFree planUnited States
2Together AI Developers on open-weight modelsUsage-basedNoneUnited States
3Scaleway Generative APIs EU companies with data-residency dutiesUsage-basedFree planFrance
4OpenAI Platform Teams wanting the vendor's own modelsUsage-basedNoneUnited States
5Anthropic Claude API Teams building on Claude modelsUsage-basedNoneUnited States
6Fireworks AI Latency-sensitive products on open modelsUsage-based (prepaid)NoneUnited States
7Groq Voice and real-time featuresUsage-basedFree planUnited States
8Mistral AI EU teams wanting a European model makerUsage-based + plansFree planFrance
9Google Gemini API Prototypes and Google Cloud shopsUsage-basedFree planUnited States
10OVHcloud AI Endpoints EU teams already on OVHcloudUsage-basedFree planFrance
11Cohere RAG and search over company documentsUsage-basedFree trialCanada
12Amazon Bedrock AWS shops needing one billUsage-basedNoneUnited States

The 12 tools, reviewed

#1 OpenRouter

A single API in front of 500+ models from 80+ providers · United States · openrouter.ai

Many models, one key Opt-in prompt logging EU in-region routing

OpenRouter takes first place because it removes the commitment that hurts most in this category: being tied to one model owner. The pricing page lists more than 500 models behind one key, and the standard plan adds routing preferences, budgets and caching. Its data page says prompts and responses are not stored unless the customer enables logging, and that retention is never on by default. Business and Enterprise customers can send requests to an EU or US host so that processing stays in that region. For a company that expects to swap models every few months, it is the cheapest way to keep that option open.

Where it falls short

The fee on top of model cost grows with spend, 5.5% on standard and 8% on Business. Each request passes through a middleman, so data handling also depends on whichever upstream provider is chosen, and OpenRouter says those providers differ on retention and training. EU in-region routing is not on the free or standard plans, and the free plan is capped at 50 requests a day.

Wrong for

A company that must contract directly with its model maker, or one sending regulated personal data without paying for in-region routing. Scaleway or a direct contract with the model owner is a cleaner fit.

Pricing: Pay as you go with a platform fee of 5.5% on the standard plan and 8% on Business; a free plan offers 25+ free models at 50 requests a day. (free plan)

Visit OpenRouter →

#2 Together AI

Open-weight model hosting with fine-tuning and dedicated capacity · United States · together.ai

Open-weight catalog Zero-retention switch Fine-tuning included

Together sits second because it is the plainest option for a team that wants open-weight models and a clean exit. Its docs describe an API compatible with the OpenAI REST API and SDKs across chat, completions, vision, image generation, speech and embeddings, so switching in takes a changed base URL. The terms put data control in a visible setting: answer No to storing prompts and allowing training and the account runs under zero data retention. The same company also trains and fine-tunes models, which helps a team that starts with prompts and later wants a tuned model.

Where it falls short

The pages we read did not state where Together's inference runs, so a buyer with residency duties must ask. Zero retention is a setting that applies only from the moment it is enabled and does not reach earlier data. The pricing page is a long list of per-model prices in several units, and we found no standing free tier in the pages read.

Wrong for

A European company that needs a stated hosting country, or a non-technical team that wants a chat interface and no code. Scaleway fits the first case; a chat product fits the second.

Pricing: Billed per million tokens with published per-model prices, cached-token discounts on some models, and dedicated capacity sold separately. (none)

Visit Together AI →

#3 Scaleway Generative APIs

Serverless open-weight model endpoints from a Paris data center · France · scaleway.com

Zero retention by default Paris-hosted Free tier

Scaleway is the strongest answer here for a buyer whose first question is where the prompt goes. Its data-privacy page says a zero-retention policy applies by default, that the content of inputs and outputs is not read or reused, and that data is not used to train base models. The FAQ says all serverless models are hosted in a data center in Paris. The contracting entity is a French simplified joint-stock company with a registered office in Paris. The API follows the OpenAI format, so existing client code works, and a free tier lets a team test before paying.

Where it falls short

The catalog is a curated set of open-weight models, so a team that needs the newest closed models from OpenAI or Anthropic will not find them. The FAQ says serverless models currently sit in one Paris site, and for single-region processing the vendor recommends the dedicated option. Failing requests may be stored for up to two weeks for debugging.

Wrong for

A team that must use a particular closed frontier model, or one that wants a gateway across many providers. A direct maker contract or OpenRouter is the better look.

Pricing: Billed per token on serverless endpoints with a free tier; dedicated deployments are available when a team wants isolation. (free plan)

Visit Scaleway Generative APIs →

#4 OpenAI Platform

The maker's own API for GPT models, audio and images · United States · platform.openai.com

Reference API EU data residency Zero-retention eligible

As the vendor whose API format became the standard, OpenAI is a safe default for a team that wants its own models and wide tooling support, but a large incumbent cannot take a top-three place. Its data-controls page says API data is not used to train models unless the customer opts in. Abuse-monitoring logs are kept for up to 30 days unless law requires longer, and a long list of endpoints, including chat completions, responses, embeddings and realtime, qualifies for zero data retention. Storage and processing are available in the US and in Europe (the EEA and Switzerland).

Where it falls short

Everything depends on one maker: model versions retire on its schedule, and a prompt tuned to one model rarely transfers unchanged. Retention defaults to a 30-day log, and zero retention and regional processing are arrangements rather than the starting point. The pages we read did not give a street address for the contracting company.

Wrong for

A company whose legal team requires an EU-only supplier, or one that wants to change models freely. Mistral, Scaleway or a gateway such as OpenRouter fits those cases better.

Pricing: Billed per token with separate input and output rates for each model; enterprise terms add residency and retention options. (none)

Visit OpenAI Platform →

#5 Anthropic Claude API

The maker's own API for Claude models · United States · anthropic.com

Own models US-only inference option Zero-retention arrangement

Anthropic sits at fifth because its terms are short and direct: they say Anthropic may not train models on customer content from the services. The data-residency page lets a request be pinned with inference_geo to the US or left global, and workspaces can restrict which geographies are allowed. A zero data retention arrangement is described in the docs, with a feature table showing what it covers. The Messages API is documented across several languages, and an OpenAI-compatible endpoint helps teams that want to test it from existing code.

Where it falls short

The only geographies documented are global and US, and the only workspace geography is the US, so a buyer who needs processing inside the EU cannot get it from the first-party API on these pages; the docs add that on Bedrock and Google Cloud the region follows the endpoint or profile chosen. Pinning to the US raises the price by 10 percent, and the pinning setting does not apply through the OpenAI-compatibility endpoint.

Wrong for

A European company that must keep inference inside the EU, or a team wanting open-weight models it could later self-host. Mistral or Scaleway are the places to start.

Pricing: Billed per token; US-only inference is priced at 1.1 times the standard rate on Claude 4.6 and later models. (none)

Visit Anthropic Claude API →

#6 Fireworks AI

Fast open-model inference with prepaid credits · United States · fireworks.ai

Zero retention for open models Prepaid credits OpenAI-format calls

Fireworks earns a mid-table place on one strong sentence: its docs say it does not log or store prompt or generation data for any open model without explicit opt-in, and that data exists only in volatile memory for the length of the request. The privacy notice repeats that inputs are not used to train models without opt-in. The base URL follows the OpenAI format. Billing is prepaid: a team buys credits, sets auto reload or a monthly limit, and cannot be surprised by an overage, which suits small companies and side projects with a firm budget.

Where it falls short

The zero-retention default covers open models only; the Responses API stores conversation data when store is true, which is its default. The pages we read did not state where inference runs, and we could not confirm the company's street address from its own pages, so only the city is given. No standing free tier appears on the pricing page.

Wrong for

A buyer who needs a closed frontier model or a stated hosting region in Europe. A maker's own API or Scaleway answers those needs more directly.

Pricing: Prepaid credits deducted per token, with Standard, Priority and Fast serving paths and a monthly spend limit option. (none)

Visit Fireworks AI →

#7 Groq

Very fast inference on custom hardware with a free plan · United States · groq.com

Fast inference Free plan No default retention

Groq's pitch is speed, and the docs back the privacy side: by default it does not retain customer data for inference requests, keeping only usage metadata. It lists text generation, transcription, speech output and image recognition on one platform, with prompt caching and structured outputs. The base URL follows the OpenAI format, which the docs describe as mostly compatible, with a short list of unsupported features. A free plan with published limits lets a team prototype first. For voice agents and autocomplete, where waiting a second is too long, it deserves a test run.

Where it falls short

The model list is smaller than the gateways and cloud platforms, and some OpenAI features are unsupported. Inputs and outputs may be logged for up to 30 days when misuse triggers it, unless the customer opts out. Batch files stay for 30 days and fine-tuning data until deleted. The legal address on its terms is a post office box in California.

Wrong for

A buyer who needs the newest closed models or EU-based contracting. OpenAI or Anthropic suits the first need; Mistral or Scaleway fits the second.

Pricing: A free plan with rate limits and pay-as-you-go billing per token on paid usage, with batch jobs available. (free plan)

Visit Groq →

#8 Mistral AI

French model maker with EU-hosted API and open-weight models · France · mistral.ai

EU-hosted by default Own open-weight models Batch discount

Mistral is a rare case of a model maker and a European contracting party in one. The privacy policy names a French company registered in Paris. Its help center states that data is hosted in the European Union by default, with a US API endpoint for those who choose it. The commercial terms say customer data is not used for training except in listed cases, such as opted-in features, feedback and preview models. The pricing page publishes per-token rates, a 50 percent batch discount and cached-input savings. A free plan and monthly API credits make a pilot inexpensive.

Where it falls short

Abuse monitoring runs on the API unless zero data retention is activated, and the help center says data can be temporarily transferred outside the EU depending on the feature, to the sub-processors in its Trust Center. Preview models can be used for training. It is one family of models, so a buyer wanting several makers behind one key needs a gateway.

Wrong for

A team that wants the broadest model choice behind one key, or one that insists on a specific closed model. OpenRouter or the maker's own API fits better.

Pricing: A free plan, paid plans with monthly API credits, and pay-as-you-go credits beyond them; Mistral Large is listed at $0.5 per million input tokens and $1.5 output. (free plan)

Visit Mistral AI →

#9 Google Gemini API

Google's own Gemini models via a free and paid API · United States · ai.google.dev

Free tier Context caching Batch discount

Google's API is the easiest place to start for free: the pricing page lists a free tier with free input and output tokens, and a paid tier with higher limits, context caching and batch pricing at half cost. The terms draw a hard line between tiers: on paid services Google says it does not use prompts or responses to improve its products, while the unpaid tier does. An OpenAI-compatible endpoint exists and the docs label it beta. Billing for paid services runs through Google Cloud, a plus for teams already there.

Where it falls short

On the free tier, submitted content is used to improve Google products, and the terms limit use for audiences in the EEA, Switzerland and the UK to paid services only. Paid services log prompts for a limited period to detect abuse, and grounding features keep data 30 days. We found no EU-only processing option on the pages read.

Wrong for

A European company sending personal data on a free key, or one needing a vendor other than a big platform. Move to paid terms, or look at Mistral or Scaleway.

Pricing: A free tier of limited models, and a paid tier billed per million tokens with context caching and a 50 percent batch discount. (free plan)

Visit Google Gemini API →

#10 OVHcloud AI Endpoints

Serverless open-model API from a French cloud provider · France · ovhcloud.com

French cloud No stored user data OpenAI specification

OVHcloud suits a company that already runs infrastructure with it and wants models inside the same account. The docs describe AI Endpoints as serverless and say user data is not stored, and that its LLM APIs follow the OpenAI specification, so existing client code works. The catalog lists open models such as gpt-oss, Llama and Qwen alongside embedding, speech and image models. The company is a French simplified joint-stock company registered in Roubaix. Anonymous requests work without an account, which allows a quick test before a Public Cloud project is created.

Where it falls short

Authenticated limits are 400 requests per minute per project and model, and keys from projects in Discovery mode without a payment method cannot use the service. The playground caps output at 1,024 tokens. The pages we read did not state the processing location of each model, and the model list is smaller than a gateway's.

Wrong for

A team that needs a wide closed-model catalog or large rate limits on day one. A gateway or the model makers fit better.

Pricing: Billed on a Public Cloud project with a payment method; anonymous use is limited to 2 requests per minute per IP and model. (free plan)

Visit OVHcloud AI Endpoints →

#11 Cohere

Enterprise language, embedding and rerank models · Canada · cohere.com

Embeddings and rerank Trial key Canadian vendor

Cohere is the retrieval specialist of the list: chat models alongside embeddings, which fits a company building search over its own documents. A trial key is free but limited, and production keys bill per token, with input and output priced apart. A compatibility API at api.cohere.ai/compatibility/v1 covers chat completions, embeddings and audio transcription for the OpenAI client libraries. The vendor is a Toronto company and its terms are governed by Ontario law. Enterprise customers can negotiate a data processing addendum and control how data is used for training.

Where it falls short

The privacy policy says inputs and outputs on the platform are generally retained 30 days for enterprise users, and that trial and research inputs may be used for improvement after de-identification. We found no zero-retention option in that policy. The compatibility API lists three endpoint groups only, and the model range is narrower than the gateways.

Wrong for

A team that needs zero retention without an enterprise contract, or one that wants the widest open-model choice. Scaleway, Fireworks or OpenRouter answer better.

Pricing: A free trial key with limits, and production keys billed per token with different input and output rates. (free trial)

Visit Cohere →

#12 Amazon Bedrock

Managed access to many makers' models inside AWS · United States · aws.amazon.com

AWS account Per-Region retention Third-party models

Bedrock suits a company whose procurement, identity and logging already live in AWS. Its documentation says model providers have no access to the accounts where models run, and therefore not to customer prompts or completions. The retention page lists modes from none, which writes nothing to durable storage, to aws_review, and lets an organization enforce a mode with IAM or service control policies. The same account can reach several makers through one bill. For a security team that wants guardrails in policy code rather than a vendor promise, that control is its main strength.

Where it falls short

Some models require retention for human review within the AWS boundary, up to 30 days, and zero retention for them needs approval through the AWS account manager. The setting is per Region and does not carry over. Setup takes an AWS account, roles and Region choices, which is heavy for a small team that wants a key and a curl command.

Wrong for

A small team without AWS experience, or one that wants the maker's newest model on release day. A direct maker API or OpenRouter is simpler.

Pricing: Billed through AWS per token by model and Region; the retention mode is set per Region at account or project level. (none)

Visit Amazon Bedrock →

What the data says about this market

Of the twelve providers ranked, eight are headquartered in the United States (OpenRouter, Together AI, Fireworks AI, Groq, OpenAI, Anthropic, Google and Amazon), three in France (Scaleway, OVHcloud and Mistral AI) and one in Canada (Cohere). No provider here is headquartered in Germany, the Netherlands or the Nordics, which means a buyer in those countries who wants a domestic supplier will not find one in this list. All twelve meter usage per token, and the vendors differ mainly in what they add on top: caching discounts, batch discounts, regional surcharges or a gateway fee.

Seven of the twelve let a developer make calls without a payment method or with a standing free allowance: OpenRouter (25 or more free models at 50 requests a day), Groq, Google, Mistral, Scaleway, OVHcloud (anonymous calls at a low rate) and Cohere (a trial key). The free tiers are not equal. Google states that on its free tier submitted content is used to improve its products, and that only paid services may be used when serving users in the European Economic Area, Switzerland or the United Kingdom. A prototype built on a free key can therefore be unusable in production for a European audience until the account is moved to a paid tier.

Where requests run is the clearest split. Scaleway serves its serverless models from a data center in Paris, Mistral hosts in the European Union by default and offers a separate US endpoint, OpenAI offers storage and processing in Europe for eligible customers, and OpenRouter offers EU in-region routing on its Business and Enterprise plans. Anthropic's documented geography options are global or US-only, and its workspace geography is US only. The category also sits inside a wider shift: ICT services rose from 9.1% of world service exports in 2013 to 14.48% in 2023 (World Bank), and a metered model call is a small unit of that trade.

The 12 ranked vendors, counted

  • Headquarters by region: North America 9, Europe 3
  • By country: United States 8, France 3, Canada 1
  • Pricing model: Usage-based 9, Usage-based (prepaid) 1, Usage-based + plans 1, Usage-based + platform fee 1
  • Free option: Free plan 6, None 5, Free trial 1

Counted from the 12 vendors on this page. More in our market data.

For the wider market behind LLM API providers, read our report The Global Shift to ICT Services, or browse all industry reports.

How to choose

  1. Replay real traffic against two or three meters

    Export a week of actual prompts, or build a sample of one hundred from your product, and note average input and output length, how often the same prefix repeats, and which calls could wait for a batch run. Then price that profile on each candidate. Prefix repetition favors providers with cached-token discounts, such as Google and Mistral; batch-friendly jobs favor those with half-price batch pricing. Add any gateway fee, any regional surcharge such as Anthropic's 1.1 times for US-only inference, and the cost of retries. The cheapest list price often loses once those lines are in, and the spreadsheet gives you the number to negotiate from.

  2. Write down where each prompt may be processed and kept

    List the kinds of text your product will send: customer names, contract clauses, support tickets, source code. For each, decide whether it may leave the European Economic Area, whether a 30-day abuse log is acceptable, and whether a human at the provider may ever read it. Then check each shortlisted provider's privacy or data-handling page against that list, and ask for a signed data processing agreement. Providers that state zero retention by default, such as Scaleway and Fireworks, answer quickly; those that route to third-party model owners need a closer look at the sub-processor list.

  3. Keep a second provider warm from the first week

    Put the base URL, API key and model name in configuration, not in code, and run a small share of traffic through a second provider from the start. Outages, deprecations and sudden price changes happen to every vendor here, and a team that has never tested its fallback finds out the problem at the worst time. A gateway such as OpenRouter makes this a configuration change; with direct providers it means maintaining two integrations. Re-run your own evaluation prompts each quarter, because the cheapest acceptable model changes faster than any ranking.

Questions and answers

What is the best LLM API provider in 2026?

OpenRouter leads this ranking because one key reaches more than 500 models from over 80 providers, prompts are not stored unless the customer opts in, and a failed provider can be routed around. Together AI is the better pick for a team that wants open-weight models and a documented zero-retention setting. Scaleway is the better pick when the prompt must stay in the EU, since it applies zero retention by default and hosts serverless models in Paris. A team committed to one maker's models should contract with that maker directly.

Is there a European alternative to the OpenAI API?

Yes, three in this list contract through French companies: Scaleway, Mistral AI and OVHcloud. Scaleway hosts its serverless models in a Paris data center and applies zero retention by default. Mistral's help center says data is hosted in the EU by default, with a US endpoint as an option. OVHcloud says its AI Endpoints do not store user data. All three accept requests close to the OpenAI format, though Mistral documents its own SDK. None offers OpenAI's own models, so prompts may need retesting.

Can I switch providers without rewriting my code?

Often, for the basic chat call. OpenRouter, Together AI, Fireworks AI, Groq, Scaleway, OVHcloud, Cohere and Google all document an OpenAI-style endpoint where you change the base URL, key and model name. Groq calls its version mostly compatible and Google labels its beta. Features beyond plain chat, such as tool use, structured outputs, batch jobs and caching, differ in names and limits, so run your own test set against each. Keep model names in configuration to make the switch a small change.

Do these providers train on my prompts?

The pages we read say no by default for most, with exceptions worth knowing. OpenAI says API data is not used for training unless you opt in. Anthropic's terms say it may not train on customer content from the services. Scaleway, Fireworks and Groq describe no retention of inference content by default. Google's free tier does use submitted content to improve its products, and Cohere says trial and research inputs may be used after de-identification. Mistral excludes preview models and feedback from its no-training promise.

How long do providers keep my requests?

It ranges from nothing to 30 days. Scaleway and Groq keep no inference content by default, with narrow misuse exceptions of up to two weeks and 30 days. OpenAI keeps abuse-monitoring logs up to 30 days unless a zero-retention arrangement applies. Cohere says enterprise inputs and outputs are generally retained 30 days. Bedrock lets the customer set a retention mode per Region, though models that require human review keep content up to 30 days. Ask for the retention clause in writing and check which endpoints it covers.

How is an LLM API priced?

By the token, with input and output counted and priced separately. Providers then add modifiers: Google and Mistral discount cached input and batch jobs by half, Anthropic charges 1.1 times the standard rate for US-only inference, and OpenRouter adds a fee of 5.5 or 8 percent on top of model cost. Fireworks uses prepaid credits. Estimate from your own traffic: average prompt length, answer length, how often a prefix repeats and how many calls can wait for a batch. List prices alone mislead.

Which providers have a free tier?

OpenRouter offers more than 25 free models at 50 requests a day. Groq, Google, Mistral and Scaleway each describe a free plan or free tier, Cohere gives a limited trial key, and OVHcloud allows anonymous calls at 2 requests per minute. Free tiers carry data terms: Google uses free-tier content to improve its products and limits audiences in the EEA, Switzerland and the UK to paid services. Use a free key for prototypes and move to paid terms before real customer text goes through it.

When does a gateway like OpenRouter make sense?

A gateway fits when you compare models often, want an automatic fallback during a provider outage, or prefer one invoice for many vendors. The cost is a platform fee of 5.5 percent on the standard plan and 8 percent on Business, plus an extra hop in the data path. OpenRouter says retention depends on each upstream provider and gives controls to avoid providers that train. For a single stable model and large volume, a direct contract is cheaper.

Can I keep processing inside the EU?

Some can. Scaleway serves serverless models from Paris, Mistral hosts in the EU by default, OpenAI offers European storage and processing to eligible customers, and OpenRouter offers EU in-region routing on Business and Enterprise plans. Anthropic documents only global and US-only inference, with US the sole workspace geography. Several vendors did not state a processing location on the pages we read, including Together AI, Fireworks and OVHcloud for individual models. Get the location in the contract, not a marketing page.

Should I use the model maker or a cloud platform like Bedrock?

The maker gets new models first and sells the clearest terms for its own models. A cloud platform makes sense when your company already buys, authenticates and logs everything in that cloud. Bedrock keeps billing and IAM policy in AWS and says model providers cannot see prompts or completions, but some models require retained data for review, and setup assumes AWS skills. A small team without that background will ship sooner with a maker's API or a gateway.

What breaks when a provider retires a model?

The model name in your code starts returning errors, or a replacement answers differently, so prompts tuned for the old model degrade. Makers such as OpenAI, Anthropic and Cohere retire their own versions on their own schedules. Providers that host open-weight models can keep older families longer but may drop them too. Defend against it by keeping a fixed evaluation set of prompts, re-running it each quarter, and holding a second provider ready so a retirement notice becomes a planned migration.