Compare AI model prices for the work you actually do
The lowest input price can produce the bigger bill. Two worked examples show what to compare instead.
A model costs $0.20 per million input tokens. Another costs $0.40. The first looks cheaper—until you ask both to write a long answer.
Compare the cost of your whole request: the text you send, the answer you expect, and how often you run it. Then check that the provider offering that price supports the features you need.
Start with one real task
Pick something you will actually run: summarise a document, rewrite a post, or answer a support question. Count the instructions, conversation history and attached text as input, not just the final question. Tokens are pieces of text; they are not a fixed number of words across languages or models.
Open the AI model catalog, choose a model, then open Cost estimate. Start with a task preset, or enter your own input tokens, output tokens and request count. Presets are editable examples, not measurements of your documents.
If you already have API usage logs, use their token counts. Otherwise, start with an estimate and replace it after a small test. A context limit tells you how much can fit, not how much every request will cost.
Why the cheaper input price can lose
These are two fictional offers, in USD per million tokens. They illustrate the calculation; they are not current prices for named models.
| Offer | Input price | Output price |
|---|---|---|
| A | $0.20 | $1.00 |
| B | $0.40 | $0.40 |
For 10,000 requests with 8,000 input tokens and 500 output tokens each, A costs $21, while B costs $34.
Change the task to 1,000 input tokens and 4,000 output tokens each, keeping 10,000 requests. A now costs $42, while B costs $20. The longer answers reverse the ranking.
total = requests × ((input tokens × input price) + (output tokens × output price)) / 1,000,000
We checked both examples on 10 October 2026 with the same token-cost function used by the catalog. Download the inputs and results. This checks the arithmetic, not model quality or a provider's actual invoice.
The simple formula assumes ordinary text pricing without discounts or tiers. Cache rates and long-context tiers can change the result. Tool calls, images, audio, per-request fees and taxes may add charges beyond the text estimate. For a real run, compare the estimate with the provider's usage record; OpenRouter documents the usage fields it returns.
Compare the same provider offer
The company that develops a model and the service that hosts its API are different choices. In the catalog, use Model creators to narrow the model family and API hosts to choose where to use it. Both allow multiple selections. The Official API label identifies a developer's direct API.
Open a model's provider list before choosing. Check its price, context and output limits together. A low price from one host and a large context window from another do not make one usable offer.
Copy the model ID from the provider you plan to call. Similar model names do not guarantee identical API identifiers. With an aggregator, also check routing: OpenRouter can route the same model to different endpoints, and the selected route affects the available features and price.
Use filters to make a shortlist
Start with requirements that would rule a model out: enough context for the input and answer, support for images if needed, and tool calling or structured output for your application. Then narrow the price range and compare a few candidates.
Speed has two parts. Time to first token describes the wait before an answer starts. Output tokens per second describes how quickly it continues. A missing measurement means unknown, not slow; a provider observation is not a promise about your connection or prompt.
Benchmark scores can help you choose what to test. They cannot tell you whether a model will follow your particular instructions. Run the same small set of representative prompts against your shortlist and check the answers, failures and actual token usage.
The catalog aims to refresh about every six hours. Check the source timestamps, especially when a price or speed looks unusually good. Filters and sorting stay in the URL, so you can copy the address to share the same view.
Repeat the comparison from an app or agent
The catalog API supports filtered searches, provider comparisons and token-cost estimates. Catalog reads use no CozyToolkit credits and work without an API key, within the published rate limit. They do not run the models for you.
The model research skill can do the lookup for an agent. Give it the workload, rather than asking for a universal “best model”: find three candidates for 10,000 document-summary requests, each with 8,000 input and 500 output tokens, and include provider IDs, estimated costs and sources.