Lower your inference bill.
Compare published rates for the models you already use, including input, output, and available cache pricing.
Use leading AI models through one OpenAI-compatible API.
Keep your existing tools and spend less on inference.
PRICING · PER 1M TOKENS
Search the catalogue, compare published rates, and explore cache pricing for your workload.
| Model | Input | Cached input | Output | Savings | Details |
|---|---|---|---|---|---|
Claude Fable 5Anthropic | $10.00$2.0080% less than OpenRouter | $1.00$0.2080% less than OpenRouter | $50.00$10.0080% less than OpenRouter | 80% less | |
Claude Haiku 4.5Anthropic | $1.10$0.2280% less than OpenRouter | $0.11$0.02280% less than OpenRouter | $5.50$1.1080% less than OpenRouter | 80% less | |
Claude Opus 4.6Anthropic | $5.00$1.0080% less than OpenRouter | $0.50$0.1080% less than OpenRouter | $25.00$5.0080% less than OpenRouter | 80% less | |
Claude Opus 4.7Anthropic | $5.00$1.0080% less than OpenRouter | $0.50$0.1080% less than OpenRouter | $25.00$5.0080% less than OpenRouter | 80% less | |
Claude Opus 4.8Anthropic | $5.00$1.0080% less than OpenRouter | $0.50$0.1080% less than OpenRouter | $25.00$5.0080% less than OpenRouter | 80% less | |
Claude Opus 5Anthropic | $5.00$1.0080% less than OpenRouter | $0.50$0.1080% less than OpenRouter | $25.00$5.0080% less than OpenRouter | 80% less | |
Claude Sonnet 4.6Anthropic | $3.00$0.6080% less than OpenRouter | $0.30$0.0680% less than OpenRouter | $15.00$3.0080% less than OpenRouter | 80% less | |
Claude Sonnet 5Anthropic | $2.00$0.4080% less than OpenRouter | $0.20$0.0480% less than OpenRouter | $10.00$2.0080% less than OpenRouter | 80% less | |
DeepSeek V4 FlashDeepSeek | $0.14$0.02880% less than OpenRouter | $0.03$0.00680% less than OpenRouter | $0.28$0.05680% less than OpenRouter | 80% less | |
DeepSeek V4 ProDeepSeek | $1.74$0.34880% less than OpenRouter | $0.14$0.02880% less than OpenRouter | $3.48$0.69680% less than OpenRouter | 80% less | |
Gemini 2.5 FlashGoogle | $0.30$0.0680% less than OpenRouter | $0.03$0.00680% less than OpenRouter | $2.50$0.5080% less than OpenRouter | 80% less | |
Gemini 2.5 ProGoogle | $1.25$0.2580% less than OpenRouter | $0.12$0.02480% less than OpenRouter | $10.00$2.0080% less than OpenRouter | 80% less |
USD per million tokens. Crossed-out rates are OpenRouter list prices. Savings show the lower of input and output savings; individual rates can differ. Expand a row for all published rates.
Compare published rates for the models you already use, including input, output, and available cache pricing.
Use your OpenAI-compatible client with a new base URL and API key. Choose a model for each request.
Tell us which models you use and how much traffic you expect, so we can plan capacity for your application.
A FAMILIAR STARTING POINT
After onboarding, update your base URL and API key in your OpenAI-compatible client. Check the model and features your application uses before switching traffic.
Read the integration docsfrom openai import OpenAI
import os
client = OpenAI(
base_url="https://api.ferrix.sh/v1",
api_key=os.environ["FERRIX_API_KEY"],
)Before you move a workload, know how pricing works, what happens to your data, and what your team needs.
Trust
Public rates, usage records, and practical checks before you move production traffic.
See what you can verifyPrivacy
Unbar never stores your API prompts or completions. Usage and billing records stay separate from conversation content.
Read the Privacy PolicyFor teams
Tell us your model needs and billing requirements. We confirm the options we can support before onboarding.
Explore team optionsLet’s talk about your AI spend
Tell us which models you use and roughly what you spend each month. An estimate is enough.
We’ll review your request and email you when we can onboard your workload.