AI Gateway
Use models through one API. Compare published rates by model, provider, and billing unit.
Explore model ratesPricing
Explore current model rates, estimate your API usage, and keep agent hosting costs in view.
Understand your costs
Model usage
Input, output, or media units at the published model rate.
Agent hosting
DOSClaw runtime and plan charges, separate from model usage.
Connected services
External providers may charge for their own products and APIs.
One place to understand AI Gateway usage, DOSClaw hosting, and business requirements.
Use models through one API. Compare published rates by model, provider, and billing unit.
Explore model ratesRun agents with knowledge, channels, and integrations. Choose hosting and review the charge before creation.
Understand agent hostingDiscuss workload, rollout, support, and procurement requirements with our team.
Contact salesAI Gateway
Rates come from the same live model catalog used by the app. Search and compare before you integrate.
Current catalog rates
All amounts are in USD per displayed unit. Input / generation is the catalog's base rate; text output is listed separately.
| Model | Billing unit | Input / generation | Text output |
|---|---|---|---|
| AutoDOS.AI | 1M tokens | Selected model rate | Selected model rate |
| DOSDOS.AI | 1M tokens | $0.07 | $0.50 |
| Qwen3.8-27BDOS.AI | 1M tokens | $0.1125 | $1.50 |
| Granite 4.0 H MicroCloudflare | 1M tokens | $0.017 | $0.112 |
| Llama 3.2 1B InstructCloudflare | 1M tokens | $0.027 | $0.201 |
| GPT-OSS 120Bdeepinfra | 1M tokens | $0.037 | $0.17 |
| GPT-OSS 20Bdeepinfra | 1M tokens | $0.04 | $0.15 |
| Llama 3.2 3B InstructCloudflare | 1M tokens | $0.0509 | $0.335 |
| Qwen3 30B A3B FP8Cloudflare | 1M tokens | $0.0509 | $0.335 |
| GLM 4.7 FlashCloudflare | 1M tokens | $0.0605 | $0.40 |
| Gemma 4 26B A4BCloudflare | 1M tokens | $0.10 | $0.30 |
| Tencent HY3deepinfra | 1M tokens | $0.14 | $0.58 |
| Xiaomi MiMo V2.5deepinfra | 1M tokens | $0.14 | $0.28 |
| GLM 5.3 FlashCloudflare | 1M tokens | $0.15 | $0.50 |
| Qwen3.8 FlashQwen | 1M tokens | $0.15 | $0.47 |
| Llama 3.1 8BCloudflare | 1M tokens | $0.152 | $0.287 |
| GPT-5.4 NanoOpenAI | 1M tokens | $0.21 | $1.31 |
| GPT-5.6 LunaOpenAI | 1M tokens | $0.21 | $1.26 |
| Gemini 3.1 Flash-LiteGoogle | 1M tokens | $0.26 | $1.58 |
| Llama 3.3 70BCloudflare | 1M tokens | $0.293 | $2.253 |
| MiniMax M2.7MiniMax | 1M tokens | $0.30 | $1.20 |
| DeepSeek V4.1 FlashAlibaba Cloud | 1M tokens | $0.315 | $1.26 |
| Gemini 3.5 Flash-LiteGoogle | 1M tokens | $0.32 | $2.63 |
| Gemma SEA-LION v4 27BCloudflare | 1M tokens | $0.351 | $0.555 |
| Mistral Small 3.1 24BCloudflare | 1M tokens | $0.351 | $0.555 |
| Qwen 3.7 PlusQwen | 1M tokens | $0.42 | $1.68 |
| DeepSeek R1 Distill Qwen 32BCloudflare | 1M tokens | $0.497 | $4.881 |
| Nemotron 3 120B A12BCloudflare | 1M tokens | $0.50 | $1.50 |
| NVIDIA Nemotron 3 Ultradeepinfra | 1M tokens | $0.50 | $2.20 |
| Qwen2.5 Coder 32BCloudflare | 1M tokens | $0.66 | $1.00 |
| QwQ 32BCloudflare | 1M tokens | $0.66 | $1.00 |
| Gemini 3.1 Flash LiveGoogle | 1M tokens | $0.79 | $4.73 |
| Gemini 3.7 FlashGoogle | 1M tokens | $0.79 | $3.94 |
| Gemini 3.8 FlashGoogle | 1M tokens | $0.79 | $3.94 |
| GPT-5.4 MiniOpenAI | 1M tokens | $0.79 | $4.73 |
| Kimi K2.6Cloudflare | 1M tokens | $0.95 | $4.00 |
| Kimi K2.7 CodeTogether AI | 1M tokens | $0.95 | $4.00 |
| Claude Haiku 4.5Anthropic | 1M tokens | $1.05 | $5.25 |
| Grok 4.3xAI | 1M tokens | $1.31 | $2.63 |
| DeepSeek V4 ProDeepSeek | 1M tokens | $1.32 | $3.96 |
| GLM-5.2Together AI | 1M tokens | $1.40 | $4.40 |
| GLM 5.3Cloudflare | 1M tokens | $1.40 | $4.40 |
| Gemini 3.5 FlashGoogle | 1M tokens | $1.58 | $9.45 |
| Qwen 3.8 MaxQwen | 1M tokens | $1.87 | $5.62 |
| Gemini 3.1 ProGoogle | 1M tokens | $2.10 | $12.60 |
| GPT-5.6 TerraOpenAI | 1M tokens | $2.10 | $12.60 |
| Grok 4.20xAI | 1M tokens | $2.10 | $6.30 |
| Qwen 3.7 MaxQwen | 1M tokens | $2.63 | $7.88 |
| Kimi K3Alibaba Cloud | 1M tokens | $3.00 | $15.00 |
| Claude Sonnet 5Anthropic | 1M tokens | $3.15 | $15.75 |
| Claude Opus 5Anthropic | 1M tokens | $5.25 | $26.25 |
| GPT-5.6 SolOpenAI | 1M tokens | $5.25 | $31.50 |
| Nano Banana 2 LiteGoogle | Unavailable | Unavailable | Not applicable |
| Nano Banana 2Google | Unavailable | Unavailable | Not applicable |
| Nano Banana ProGoogle | Unavailable | Unavailable | Not applicable |
| Wan 2.2 T2I FlashAlibaba Cloud | Unavailable | Unavailable | Not applicable |
| Wan 2.7 Image-to-VideoAlibaba Cloud | 1K seconds | $100.00 | Not applicable |
| Wan 2.7 Text-to-VideoAlibaba Cloud | 1K seconds | $100.00 | Not applicable |
| Qwen3 Embedding 4BDOS.AI | 1M tokens | $0.02 | Not applicable |
| Qwen3 Embedding 0.6BCloudflare | 1M tokens | $0.0118 | Not applicable |
| Gemini Embedding 001Google | 1M tokens | $0.02 | Not applicable |
| BGE Small EN v1.5Cloudflare | 1M tokens | $0.0202 | Not applicable |
| BGE Base EN v1.5Cloudflare | 1M tokens | $0.0666 | Not applicable |
| BGE Large EN v1.5Cloudflare | 1M tokens | $0.204 | Not applicable |
The table shows each model's published base rate. Provider-specific offerings and additional billing details are available on the model page. Rates refresh every five minutes while this page is active.
Choose a text model and enter your expected input and output tokens.
An estimate using current base rates, not a quote. Excludes agent hosting, provider-specific pricing, caching adjustments, and other service fees.
DOSClaw
Choose your DOSClaw setup first, then budget for the models and connected services your agent uses.
Hosting charges depend on the instance and your account's plan. Review current prices and included entitlements in the console before confirming a purchase.
View current plans in the appModel requests use the applicable model rate and account entitlements. External integrations can have their own subscriptions or API fees. Hosting does not make every model or connected service free.
Explore DOSClawDOSafe
Entity and URL checks, AI-content detection, and biometric verification. Paid partner keys are charged per successful call after their daily allowance is exceeded.
Service rates are temporarily unavailable. Refresh to try again.
Entity and URL checks, reports, and webhooks.
Unavailable
USD per successful over-limit call
Text, image, video, plagiarism, and document analysis.
Unavailable
USD per successful over-limit call
Voice, face, and multimodal identity verification.
Unavailable
USD per successful over-limit call
These rates are loaded from the service pricing API. Allowances, scopes, and endpoint availability depend on your key and service configuration. Refreshes every five minutes while this page is active.
Manage API keysUsage is calculated from the model's billing unit and applicable input and output rates. Add credits and review recorded usage and charges in your billing dashboard.
The auto model uses the rate of the model selected for that request. It does not have one fixed input and output price. Review the model catalog and request usage for details.
Cache support and charges depend on the provider and model. The table lists base model rates; cache reads, writes, or storage can follow separate billing rules. Review the model and usage details.
Your provider bills requests using your key. DOS currently charges a 4.5% platform fee on the equivalent usage cost at the model's retail rates, separate from your provider's charges. Review BYOK terms and recorded usage in the app.
No. Hosting and model usage are different costs. Your plan may include specific entitlements; check your current plan and selected model before deploying.
The billing dashboard shows available payment methods, credits, and usage. Contact sales to discuss your organization's billing requirements.
No. Eligible paid partner keys use per-call overage rates after their daily allowance. Only successful over-limit calls are charged; your key's limits and enabled scopes still apply.
Start with the models you need. Review your account's billing options before adding funds or deploying agents.