Models (64)

All Models

DOS.AI
LLM

Auto

DOS.AI Smart Router classifies each request by complexity. Simple and Medium requests use DOS.AI; Complex requests can use an eligible paid route configured in the live catalog.

Smart Router128K context
autodynamic pricing
DOS.AI
LLMVisionCode

DOS

DOS curated meta-model for DOSClaw and first-party DOS products.

Curated models262K context
$0.07 / $0.50Per 1M tokens
DOS.AI
LLMVisionCode

Qwen3.8-27B

DOS.AI hosted Qwen3.8-27B for reasoning, coding, tool use, and multilingual text generation.

27B Dense262K context
$0.1125 / $1.50Per 1M tokens
Cloudflare
LLM

Granite 4.0 H Micro

131K context
$0.017 / $0.112Per 1M tokens
Cloudflare
LLM

Llama 3.2 1B Instruct

60K context
$0.027 / $0.201Per 1M tokens
deepinfra
LLMCode

GPT-OSS 120B

OpenAI open-weight 120B, strong general reasoning at a fraction of hosted GPT pricing.

128K context
$0.037 / $0.17Per 1M tokens
deepinfra
LLMCode

GPT-OSS 20B

OpenAI open-weight 20B, the cheapest capable chat model in the catalog.

128K context
$0.04 / $0.15Per 1M tokens
Cloudflare
LLM

Llama 3.2 3B Instruct

80K context
$0.0509 / $0.335Per 1M tokens
Cloudflare
LLM

Qwen3 30B A3B FP8

33K context
$0.0509 / $0.335Per 1M tokens
Cloudflare
LLM

GLM 4.7 Flash

128K context
$0.0605 / $0.40Per 1M tokens
Cloudflare
LLM

Gemma 4 26B A4B

256K context
$0.10 / $0.30Per 1M tokens
deepinfra
LLM

Tencent HY3

262K context
$0.14 / $0.58Per 1M tokens
deepinfra
LLM

Xiaomi MiMo V2.5

262K context
$0.14 / $0.28Per 1M tokens
Cloudflare
LLM

GLM 5.3 Flash

1M context
$0.15 / $0.50Per 1M tokens
Qwen
LLM

Qwen3.8 Flash

1M context
$0.15 / $0.47Per 1M tokens
Cloudflare
LLM

Llama 3.1 8B

32K context
$0.152 / $0.287Per 1M tokens
OpenAI
LLMVision

GPT-5.4 Nano

Cheapest GPT-5.4-class model for simple high-volume tasks

Nano400K context
$0.21 / $1.31Per 1M tokens
OpenAI
LLMVisionCode

GPT-5.6 Luna

OpenAI most economical GPT-5.6 model, built for efficient high-volume workloads, with vision and a 1.05M context window.

Efficient1M context
$0.21 / $1.26Per 1M tokens
Google
LLMVision

Gemini 3.1 Flash-Lite

Fastest and most cost-efficient Gemini 3 model

Flash-Lite1M context
$0.26 / $1.58Per 1M tokens
Cloudflare
LLM

Llama 3.3 70B

24K context
$0.293 / $2.253Per 1M tokens
MiniMax
LLM

MiniMax M2.7

205K context
$0.30 / $1.20Per 1M tokens
Alibaba CloudNew
LLM

DeepSeek V4.1 Flash

MoE (flash tier)1M context
$0.315 / $1.26Per 1M tokens
Google
LLMVision

Gemini 3.5 Flash-Lite

Fast, cost-effective Flash-Lite (GA). 1M context. Closest Gemini peer to dos-ai (Qwen3.6-35B-A3B). [PROMO] Free when used BY a DOSClaw agent on a paid plan (agent traffic only - direct API usage is billed normally). Limited-time.

Flash-Lite1M context
$0.32 / $2.63Per 1M tokens
Cloudflare
LLM

Gemma SEA-LION v4 27B

128K context
$0.351 / $0.555Per 1M tokens
Cloudflare
LLM

Mistral Small 3.1 24B

128K context
$0.351 / $0.555Per 1M tokens
Qwen
LLMVisionCode

Qwen 3.7 Plus

Qwen 3.7 Plus - multimodal (text+image), 1M context

Plus1M context
$0.42 / $1.68Per 1M tokens
Cloudflare
LLM

DeepSeek R1 Distill Qwen 32B

80K context
$0.497 / $4.881Per 1M tokens
Cloudflare
LLM

Nemotron 3 120B A12B

256K context
$0.50 / $1.50Per 1M tokens
deepinfra
LLM

NVIDIA Nemotron 3 Ultra

262K context
$0.50 / $2.20Per 1M tokens
Cloudflare
LLM

Qwen2.5 Coder 32B

33K context
$0.66 / $1.00Per 1M tokens
Cloudflare
LLM

QwQ 32B

24K context
$0.66 / $1.00Per 1M tokens
Google
LLMVision

Gemini 3.1 Flash Live

Real-time voice and dialogue model

Flash Live1M context
$0.79 / $4.73Per 1M tokens
Google
LLMVisionCode

Gemini 3.7 Flash

Google hybrid reasoning model combining fast multimodal inference with controllable thinking capabilities, a 1M context window, and up to 64K output tokens.

Flash (Hybrid)1M context
$0.79 / $3.94Per 1M tokens
GoogleNew
LLMVisionCode

Gemini 3.8 Flash

Google's newest Gemini Flash model, with a 1M token context window.

Flash (Hybrid)1M context
$0.79 / $3.94Per 1M tokens
OpenAI
LLMVision

GPT-5.4 Mini

Strong mini model for coding, computer use, and sub-agents

Mini400K context
$0.79 / $4.73Per 1M tokens
Cloudflare
LLM

Kimi K2.6

262K context
$0.95 / $4.00Per 1M tokens
Together AI
LLMCode

Kimi K2.7 Code

Moonshot Kimi K2.7 Code, coding-focused with a 256K context window.

262K context
$0.95 / $4.00Per 1M tokens
Anthropic
LLMVision

Claude Haiku 4.5

Fastest and most compact Claude model

Haiku200K context
$1.05 / $5.25Per 1M tokens
xAI
LLMVisionCode

Grok 4.3

xAI flagship reasoning model - 1M context, text+image (replaces grok-4.1-fast)

Flagship1M context
$1.31 / $2.63Per 1M tokens
DeepSeek
LLMCode

DeepSeek V4 Pro

Near-frontier quality at ~1/6 the cost of Opus 4.7 / GPT-5.5

MoE (pro tier)1M context
$1.32 / $3.96Per 1M tokens
Together AI
LLMCode

GLM-5.2

Zhipu GLM-5.2 with a 1M context window, strong Chinese and English coding.

1M context
$1.40 / $4.40Per 1M tokens
Cloudflare
LLM

GLM 5.3

1M context
$1.40 / $4.40Per 1M tokens
Google
LLMVision

Gemini 3.5 Flash

Latest Gemini Flash - frontier performance, standard tier

Flash1M context
$1.58 / $9.45Per 1M tokens
Qwen
LLMVisionCode

Qwen 3.8 Max

Qwen 3.8 Max - Alibaba's 2.4T-parameter MoE flagship, 1M context, multimodal input (text+image+video)

1M context
$1.87 / $5.62Per 1M tokens
Google
LLMVisionCode

Gemini 3.1 Pro

Google's most advanced reasoning model for complex tasks

Pro1M context
$2.10 / $12.60Per 1M tokens
OpenAI
LLMVisionCode

GPT-5.6 Terra

OpenAI mid tier of the GPT-5.6 family, balancing intelligence and cost with vision, tool calling, and a 1.05M context window.

Balanced1M context
$2.10 / $12.60Per 1M tokens
xAI
LLMVision

Grok 4.20

xAI flagship reasoning model with 2M context

Flagship2M context
$2.10 / $6.30Per 1M tokens
Qwen
LLMCode

Qwen 3.7 Max

Qwen 3.7 Max tier - 1M context (standard price; OpenRouter shows a 50%-off promo)

Max1M context
$2.63 / $7.88Per 1M tokens
Alibaba Cloud
LLM

Kimi K3

1M context
$3.00 / $15.00Per 1M tokens
Anthropic
LLMVisionCode

Claude Sonnet 5

Anthropic balanced model with near-flagship coding and agentic quality at Sonnet pricing, a 1M context window, and up to 128K output tokens.

Sonnet1M context
$3.15 / $15.75Per 1M tokens
Anthropic
LLMVisionCode

Claude Opus 5

Anthropic flagship model for complex agentic coding and enterprise work, with a 1M context window and up to 128K synchronous output tokens.

Opus1M context
$5.25 / $26.25Per 1M tokens
OpenAI
LLMVisionCode

GPT-5.6 Sol

OpenAI frontier model for complex professional work, with maximum reasoning, vision, and a 1.05M context window.

Frontier1M context
$5.25 / $31.50Per 1M tokens
Google
Image

Nano Banana 2 Lite

Google fast, low-cost image-output Gemini model. Output fixed at 1K resolution.

33K context
See pricing
Google
Image

Nano Banana 2

Google mid-tier image-output Gemini model. Supports 0.5K/1K/2K/4K output resolutions (this catalog offering exposes up to 2K; see spec for 4K rationale).

33K context
See pricing
Google
Image

Nano Banana Pro

Google premium image-output Gemini model. 1K and 2K are priced identically; this catalog offering exposes up to 2K.

33K context
See pricing
Alibaba Cloud
Image

Wan 2.2 T2I Flash

Alibaba DashScope fast text-to-image model (Wanx family). Submit+poll upstream, existing generateDashScopeImage pipeline, now reachable from the public /v1/images/generations route.

See pricing
Alibaba Cloud
Video

Wan 2.7 Image-to-Video

Video generation from image + text prompt via Alibaba Wan 2.7. Pricing: per 1000 seconds.

Image-to-Video
$100.00Per 1K seconds
Alibaba Cloud
Video

Wan 2.7 Text-to-Video

Video generation from text prompt via Alibaba Wan 2.7. Duration 2-15s, 1080P, native audio. Pricing: per 1000 seconds.

Text-to-Video
$100.00Per 1K seconds
DOS.AI
Embedding

Qwen3 Embedding 4B

Self-hosted multilingual text embedding (2560 dims), strong on Vietnamese. Priced at our OpenRouter fallback cost.

33K context
$0.02 / $0.00Per 1M tokens
Cloudflare
Embedding

Qwen3 Embedding 0.6B

8K context
$0.0118 / $0.00Per 1M tokens
Google
Embedding

Gemini Embedding 001

Google high-dimensional embedding model optimized for semantic search and Retrieval-Augmented Generation (RAG).

Embedding2K context
$0.02 / $0.00Per 1M tokens
Cloudflare
Embedding

BGE Small EN v1.5

1K context
$0.0202 / $0.00Per 1M tokens
Cloudflare
Embedding

BGE Base EN v1.5

1K context
$0.0666 / $0.00Per 1M tokens
Cloudflare
Embedding

BGE Large EN v1.5

1K context
$0.204 / $0.00Per 1M tokens

Ready to get started?

Start building with $5 in free credits. No credit card required.