Models (64)
Featured Models
Claude Opus 5
Anthropic flagship model for complex agentic coding and enterprise work, with a 1M context window and up to 128K synchronous output tokens.
Claude Sonnet 5
Anthropic balanced model with near-flagship coding and agentic quality at Sonnet pricing, a 1M context window, and up to 128K output tokens.
Gemini 3.8 Flash
Google's newest Gemini Flash model, with a 1M token context window.
DeepSeek V4.1 Flash
Qwen3.8-27B
DOS.AI hosted Qwen3.8-27B for reasoning, coding, tool use, and multilingual text generation.
GPT-5.6 Luna
OpenAI most economical GPT-5.6 model, built for efficient high-volume workloads, with vision and a 1.05M context window.
All Models
Auto
DOS.AI Smart Router classifies each request by complexity. Simple and Medium requests use DOS.AI; Complex requests can use an eligible paid route configured in the live catalog.
DOS
DOS curated meta-model for DOSClaw and first-party DOS products.
Qwen3.8-27B
DOS.AI hosted Qwen3.8-27B for reasoning, coding, tool use, and multilingual text generation.
Granite 4.0 H Micro
Llama 3.2 1B Instruct
GPT-OSS 120B
OpenAI open-weight 120B, strong general reasoning at a fraction of hosted GPT pricing.
GPT-OSS 20B
OpenAI open-weight 20B, the cheapest capable chat model in the catalog.
Llama 3.2 3B Instruct
Qwen3 30B A3B FP8
GLM 4.7 Flash
Gemma 4 26B A4B
Tencent HY3
Xiaomi MiMo V2.5
GLM 5.3 Flash
Qwen3.8 Flash
Llama 3.1 8B
GPT-5.4 Nano
Cheapest GPT-5.4-class model for simple high-volume tasks
GPT-5.6 Luna
OpenAI most economical GPT-5.6 model, built for efficient high-volume workloads, with vision and a 1.05M context window.
Gemini 3.1 Flash-Lite
Fastest and most cost-efficient Gemini 3 model
Llama 3.3 70B
MiniMax M2.7
DeepSeek V4.1 Flash
Gemini 3.5 Flash-Lite
Fast, cost-effective Flash-Lite (GA). 1M context. Closest Gemini peer to dos-ai (Qwen3.6-35B-A3B). [PROMO] Free when used BY a DOSClaw agent on a paid plan (agent traffic only - direct API usage is billed normally). Limited-time.
Gemma SEA-LION v4 27B
Mistral Small 3.1 24B
Qwen 3.7 Plus
Qwen 3.7 Plus - multimodal (text+image), 1M context
DeepSeek R1 Distill Qwen 32B
Nemotron 3 120B A12B
NVIDIA Nemotron 3 Ultra
Qwen2.5 Coder 32B
QwQ 32B
Gemini 3.1 Flash Live
Real-time voice and dialogue model
Gemini 3.7 Flash
Google hybrid reasoning model combining fast multimodal inference with controllable thinking capabilities, a 1M context window, and up to 64K output tokens.
Gemini 3.8 Flash
Google's newest Gemini Flash model, with a 1M token context window.
GPT-5.4 Mini
Strong mini model for coding, computer use, and sub-agents
Kimi K2.6
Kimi K2.7 Code
Moonshot Kimi K2.7 Code, coding-focused with a 256K context window.
Claude Haiku 4.5
Fastest and most compact Claude model
Grok 4.3
xAI flagship reasoning model - 1M context, text+image (replaces grok-4.1-fast)
DeepSeek V4 Pro
Near-frontier quality at ~1/6 the cost of Opus 4.7 / GPT-5.5
GLM-5.2
Zhipu GLM-5.2 with a 1M context window, strong Chinese and English coding.
GLM 5.3
Gemini 3.5 Flash
Latest Gemini Flash - frontier performance, standard tier
Qwen 3.8 Max
Qwen 3.8 Max - Alibaba's 2.4T-parameter MoE flagship, 1M context, multimodal input (text+image+video)
Gemini 3.1 Pro
Google's most advanced reasoning model for complex tasks
GPT-5.6 Terra
OpenAI mid tier of the GPT-5.6 family, balancing intelligence and cost with vision, tool calling, and a 1.05M context window.
Grok 4.20
xAI flagship reasoning model with 2M context
Qwen 3.7 Max
Qwen 3.7 Max tier - 1M context (standard price; OpenRouter shows a 50%-off promo)
Kimi K3
Claude Sonnet 5
Anthropic balanced model with near-flagship coding and agentic quality at Sonnet pricing, a 1M context window, and up to 128K output tokens.
Claude Opus 5
Anthropic flagship model for complex agentic coding and enterprise work, with a 1M context window and up to 128K synchronous output tokens.
GPT-5.6 Sol
OpenAI frontier model for complex professional work, with maximum reasoning, vision, and a 1.05M context window.
Nano Banana 2 Lite
Google fast, low-cost image-output Gemini model. Output fixed at 1K resolution.
Nano Banana 2
Google mid-tier image-output Gemini model. Supports 0.5K/1K/2K/4K output resolutions (this catalog offering exposes up to 2K; see spec for 4K rationale).
Nano Banana Pro
Google premium image-output Gemini model. 1K and 2K are priced identically; this catalog offering exposes up to 2K.
Wan 2.2 T2I Flash
Alibaba DashScope fast text-to-image model (Wanx family). Submit+poll upstream, existing generateDashScopeImage pipeline, now reachable from the public /v1/images/generations route.
Wan 2.7 Image-to-Video
Video generation from image + text prompt via Alibaba Wan 2.7. Pricing: per 1000 seconds.
Wan 2.7 Text-to-Video
Video generation from text prompt via Alibaba Wan 2.7. Duration 2-15s, 1080P, native audio. Pricing: per 1000 seconds.
Qwen3 Embedding 4B
Self-hosted multilingual text embedding (2560 dims), strong on Vietnamese. Priced at our OpenRouter fallback cost.
Qwen3 Embedding 0.6B
Gemini Embedding 001
Google high-dimensional embedding model optimized for semantic search and Retrieval-Augmented Generation (RAG).
BGE Small EN v1.5
BGE Base EN v1.5
BGE Large EN v1.5
Ready to get started?
Start building with $5 in free credits. No credit card required.