DeepSeek V4 Flash
by DeepSeek
1M-context fast tier replacing DeepSeek V3
Run queries immediately, pay only for usage
Per 1M tokens
About this model
DeepSeek V4 Flash is a Mixture-of-Experts model tuned for fast reasoning. It supports math, logic, code generation, and a 1M-token context. Current input and output rates are shown from the live catalog.
Capabilities
Use Cases
- High-volume reasoning
- Code development
- Data analysis
- Batch processing
Providers
One model, several sources. DOS routes to the first that answers and falls back automatically, or you can pin a provider per request.
| Provider | Input /1M | Output /1M | Context | Served from |
|---|---|---|---|---|
DeepSeek (via Alibaba)Available | $0.15 | $0.29 | 1M | SG |
Model Details
- Provider
- DeepSeek
- Model ID
- deepseek-v4-flash
- Parameters
- MoE (fast tier)
- Context Length
- 1M tokens
- Category
- chat
API Usage
Use the DOS API to integrate DeepSeek V4 Flash into your applications. Our API is compatible with OpenAI's client libraries for easy migration.
Model ID
deepseek-v4-flashPython
from dos import DOS
client = DOS()
response = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[
{"role": "user", "content": "Hello, how are you?"}
]
)
print(response.choices[0].message.content)cURL
curl https://api.dos.ai/v1/chat/completions \
-H "Authorization: Bearer $DOS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-flash",
"messages": [
{"role": "user", "content": "Hello, how are you?"}
]
}'Node.js
import DOS from 'dos-ai';
const client = new DOS();
const response = await client.chat.completions.create({
model: "deepseek-v4-flash",
messages: [
{ role: "user", content: "Hello, how are you?" }
]
});
console.log(response.choices[0].message.content);Related Models
DOS.AI Auto
Smart routing - automatically picks the best model for your request. Free for simple tasks, paid models for complex ones.
DOS.AI
DOS.AI's hosted Qwen 3.5 35B-A3B model for text generation.
Gemini 3.7 Flash
Google hybrid reasoning model combining fast multimodal inference with controllable thinking capabilities, a 1M context window, and up to 64K output tokens.