DeepSeek V4 Flash

by DeepSeek

1M-context fast tier replacing DeepSeek V3

Parameters
MoE (fast tier)
Context Length
1M
Category
chat
Available Serverless

Run queries immediately, pay only for usage

$0.15in|$0.29out

Per 1M tokens

Try this modelView documentation

About this model

DeepSeek V4 Flash is a Mixture-of-Experts model tuned for fast reasoning. It supports math, logic, code generation, and a 1M-token context. Current input and output rates are shown from the live catalog.

Capabilities

ReasoningMath and logicCode generationLong context (1M)Efficient inference

Use Cases

  • High-volume reasoning
  • Code development
  • Data analysis
  • Batch processing

Providers

One model, several sources. DOS routes to the first that answers and falls back automatically, or you can pin a provider per request.

ProviderInput /1MOutput /1MContextServed from
DeepSeek (via Alibaba)Available
$0.15$0.291MSG

Model Details

Provider
DeepSeek
Model ID
deepseek-v4-flash
Parameters
MoE (fast tier)
Context Length
1M tokens
Category
chat

API Usage

Use the DOS API to integrate DeepSeek V4 Flash into your applications. Our API is compatible with OpenAI's client libraries for easy migration.

Model ID

deepseek-v4-flash

Python

python
from dos import DOS

client = DOS()

response = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[
        {"role": "user", "content": "Hello, how are you?"}
    ]
)

print(response.choices[0].message.content)

cURL

bash
curl https://api.dos.ai/v1/chat/completions \
  -H "Authorization: Bearer $DOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [
      {"role": "user", "content": "Hello, how are you?"}
    ]
  }'

Node.js

javascript
import DOS from 'dos-ai';

const client = new DOS();

const response = await client.chat.completions.create({
  model: "deepseek-v4-flash",
  messages: [
    { role: "user", content: "Hello, how are you?" }
  ]
});

console.log(response.choices[0].message.content);