Gemini 2.0 Flash

by Google

Google next-generation multimodal Flash model offering sub-second latency, multimodal understanding, and rich tool use across a 1M context.

Parameters
Flash
Context Length
1M
Category
chat
Available Serverless

Run queries immediately, pay only for usage

$0.11in|$0.42out

Per 1M Tokens

Try this modelView documentation

About this model

Gemini 2.0 Flash combines next-generation multimodal understanding with ultra-fast inference speed and a 1M context window, making it the ideal engine for production agentic applications.

Capabilities

Sub-second latencyVision & audio1M contextTool callingCode generation

Use Cases

  • Real-time AI agents
  • High-throughput chat
  • Vision analysis
  • Fast tool calling

Model Details

Provider
Google
Model ID
gemini-2.0-flash
Parameters
Flash
Context Length
1M tokens
Category
chat

API Usage

Use the DOS API to integrate Gemini 2.0 Flash into your applications. Our API is compatible with OpenAI's client libraries for easy migration.

Model ID

gemini-2.0-flash

Python

python
from dos import DOS

client = DOS()

response = client.chat.completions.create(
    model="gemini-2.0-flash",
    messages=[
        {"role": "user", "content": "Hello, how are you?"}
    ]
)

print(response.choices[0].message.content)

cURL

bash
curl https://api.dos.ai/v1/chat/completions \
  -H "Authorization: Bearer $DOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-2.0-flash",
    "messages": [
      {"role": "user", "content": "Hello, how are you?"}
    ]
  }'

Node.js

javascript
import DOS from 'dos-ai';

const client = new DOS();

const response = await client.chat.completions.create({
  model: "gemini-2.0-flash",
  messages: [
    { role: "user", content: "Hello, how are you?" }
  ]
});

console.log(response.choices[0].message.content);