Back to models

Gemini 3.7 Flash

ZDRAvailable Serverless
Google•

Gemini 3.7 Flash is Google's hybrid reasoning model, seamlessly combining rapid response with deep, controllable thinking. It features a 1M-token context window, advanced multimodal capabilities (text, code, image), and high agentic performance. Available through the DOS.AI gateway with unified billing and OpenAI compatibility.

Capabilities:Hybrid reasoningVisionLong context (1M)Tool callingCode generationStreaming
Modalities
Input / Output Price
$0.75/$3.75/ 1M
Context Length
1M tokens
Released
Feb 24, 2026

Providers

Different companies host the same model. DOS.AI routes your request to one of them based on the routing mode you pick — Balanced (price + speed), Nitro (fastest), or Exacto (highest tool-calling accuracy).

Reasoning effortⓘAll
ProviderInput /1MOutput /1M
Latency
ThroughputUptimeContextRetention
Google
$0.75$3.75
1.64 s
-
100.0%
1MNot retainedZDR

Pricing

Rates from the DOS.AI catalog for the billing units shown. Caching and provider discounts mean the price actually paid is often below the listed one.

Weighted Averagei
Weighted Avg Input Price
$0.78997607539/ 1M tokens
Equal to list price
Weighted Avg Output Price
$3.91823899371/ 1M tokens
Equal to list price
Price History1 verified endpoints
Weighted Average($0.79)
$0.00$0.31$0.62$0.9430d ago24d ago18d ago12d ago6d agoToday
Provider↕
Effective in /M↕
Effective out /M↕
Cache read /M↕
Latency↕
Throughput↕
Uptime↕
Savings↕
Google
$0.79
$3.92-
1.64 s
-
100.0%
-
Based on settled billing for the last 30d.Effective pricing represents the average price paid for requests to this endpoint, factoring in caching, discounts, and tiered pricing. Listed pricing shows the posted provider list prices.

Observed performance

Provider results over the last 7d.

Updated Oct 1, 12:15 AM UTC
Google
54 samples
Uptime
100.0%
P50 response
1.64 s
P95 response
2.21 s
Latency spread

Uptime is based on observed valid provider attempts. Not an SLA.

Latency is measured to provider response headers, not end-to-end completion.

Observed reliability

Success rate across observed provider requests. This is historical sample data, not live health or an uptime guarantee.

Request observations

Last 7 days

Observed success rate
100.00%
Observed requests
54

Benchmarks

Dated third-party measurements with source links.

No verified benchmark snapshot is available for this model.

Observed gateway latency

Window: 7d

Lowest provider median TTFT
-
Lowest provider median latency
1.64 s
Lowest provider P95 latency
2.21 s
Providers with latency samples
1

API Integration

Use the DOS API to integrate Gemini 3.7 Flash into your applications. Compatible with standard OpenAI SDKs for easy drop-in migration.

</>Run inference
curl https://api.dos.ai/v1/chat/completions \
  -H "Authorization: Bearer $DOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.7-flash",
    "messages": [
      {
        "role": "user",
        "content": "Hello!"
      }
    ]
  }'