Gemini 3.1 Flash-Lite
ZDRFastest and most cost-efficient Gemini 3 model
Providers
Different companies host the same model. DOS.AI routes your request to one of them based on the routing mode you pick — Balanced (price + speed), Nitro (fastest), or Exacto (highest tool-calling accuracy).
| Provider | Input /1M | Output /1M | Latency | Uptime |
|---|---|---|---|---|
Google | $0.25 | $1.50 | 761 ms | 100.0% |
Pricing
Rates from the DOS.AI catalog for the billing units shown. Caching and provider discounts mean the price actually paid is often below the listed one.
Provider↕ | Effective in /M↕ | Effective out /M↕ | Cache read /M↕ | Latency↕ | Uptime↕ | Savings↕ |
|---|---|---|---|---|---|---|
Google | $0.25 | $1.50 | - | 761 ms | 100.0% | - |
Observed performance
Provider results over the last 7d.
- Uptime
- 100.0%
- P50 response
- 761 ms
- P95 response
- 1.07 s
- Latency spread
Uptime is based on observed valid provider attempts. Not an SLA.
Latency is measured to provider response headers, not end-to-end completion.
Observed reliability
Success rate across observed provider requests. This is historical sample data, not live health or an uptime guarantee.
Request observations
Last 7 days
- Observed success rate
- 100.00%
- Observed requests
- 47
Benchmarks
Dated third-party measurements with source links.
No verified benchmark snapshot is available for this model.
Observed gateway latency
Window: 7d
- Lowest provider median TTFT
- -
- Lowest provider median latency
- 761 ms
- Lowest provider P95 latency
- 1.07 s
- Providers with latency samples
- 1
API Integration
Use the DOS API to integrate Gemini 3.1 Flash-Lite into your applications. Compatible with standard OpenAI SDKs for easy drop-in migration.
curl https://api.dos.ai/v1/chat/completions \
-H "Authorization: Bearer $DOS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.1-flash-lite",
"messages": [
{
"role": "user",
"content": "Hello!"
}
]
}'