Gemini 1.5 Flash-8B

by Google

Google ultra-lightweight 8B multimodal model engineered for high-volume text and image processing at minimal cost.

Parameters
Flash-8B
Context Length
1M
Category
chat
Available Serverless

Run queries immediately, pay only for usage

$0.04in|$0.16out

Per 1M Tokens

Try this modelView documentation

About this model

Gemini 1.5 Flash-8B is an ultra-lightweight 8-billion parameter model engineered for high-frequency queries, simple classification, and high-throughput content workflows.

Capabilities

Ultra-lightweightFast inference1M contextEconomical

Use Cases

  • High-throughput extraction
  • Simple chat
  • Routing & classification

Model Details

Provider
Google
Model ID
gemini-1.5-flash-8b
Parameters
Flash-8B
Context Length
1M tokens
Category
chat

API Usage

Use the DOS API to integrate Gemini 1.5 Flash-8B into your applications. Our API is compatible with OpenAI's client libraries for easy migration.

Model ID

gemini-1.5-flash-8b

Python

python
from dos import DOS

client = DOS()

response = client.chat.completions.create(
    model="gemini-1.5-flash-8b",
    messages=[
        {"role": "user", "content": "Hello, how are you?"}
    ]
)

print(response.choices[0].message.content)

cURL

bash
curl https://api.dos.ai/v1/chat/completions \
  -H "Authorization: Bearer $DOS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-1.5-flash-8b",
    "messages": [
      {"role": "user", "content": "Hello, how are you?"}
    ]
  }'

Node.js

javascript
import DOS from 'dos-ai';

const client = new DOS();

const response = await client.chat.completions.create({
  model: "gemini-1.5-flash-8b",
  messages: [
    { role: "user", content: "Hello, how are you?" }
  ]
});

console.log(response.choices[0].message.content);