Gemini 3.7 Flash

Google

Google's newest Flash-class model: a million-token context, native image and video input, and reasoning tuned for agentic coding at Flash speed.

Model highlights
  • Million-token context
  • Native image and video input
  • Adjustable thinking levels
  • Tool calling and structured output
  • Flash-tier speed
Modalities
IntextimageOuttext
Input / output
$0.75 / $3.75USD per 1M tokens
Context
1M
Providers
2 live

Chat with Gemini 3.7 Flash free →

Playground

Preparing playground

Providers

primary6.5s100%
emergency——

Performance

Request success
100%
Time to first token
2.3s
Generation speed
327.9tok/s

Response latency estimates

Time to first token

Generation speed · tok/s

Request success

Text-model timing and speed, plus API request success. Daily latency percentiles are estimates. Your request logs

API

POST/api/v1/chat/completionsOpenAI-compatible
gemini-3.7-flash
curl https://www.ninjachat.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NINJACHAT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.7-flash",
    "messages": [{ "role": "user", "content": "Hello!" }],
    "stream": true
  }'
messagestemperaturemax_completion_tokenstop_pstopfrequency_penaltypresence_penaltyseedstreamuserroutingtoolstool_choiceresponse_formatimage_url content partsreasoningreasoning_effort

Pricing

Input
$0.75/M tokens
Cached input
$0.075/M tokens
Output
$3.75/M tokens

Published USD rates. Actual cost depends on input, output, cache usage and request options.

Further reading