Providers

Google Vertex AI

Verified

Models, pricing, and routing options for Google Vertex AI.

Models on Google Vertex AI

Open in catalog

2 of 2

Gemini 2.5 Flash-LiteChatgemini-2.5-flash-lite1M contextRoutefallback$0.10 / $0.40per M input / output
Gemini 3.8 FlashChatgemini-3.8-flash1M contextRoutefallback$0.75 / $3.75per M input / output

Performance

Request success
100%
Time to first token
14.7s
Generation speed
108.9tok/s

Response latency estimates

Time to first token

Generation speed · tok/s

Request success

Text-model timing and speed, plus API request success. Daily latency percentiles are estimates.

Direct routingCall Google Vertex AI explicitly
vertex· routing
# pin every request to Google Vertex AI
curl https://www.ninjachat.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NINJACHAT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-2.5-flash-lite",
    "messages": [{ "role": "user", "content": "Hello!" }],
    "routing": { "providers": { "only": ["vertex"] } }
  }'