Providers

Groq

Verified

Custom LPU silicon purpose-built for inference — ultra-low latency at very high token speeds.

Models on Groq

Open in catalog

4 of 4

GPT-OSS 120BChatgpt-oss-120b131K contextWeights Routefallback$0.35 / $0.75per M input / output
GPT-OSS 20BChatgpt-oss-20b131K contextWeights Routeprimary$0.07 / $0.45per M input / output
Qwen 3.6 27BChatqwen-3.6-27b131K contextWeights Routefallback$0.60 / $3.20per M input / output
QwQ 32BChatqwq-32b32K contextWeights Routeprimary$0.29 / $0.59per M input / output
Direct routingCall Groq explicitly
groq· routing
# pin every request to Groq
curl https://www.ninjachat.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NINJACHAT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-oss-20b",
    "messages": [{ "role": "user", "content": "Hello!" }],
    "routing": { "providers": { "only": ["groq"] } }
  }'