Providers

Together AI

Verified

Models, pricing, and routing options for Together AI.

Models on Together AI

Open in catalog

5 of 5

DeepSeek V4 FlashChatdeepseek-v4-flash1M contextWeights Routefallback$0.14 / $0.28per M input / output
GLM 5.3 FlashChatglm-5.3-flash1M contextRoutefallback$0.15 / $0.50per M input / output
GPT-OSS 120BChatgpt-oss-120b131K contextWeights Routefallback$0.35 / $0.75per M input / output
Kimi K3Chatkimi-k31M contextRoutefallback$3.00 / $15.00per M input / output
MiniMax M3Chatminimax-m3512K contextRoutefallback$0.30 / $1.20per M input / output

Performance

Request success
100%
Time to first token
4.2s
Generation speed
1039.6tok/s

Response latency estimates

Time to first token

Generation speed · tok/s

Request success

Text-model timing and speed, plus API request success. Daily latency percentiles are estimates.

Direct routingCall Together AI explicitly
together· routing
# pin every request to Together AI
curl https://www.ninjachat.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NINJACHAT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [{ "role": "user", "content": "Hello!" }],
    "routing": { "providers": { "only": ["together"] } }
  }'