Providers

Baseten

Verified

A production inference platform serving open-weight models on dedicated and shared GPU fleets.

Models on Baseten

Open in catalog

13 of 13

DeepSeek V4 Flash 0731Chatdeepseek-v4-flash-07311M contextWeights Routeemergency$0.22 / $0.66per M input / output
DeepSeek V4 ProChatdeepseek-v4-pro1M contextWeights Routefallback$1.74 / $3.48per M input / output
DeepSeek V4 Pro 0813Chatdeepseek-v4-pro-08131M contextRouteemergency$1.32 / $3.96per M input / output
GLM 4.7Chatglm-4.7200K contextWeights Routefallback$0.60 / $2.20per M input / output
GLM 5.2Chatglm-5.21M contextWeights Routeemergency$1.40 / $4.40per M input / output
GLM 5.2 FastChatglm-5.2-fast1M contextRoutefallback$2.10 / $6.60per M input / output
GPT-OSS 120BChatgpt-oss-120b131K contextWeights Routefallback$0.35 / $0.75per M input / output
InklingChatinkling1M contextRoutefallback$1.00 / $4.05per M input / output
Inkling SmallChatinkling-small524K contextRoutefallback$0.50 / $1.20per M input / output
Kimi K2.6Chatkimi-k2.6262K contextRouteemergency$0.95 / $4.00per M input / output
Kimi K2.7 CodeChatkimi-k2.7-code262K contextRouteemergency$0.95 / $4.00per M input / output
Kimi K3Chatkimi-k31M contextRouteemergency$3.00 / $15.00per M input / output
Nemotron 3 UltraChatnemotron-3-ultra262K contextRouteemergency$0.90 / $2.40per M input / output

Performance

Request success
100%
Time to first token
13.0s
Generation speed
326.6tok/s

Response latency estimates

Time to first token

Generation speed · tok/s

Request success

Text-model timing and speed, plus API request success. Daily latency percentiles are estimates.

Direct routingCall Baseten explicitly
baseten· routing
# pin every request to Baseten
curl https://www.ninjachat.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NINJACHAT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-pro",
    "messages": [{ "role": "user", "content": "Hello!" }],
    "routing": { "providers": { "only": ["baseten"] } }
  }'