Providers

Fireworks

Verified

A speed-focused inference cloud for open models, with strong function-calling and structured-output support.

Models on Fireworks

Open in catalog

19 of 19

DeepSeek V4 Flash 0731Chatdeepseek-v4-flash-07311M contextWeights Routeprimary$0.22 / $0.66per M input / output
DeepSeek V4 Pro 0813Chatdeepseek-v4-pro-08131M contextRouteprimary$1.32 / $3.96per M input / output
GLM 5.2Chatglm-5.21M contextWeights Routeprimary$1.40 / $4.40per M input / output
GLM 5.2 FastChatglm-5.2-fast1M contextRouteprimary$2.10 / $6.60per M input / output
GLM 5.3Chatglm-5.31M contextRoutefallback$1.40 / $4.40per M input / output
GLM 5.3 FlashChatglm-5.3-flash1M contextRoutefallback$0.15 / $0.50per M input / output
GPT-OSS 120BChatgpt-oss-120b131K contextWeights Routeprimary$0.35 / $0.75per M input / output
InklingChatinkling1M contextRouteprimary$1.00 / $4.05per M input / output
Kimi K2.6Chatkimi-k2.6262K contextRouteprimary$0.95 / $4.00per M input / output
Kimi K2.7 CodeChatkimi-k2.7-code262K contextRouteprimary$0.95 / $4.00per M input / output
Kimi K3Chatkimi-k31M contextRouteprimary$3.00 / $15.00per M input / output
Kimi K3 FastChatkimi-k3-fast1M contextRouteprimary$4.50 / $22.50per M input / output
MiniMax M3Chatminimax-m3512K contextRouteprimary$0.30 / $1.20per M input / output
Muse Glimmer 30BChatmuse-glimmer-30b131K contextRouteprimary$0.35 / $1.50per M input / output
Nemotron 3 UltraChatnemotron-3-ultra262K contextRouteprimary$0.90 / $2.40per M input / output
Nemotron 3.5 LightningChatnemotron-3.5-lightning262K contextRouteprimary$0.08 / $0.20per M input / output
Qwen 3.7 PlusChatqwen-3.7-plus262K contextRouteprimary$0.50 / $3.00per M input / output
Qwen 3.8 2.4T A95BChatqwen-3.8-2.4t262K contextRouteprimary$2.00 / $6.00per M input / output
Qwen 3.8 MaxChatqwen-3.8-max262K contextRoutefallback$2.00 / $6.00per M input / output

Performance

Request success
100%
Time to first token
4.8s
Generation speed
164.8tok/s

Response latency estimates

Time to first token

Generation speed · tok/s

Request success

Text-model timing and speed, plus API request success. Daily latency percentiles are estimates.

Direct routingCall Fireworks explicitly
fireworks· routing
# pin every request to Fireworks
curl https://www.ninjachat.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NINJACHAT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash-0731",
    "messages": [{ "role": "user", "content": "Hello!" }],
    "routing": { "providers": { "only": ["fireworks"] } }
  }'