Qwen's fast multimodal flagship for agent loops, coding, and tool use.
Qwen's sparse 35B/3B-active multimodal model for efficient coding and agents.
| Provider | Context | Price · in / out | p50 | p95 | Uptime |
|---|---|---|---|---|---|
| ddeepinfraNo trainingNo training on your data | 262K | $0.25 / $1.14/MTok | 24.6s | 44.5s | 100% |
routing.strategyproviders.onlyproviders.excludeproviders.ordermax_cost_usddata_policydocs →Collecting — charts appear after two days.
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NINJACHAT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-3.6-35b-a3b",
"messages": [{ "role": "user", "content": "Hello!" }],
"stream": true
}'Qwen's sparse 35B/3B-active multimodal model for efficient coding and agents.
Qwen 3.6 35B A3B costs $0.25/M input tokens and $1.14/M output tokens.
Qwen 3.6 35B A3B accepts text and image and returns text.
POST /api/v1/chat/completions with `"model": "qwen-3.6-35b-a3b"`.
Qwen 3.7 Plus, Qwen 3.8 2.4T A95B, Qwen 3.8 27B, Qwen 3.8 Max, Qwen3 Next 80B A3B, QwQ 32B.
Qwen's efficient 27B model on independently operated DeepInfra and GMI rails.
Qwen's 2.4T-parameter sparse-MoE frontier model for autonomous long-horizon work.
Qwen's efficient 80B/3B-active instruct model on two independent direct rails.