Qwen's 122B/10B-active multimodal MoE, balancing stronger quality with efficient inference.
Qwen's efficient 8B vision-language model with tools, structured outputs, and 262K context.
| Provider | Context | Price · in / out | p50 | p95 | Uptime |
|---|---|---|---|---|---|
| ddeepinfraNo trainingNo training on your data | 262K | $0.33 / $0.84/MTok | 302ms | 415ms | 100% |
routing.strategyproviders.onlyproviders.excludeproviders.ordermax_cost_usddata_policydocs →Collecting — charts appear after two days.
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NINJACHAT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-vl-8b",
"messages": [{ "role": "user", "content": "Hello!" }],
"stream": true
}'Qwen's efficient 8B vision-language model with tools, structured outputs, and 262K context.
Qwen3 VL 8B costs $0.33/M input tokens and $0.84/M output tokens.
Qwen3 VL 8B accepts text and image and returns text.
POST /api/v1/chat/completions with `"model": "qwen3-vl-8b"`.
Qwen 3.5 122B A10B, Qwen 3.5 27B, Qwen 3.5 35B A3B, Qwen 3.5 397B A17B, Qwen 3.6 27B, Qwen 3.6 35B A3B, Qwen 3.7 Plus, Qwen 3.8 2.4T A95B, Qwen 3.8 27B, Qwen 3.8 Max, Qwen3 235B A22B Instruct, Qwen3 Coder 30B A3B, Qwen3 Next 80B A3B, Qwen3 VL 235B A22B Instruct, Qwen3 VL 30B A3B Instruct, Qwen3 VL 4B, QwQ 32B.
Qwen's efficient 35B MoE activates only 3B parameters per token while supporting tools and vision.
Qwen's 397B/17B-active multimodal flagship with broad multilingual, coding, and agentic capability.
Qwen's compact multimodal reasoning model on a cost-efficient DeepInfra rail with ultra-fast Groq failover.