Qwen's 122B/10B-active multimodal MoE, balancing stronger quality with efficient inference.
Qwen's large sparse vision-language flagship for high-accuracy OCR, visual reasoning, coding, and tool orchestration.
| Provider | Context | Price · in / out | p50 | p95 | First token | Uptime | tok/s |
|---|---|---|---|---|---|---|---|
| ddeepinfraNo trainingNo training on your data | 262K | $0.35 / $1.06/MTok | 1.8s | 2.9s | 464ms | 100% | 50 |
routing.strategyproviders.onlyproviders.excludeproviders.ordermax_cost_usddata_policydocs →Collecting — charts appear after two days.
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NINJACHAT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-vl-235b-a22b",
"messages": [{ "role": "user", "content": "Hello!" }],
"stream": true
}'Qwen's large sparse vision-language flagship for high-accuracy OCR, visual reasoning, coding, and tool orchestration.
Qwen3 VL 235B A22B Instruct costs $0.35/M input tokens and $1.06/M output tokens.
Qwen3 VL 235B A22B Instruct accepts text and image and returns text.
POST /api/v1/chat/completions with `"model": "qwen3-vl-235b-a22b"`.
Qwen 3.5 122B A10B, Qwen 3.5 27B, Qwen 3.5 397B A17B, Qwen 3.6 27B, Qwen 3.6 35B A3B, Qwen 3.7 Plus, Qwen 3.8 2.4T A95B, Qwen 3.8 27B, Qwen 3.8 Max, Qwen3 Next 80B A3B, Qwen3 VL 30B A3B Instruct, QwQ 32B.
Qwen's 397B/17B-active multimodal flagship with broad multilingual, coding, and agentic capability.
Qwen's compact multimodal reasoning model on a cost-efficient DeepInfra rail with ultra-fast Groq failover.
Qwen's sparse 35B/3B-active multimodal model for efficient coding and agents.