NVIDIA's efficient 3B-active hybrid MoE reasoning model on DeepInfra's five-cent input rail.
NVIDIA's compact multimodal reasoning model for cost-efficient visual agents and tool workflows.
| Provider | Context | Price · in / out | p50 | p95 | Uptime |
|---|---|---|---|---|---|
| ddigitaloceanNo trainingNo training on your data. Zero-retention available | 128K | $0.35 / $0.75/MTok | 1.2s | 2.1s | 100% |
routing.strategyproviders.onlyproviders.excludeproviders.ordermax_cost_usddata_policydocs →Collecting — charts appear after two days.
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NINJACHAT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nemotron-nano-12b-v2-vl",
"messages": [{ "role": "user", "content": "Hello!" }],
"stream": true
}'NVIDIA's compact multimodal reasoning model for cost-efficient visual agents and tool workflows.
Nemotron Nano 12B v2 VL costs $0.35/M input tokens and $0.75/M output tokens.
Nemotron Nano 12B v2 VL accepts text and image and returns text.
POST /api/v1/chat/completions with `"model": "nemotron-nano-12b-v2-vl"`.
Nemotron 3 Nano 30B A3B, Nemotron 3 Nano Omni, Nemotron 3 Super, Nemotron 3 Ultra, Nemotron 3.5 Lightning.
NVIDIA's 120B/12B-active hybrid MoE for compute-efficient multi-agent reasoning and tool workflows.
NVIDIA's 550B/55B-active frontier reasoning model for demanding agent workloads.
NVIDIA's 3B-active hybrid Mamba-Transformer reasoning model at five cents per input MTok.