NVIDIA's compact Omni model for image-aware reasoning and function-driven agent workloads.
| Provider | Context | Price · in / out | p50 | p95 | Uptime |
|---|---|---|---|---|---|
| ddeepinfraNo trainingNo training on your data | 262K | $0.20 / $0.35/MTok | 816ms | 816ms | 100% |
routing.strategyproviders.onlyproviders.excludeproviders.ordermax_cost_usddata_policydocs →Collecting — charts appear after two days.
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NINJACHAT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nemotron-3-nano",
"messages": [{ "role": "user", "content": "Hello!" }],
"stream": true
}'NVIDIA's efficient 3B-active hybrid MoE reasoning model on DeepInfra's five-cent input rail.
Nemotron 3 Nano 30B A3B costs $0.20/M input tokens and $0.35/M output tokens.
Nemotron 3 Nano 30B A3B accepts text and returns text.
POST /api/v1/chat/completions with `"model": "nemotron-3-nano"`.
Nemotron 3 Nano Omni, Nemotron 3 Super, Nemotron 3 Ultra, Nemotron 3.5 Lightning, Nemotron Nano 12B v2 VL.
NVIDIA's 550B/55B-active frontier reasoning model for demanding agent workloads.
NVIDIA's 3B-active hybrid Mamba-Transformer reasoning model at five cents per input MTok.
NVIDIA's compact multimodal reasoning model for cost-efficient visual agents and tool workflows.