NVIDIA's efficient 3B-active hybrid MoE reasoning model on DeepInfra's five-cent input rail.
NVIDIA's compact Omni model for image-aware reasoning and function-driven agent workloads.
| Provider | Context | Price · in / out | p50 | p95 | Uptime |
|---|---|---|---|---|---|
| ddigitaloceanNo trainingNo training on your data. Zero-retention available | 66K | $0.65 / $1.08/MTok | 1.2s | 1.4s | 100% |
routing.strategyproviders.onlyproviders.excludeproviders.ordermax_cost_usddata_policydocs →Collecting — charts appear after two days.
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NINJACHAT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nemotron-3-nano-omni",
"messages": [{ "role": "user", "content": "Hello!" }],
"stream": true
}'NVIDIA's compact Omni model for image-aware reasoning and function-driven agent workloads.
Nemotron 3 Nano Omni costs $0.65/M input tokens and $1.08/M output tokens.
Nemotron 3 Nano Omni accepts text and image and returns text.
POST /api/v1/chat/completions with `"model": "nemotron-3-nano-omni"`.
Nemotron 3 Nano 30B A3B, Nemotron 3 Super, Nemotron 3 Ultra, Nemotron 3.5 Lightning, Nemotron Nano 12B v2 VL.
NVIDIA's 550B/55B-active frontier reasoning model for demanding agent workloads.
NVIDIA's 3B-active hybrid Mamba-Transformer reasoning model at five cents per input MTok.
NVIDIA's compact multimodal reasoning model for cost-efficient visual agents and tool workflows.