NinjaChat
ModelsAPIPricingToolsBlog
Sign inDashboard→
ModelsAPIPricingToolsBlog
Dashboard →
  1. Home›
  2. API models›
  3. Nemotron Nano 9B V2

Nemotron Nano 9B V2

Amazon Bedrock/
StreamingJSON modeTool callingReasoningLong context
Get an API key

NVIDIA's compact 9B controllable-reasoning model with fast tool calls.

Modalities
text→text
Price
$0.06 / $0.23/M tok
Context
131K
Providers
Live
PlaygroundProvidersAPIRecipesPricingRequest logs

Playground

Preparing playground

Providers

API

POST/api/v1/chat/completionsOpenAI-compatible
nemotron-nano-9b-v2
curl https://www.ninjachat.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NINJACHAT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nemotron-nano-9b-v2",
    "messages": [{ "role": "user", "content": "Hello!" }],
    "stream": true
  }'
messagestemperaturemax_completion_tokenstop_pstopfrequency_penaltypresence_penaltyseedstreamuserroutingtoolstool_choiceresponse_formatreasoningreasoning_effort

Recipes for Nemotron Nano 9B V2

Each request below is generated from what Nemotron Nano 9B V2 serves on NinjaChat today — the same capabilities and parameters that GET /models reports — so it runs as written with your key.

Stream tokens

Show text as it is generated instead of waiting for the whole completion.

nemotron-nano-9b-v2· stream tokens
# export NINJACHAT_API_KEY="nj_sk_..."   (Developers → Keys)
curl -N https://www.ninjachat.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NINJACHAT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nemotron-nano-9b-v2",
    "messages": [{ "role": "user", "content": "Write a haiku about deploy day." }],
    "stream": true,
    "stream_options": { "include_usage": true }
  }'
  • The response is Server-Sent Events: each data: line carries one chat.completion.chunk and the stream ends with data: [DONE]. Both SDKs parse that for you and stop at the sentinel.
  • With stream_options.include_usage the last chunk before [DONE] has an empty choices array and a usage object with the billed token counts, so guard on choices.length before reading a delta.
  • Keep curl -N so the buffer is not held back; the final usage chunk also carries cost_usd for this request.

Call a tool

Let the model decide when to call your function and hand you typed arguments.

nemotron-nano-9b-v2· call a tool
# export NINJACHAT_API_KEY="nj_sk_..."   (Developers → Keys)
curl https://www.ninjachat.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NINJACHAT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nemotron-nano-9b-v2",
    "messages": [{ "role": "user", "content": "Do I need an umbrella in Lisbon today?" }],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Current weather for a city",
        "parameters": {
          "type": "object",
          "properties": { "city": { "type": "string" } },
          "required": ["city"]
        }
      }
    }],
    "tool_choice": "auto"
  }'
  • tool_choice: "auto" lets the model answer directly when no tool is needed; "required" forces at least one call, and { "type": "function", "function": { "name": "get_weather" } } pins a specific one.
  • Arguments arrive as a JSON string in function.arguments, never as an object — parse before use. Send your result back as a tool message with the same tool_call_id, then call again for the natural-language answer.
  • Up to 32 tools per request; set parallel_tool_calls: false when your functions must run one at a time.

Get strict JSON

Have the model return an object that matches your schema, so you can parse it without cleanup.

nemotron-nano-9b-v2· get strict json
# export NINJACHAT_API_KEY="nj_sk_..."   (Developers → Keys)
curl https://www.ninjachat.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NINJACHAT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nemotron-nano-9b-v2",
    "messages": [{ "role": "user", "content": "Extract the invoice: ACME Ltd billed 1,240.50 EUR." }],
    "response_format": {
      "type": "json_schema",
      "json_schema": {
        "name": "invoice",
        "strict": true,
        "schema": {
          "type": "object",
          "properties": {
            "vendor": { "type": "string" },
            "total": { "type": "number" },
            "currency": { "type": "string" }
          },
          "required": ["vendor", "total", "currency"],
          "additionalProperties": false
        }
      }
    }
  }'
  • json_schema with strict: true constrains the output to your schema; mark every property required and set additionalProperties: false so the object is exactly what you parse.
  • { "type": "json_object" } is the looser form — valid JSON with no schema. Whichever you use, message.content is the JSON string; parse it, do not regex it.

Pricing

Input
$0.06/M tokens
Cached input
$0.006/M tokens
Output
$0.23/M tokens

More from Nemotron

Model
ProvidersContextLatencyUptime30d volumePrice · in / out
Nemotron 3 Nano 30B A3B
nemotron-3-nano
262K$0.05/$0.20
Nemotron 3 Nano Omni
nemotron-3-nano-omni
66K$0.50/$0.90
Nemotron 3 Super
nemotron-3-super
+1262K$0.30/$0.65
Nemotron 3 Ultra
nemotron-3-ultra
+1262K$0.90/$2.40
Nemotron 3.5 Lightning
nemotron-3.5-lightning
262K$0.08/$0.20
Nemotron Nano 12B v2 VL
nemotron-nano-12b-v2-vl
128K$0.20/$0.60
NinjaChat

Every AI. One app.

Download on the App Store

Product

  • Dashboard
  • ninja for iMessage
  • Pricing
  • Free AI Tools
  • Affiliate Program
  • iOS App

Developers

  • API
  • API Models
  • MCP / Agents
  • API Docs

Models

  • Model Council
  • Seed 1.8
  • Gemini 2.5 Flash
  • Gemini 2.5 Pro
  • Gemini 3 Flash Preview
  • View All Models

Company

  • Blog
  • Uncensored AI
  • Community
  • Careers
  • Support
  • Privacy Policy
  • Terms of Service

Copyright © 2026 NinjaChat AI. Product of Bloon All Rights Reserved.