NVIDIA's compact 9B controllable-reasoning model with fast tool calls.
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NINJACHAT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nemotron-nano-9b-v2",
"messages": [{ "role": "user", "content": "Hello!" }],
"stream": true
}'Each request below is generated from what Nemotron Nano 9B V2 serves on NinjaChat today — the same capabilities and parameters that GET /models reports — so it runs as written with your key.
Show text as it is generated instead of waiting for the whole completion.
# export NINJACHAT_API_KEY="nj_sk_..." (Developers → Keys)
curl -N https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NINJACHAT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nemotron-nano-9b-v2",
"messages": [{ "role": "user", "content": "Write a haiku about deploy day." }],
"stream": true,
"stream_options": { "include_usage": true }
}'data: line carries one chat.completion.chunk and the stream ends with data: [DONE]. Both SDKs parse that for you and stop at the sentinel.stream_options.include_usage the last chunk before [DONE] has an empty choices array and a usage object with the billed token counts, so guard on choices.length before reading a delta.curl -N so the buffer is not held back; the final usage chunk also carries cost_usd for this request.Let the model decide when to call your function and hand you typed arguments.
# export NINJACHAT_API_KEY="nj_sk_..." (Developers → Keys)
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NINJACHAT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nemotron-nano-9b-v2",
"messages": [{ "role": "user", "content": "Do I need an umbrella in Lisbon today?" }],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": { "city": { "type": "string" } },
"required": ["city"]
}
}
}],
"tool_choice": "auto"
}'tool_choice: "auto" lets the model answer directly when no tool is needed; "required" forces at least one call, and { "type": "function", "function": { "name": "get_weather" } } pins a specific one.function.arguments, never as an object — parse before use. Send your result back as a tool message with the same tool_call_id, then call again for the natural-language answer.parallel_tool_calls: false when your functions must run one at a time.Have the model return an object that matches your schema, so you can parse it without cleanup.
# export NINJACHAT_API_KEY="nj_sk_..." (Developers → Keys)
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NINJACHAT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nemotron-nano-9b-v2",
"messages": [{ "role": "user", "content": "Extract the invoice: ACME Ltd billed 1,240.50 EUR." }],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "invoice",
"strict": true,
"schema": {
"type": "object",
"properties": {
"vendor": { "type": "string" },
"total": { "type": "number" },
"currency": { "type": "string" }
},
"required": ["vendor", "total", "currency"],
"additionalProperties": false
}
}
}
}'json_schema with strict: true constrains the output to your schema; mark every property required and set additionalProperties: false so the object is exactly what you parse.{ "type": "json_object" } is the looser form — valid JSON with no schema. Whichever you use, message.content is the JSON string; parse it, do not regex it.