Same OpenAI-compatible mechanics as Fireworks' serverless endpoint — shorter model ids, one prepaid per-token balance, and frontier closed models next to your open ones.
api.fireworks.ai/inference/v1www.ninjachat.ai/api/v1accounts/fireworks/models/deepseek-v3p2→deepseek-v3.2accounts/fireworks/models/qwq-32b→qwq-32baccounts/fireworks/models/kimi-k2-instruct→kimi-k2.6from openai import OpenAI
client = OpenAI(
base_url="https://api.fireworks.ai/inference/v1",
api_key="fw_...",
)
resp = client.chat.completions.create(
model="accounts/fireworks/models/deepseek-v3p2",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)from openai import OpenAI
client = OpenAI(
base_url="https://www.ninjachat.ai/api/v1", # ← changed
api_key="nj_sk_YOUR_KEY", # ← changed
)
resp = client.chat.completions.create(
model="deepseek-v3.2",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)Get an nj_sk_ key at /developers and add a credit pack. Balances are prepaid, and failed generations are refunded.
# Sanity check: list every model your key can call (no auth needed)
curl https://www.ninjachat.ai/api/v1/models
# First authenticated request
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer nj_sk_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-5", "messages": [{"role": "user", "content": "Say hi"}]}'Fireworks ids carry the deployment path. NinjaChat slugs are just the model: accounts/fireworks/models/deepseek-v3p2 → deepseek-v3.2. Mapping table below; anything not listed can be checked live via GET /v1/models.
Same SDK code. Streaming arrives as SSE deltas, and tool_calls round-trip with role: "tool" messages.
# Streaming
stream = client.chat.completions.create(
model="claude-sonnet-4.6",
messages=[{"role": "user", "content": "Write a haiku about ninjas"}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
# Tool calling
resp = client.chat.completions.create(
model="gpt-5",
messages=[{"role": "user", "content": "Weather in Tokyo?"}],
tools=[{
"type": "function",
"function": {
"name": "get_weather",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}],
)
call = resp.choices[0].message.tool_calls[0]
# ...run your function, then send the result back:
followup = client.chat.completions.create(
model="gpt-5",
messages=[
{"role": "user", "content": "Weather in Tokyo?"},
resp.choices[0].message,
{"role": "tool", "tool_call_id": call.id, "content": "22°C, clear"},
],
)NinjaChat's v1 API covers chat, images, video, and search. If your pipeline embeds documents or calls a moderation endpoint, keep those calls on your current provider — only the completion traffic needs to move.
Every API key gets 60 requests per minute (video submissions are throttled harder because each one is a long-running job). 429 responses include Retry-After. Need more? Contact us from the console.
You buy a credit pack up front instead of getting a surprise invoice. Failed generations are automatically refunded, and you can set monthly spend limits per account, key, or project.
Chat bills per token — published $/MTok input and output rates for every model on GET /v1/models (cache-read and long-context tiers included). Requests preauthorize an estimated maximum and settle to actual usage; a typical request runs from ≈$0.002 on open models to ≈$0.05 on Claude Opus (gpt-5.4-pro ≈$0.33). No length guard — long context just needs balance.
Fireworks' on-demand deployments, LoRA fine-tuning, and performance-tuned dedicated instances have no equivalent here — NinjaChat is serverless multi-provider inference. Keep custom-model workloads on Fireworks and migrate the standard-model traffic.
Fireworks serves embeddings models; NinjaChat doesn't. Point an embeddings-only client at Fireworks and your completions client at NinjaChat — the OpenAI SDK supports multiple client instances with different base_urls.
For the client, yes: base_url and api_key. Model ids also shrink from accounts/fireworks/models/deepseek-v3p2 to deepseek-v3.2 — a find-and-replace, with the table on this page as the reference.
OpenAI-style response_format json_object and json_schema are enforced server-side. Fireworks' custom grammar mode (GBNF) has no equivalent — express constraints as a JSON schema instead.
NinjaChat routes to upstream providers rather than running its own FireAttention stack, so raw tokens/sec on open models can be lower than a tuned Fireworks deployment. What you gain is one key across open + closed models, image, and video. Benchmark your workload with the free playground before committing.
No — NinjaChat has no fine-tuning or custom-checkpoint hosting. Standard model traffic migrates; custom checkpoints stay on Fireworks.