- OpenAI SDK (python + node)
- SSE streaming
- Tool calling
- response_format JSON modes
- models.list
- Idempotency-Key
Keep the OpenAI SDK. Reach every major lab.
Change the base URL and key. GPT stays; Claude, Gemini, image, and video join it.
base_url="https://api.openai.com/v1"+base_url="https://www.ninjachat.ai/api/v1"−api_key=OPENAI_API_KEY+api_key=NINJACHAT_API_KEYclient.chat.completions.create(…)The client stays. The model boundary moves.
chat.completionsA provider failure becomes another attempt.
- 01
Anthropicrate limit429 - 02
OpenAIgpt-5.4200
Move the calls you can. Keep the ones you can’t.
01Create a key and add creditsrequired
Get an nj_sk_ key at /developers and add a credit pack. Balances are prepaid, and failed generations are refunded.
# Sanity check: list every model your key can call (no auth needed)
curl https://www.ninjachat.ai/api/v1/models
# First authenticated request
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer nj_sk_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-5", "messages": [{"role": "user", "content": "Say hi"}]}'02Change the base URL and key — nothing elserequired
Point the OpenAI SDK at https://www.ninjachat.ai/api/v1. Streaming, tools, JSON modes, and models.list keep working.
from openai import OpenAI
client = OpenAI() # key from OPENAI_API_KEY
resp = client.chat.completions.create(
model="gpt-5",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)from openai import OpenAI
client = OpenAI(
base_url="https://www.ninjachat.ai/api/v1", # ← changed
api_key="nj_sk_YOUR_KEY", # ← changed
)
resp = client.chat.completions.create(
model="gpt-5",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)03Verify streaming and tool callingwhen used
Same SDK code. Streaming arrives as SSE deltas, and tool_calls round-trip with role: "tool" messages.
# Streaming
stream = client.chat.completions.create(
model="claude-sonnet-4.6",
messages=[{"role": "user", "content": "Write a haiku about ninjas"}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
# Tool calling
resp = client.chat.completions.create(
model="gpt-5",
messages=[{"role": "user", "content": "Weather in Tokyo?"}],
tools=[{
"type": "function",
"function": {
"name": "get_weather",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}],
)
call = resp.choices[0].message.tool_calls[0]
# ...run your function, then send the result back:
followup = client.chat.completions.create(
model="gpt-5",
messages=[
{"role": "user", "content": "Weather in Tokyo?"},
resp.choices[0].message,
{"role": "tool", "tool_call_id": call.id, "content": "22°C, clear"},
],
)04Move image generation to /v1/images/generationswhen used
If you use OpenAI's Images API, the equivalent here is POST /v1/images/generations — synchronous JSON with hosted URLs instead of base64 by default. gpt-image-2 is available directly, alongside FLUX, Imagen 4, and Nano Banana on the same key.
# Before — OpenAI Images API
img = client.images.generate(
model="gpt-image-1",
prompt="a ninja cat, studio lighting",
)# pip install ninjachat
import os
from ninjachat import NinjaChat
client = NinjaChat(api_key=os.environ["NINJACHAT_API_KEY"])
image = client.images.generate(
model="gpt-image-2",
prompt="a ninja cat, studio lighting",
n=1,
)
print(image["data"][0]["url"])GPT names stay. Other labs become one field.
gpt-image-1 / gpt-image-2gpt-image-2 (POST /v1/images/generations)$0.128 / image— (not on OpenAI)claude-opus-4.6, gemini-3.1-pro, grok-4…$0.009–$0.05Compatibility without pretending the APIs are identical.
01No /v1/embeddings and no moderations endpoint
NinjaChat's v1 API covers chat, images, video, and search. If your pipeline embeds documents or calls a moderation endpoint, keep those calls on your current provider — only the completion traffic needs to move.
0260 requests/min default rate limit
Every API key gets 60 requests per minute (video submissions are throttled harder because each one is a long-running job). 429 responses include Retry-After. Need more? Contact us from the console.
03Prepaid credits, not postpaid billing
You buy a credit pack up front instead of getting a surprise invoice. Failed generations are automatically refunded, and you can set monthly spend limits per account, key, or project.
04Metered $/MTok chat pricing
Chat bills per token — published $/MTok input and output rates for every model on GET /v1/models (cache-read and long-context tiers included). Requests preauthorize an estimated maximum and settle to actual usage; a typical request runs from ≈$0.002 on open models to ≈$0.05 on Claude Opus (gpt-5.4-pro ≈$0.33). No length guard — long context just needs balance.
05Chat Completions and Responses — no Assistants or Realtime API
NinjaChat implements the Chat Completions and Responses surfaces (plus native images/video/search). Code built on the Assistants API, fine-tuning, or audio/realtime endpoints won't port — chat.completions and responses code ports unchanged.
Choose based on what your product needs.
Use it when the OpenAI SDK should reach multiple labs, media models, fallbacks, and one shared balance.
Keep it when Realtime, Assistants, fine-tuning, or first-party OpenAI primitives are core to the product.
OpenAI migration questions.
Do I have to change my openai SDK code to migrate?
No. Change base_url to https://www.ninjachat.ai/api/v1 and api_key to an nj_sk_ key. chat.completions.create, streaming loops, tool-calling round-trips, and models.list all keep working with the official openai package in Python and Node.
Can I still call GPT models after leaving the OpenAI API?
Yes — gpt-5.4, gpt-5.4-pro, gpt-5, gpt-5-nano, and o3-mini are all served under their own names, alongside Claude, Gemini, Grok, and DeepSeek on the same key.
What happens to my JSON mode / structured outputs?
response_format {type: "json_object"} and {type: "json_schema"} are both enforced server-side. One caveat: json_schema combined with stream: true returns a 400 — use non-streaming for strict-schema outputs.
Is there an equivalent of the OpenAI embeddings endpoint?
No. There is no /v1/embeddings — keep embedding calls on OpenAI (or any embeddings provider) and move only completion traffic. Mixing providers is exactly what the OpenAI SDK's base_url parameter is for.
How does billing differ from OpenAI's?
OpenAI meters tokens and invoices you. NinjaChat is prepaid credits with the same per-token billing model — published $/MTok rates for every lab's models on one balance, holds that settle to actual usage, and automatic refunds on failed requests.