Keep your open-source models — DeepSeek, Qwen, Kimi, GLM — and add the closed frontier (GPT-5, Claude, Gemini) plus image and video, without changing SDKs.
api.together.xyz/v1www.ninjachat.ai/api/v1deepseek-ai/DeepSeek-V3.2→deepseek-v3.2moonshotai/Kimi-K2-Instruct→kimi-k2.6black-forest-labs/FLUX.1-pro→flux-2-profrom openai import OpenAI
client = OpenAI(
base_url="https://api.together.xyz/v1",
api_key="tgp_...",
)
resp = client.chat.completions.create(
model="deepseek-ai/DeepSeek-V3.2",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)from openai import OpenAI
client = OpenAI(
base_url="https://www.ninjachat.ai/api/v1", # ← changed
api_key="nj_sk_YOUR_KEY", # ← changed
)
resp = client.chat.completions.create(
model="deepseek-v3.2",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)Get an nj_sk_ key at /developers and add a credit pack. Balances are prepaid, and failed generations are refunded.
# Sanity check: list every model your key can call (no auth needed)
curl https://www.ninjachat.ai/api/v1/models
# First authenticated request
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer nj_sk_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-5", "messages": [{"role": "user", "content": "Say hi"}]}'Together uses full Hugging Face-style repo paths. NinjaChat slugs are short and stable — deepseek-ai/DeepSeek-V3.2 becomes deepseek-v3.2, and the slug keeps pointing at the current serving of that model, so a provider-side redeploy never breaks your code. Full table below.
Same SDK code. Streaming arrives as SSE deltas, and tool_calls round-trip with role: "tool" messages.
# Streaming
stream = client.chat.completions.create(
model="claude-sonnet-4.6",
messages=[{"role": "user", "content": "Write a haiku about ninjas"}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
# Tool calling
resp = client.chat.completions.create(
model="gpt-5",
messages=[{"role": "user", "content": "Weather in Tokyo?"}],
tools=[{
"type": "function",
"function": {
"name": "get_weather",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}],
)
call = resp.choices[0].message.tool_calls[0]
# ...run your function, then send the result back:
followup = client.chat.completions.create(
model="gpt-5",
messages=[
{"role": "user", "content": "Weather in Tokyo?"},
resp.choices[0].message,
{"role": "tool", "tool_call_id": call.id, "content": "22°C, clear"},
],
)Together's images.generate endpoint maps to POST /v1/images/generations. Same FLUX family, synchronous response with hosted URLs.
# Before — Together images
img = client.images.generate(
model="black-forest-labs/FLUX.1-pro",
prompt="a ninja cat, studio lighting",
)# pip install ninjachat
import os
from ninjachat import NinjaChat
client = NinjaChat(api_key=os.environ["NINJACHAT_API_KEY"])
image = client.images.generate(
model="flux-2-pro",
prompt="a ninja cat, studio lighting",
n=1,
)
print(image["data"][0]["url"])NinjaChat's v1 API covers chat, images, video, and search. If your pipeline embeds documents or calls a moderation endpoint, keep those calls on your current provider — only the completion traffic needs to move.
Every API key gets 60 requests per minute (video submissions are throttled harder because each one is a long-running job). 429 responses include Retry-After. Need more? Contact us from the console.
You buy a credit pack up front instead of getting a surprise invoice. Failed generations are automatically refunded, and you can set monthly spend limits per account, key, or project.
Chat bills per token — published $/MTok input and output rates for every model on GET /v1/models (cache-read and long-context tiers included). Requests preauthorize an estimated maximum and settle to actual usage; a typical request runs from ≈$0.002 on open models to ≈$0.05 on Claude Opus (gpt-5.4-pro ≈$0.33). No length guard — long context just needs balance.
Together's infra products (dedicated instances, fine-tuning jobs, GPU clusters) have no NinjaChat equivalent — this is a serverless inference API only. If you fine-tune on Together, keep that workload there and route inference wherever it's cheapest.
Together serves hundreds of open checkpoints. NinjaChat serves the popular ones (DeepSeek V3.2, Qwen, Kimi K2.6, and GLM 5). Check GET /v1/models for the live list before migrating a niche model.
The SDK code, yes — it's the same OpenAI-compatible surface (chat.completions, streaming, tools). You change base_url, api_key, and shorten model ids from repo paths to slugs (deepseek-ai/DeepSeek-V3.2 → deepseek-v3.2).
Consolidation: one key and one prepaid balance for open models AND GPT-5/Claude/Gemini AND image/video generation. Metered $/MTok pricing means you pay for the tokens you actually use, with published rates for every model on GET /v1/models.
No — there's no /v1/embeddings or rerank endpoint. Keep those on Together (or another provider) and move the completion traffic; both SDKs coexist fine since each client instance has its own base_url.
POST /v1/images/generations serves the current FLUX family (flux-2-pro $0.03, flux-2-klein $0.014, flux-kontext-max $0.08, and more) plus Imagen 4, Nano Banana, Recraft, and gpt-image-2 — synchronous, returning hosted URLs.