Keep Claude — Opus 4.6, Sonnet 4.6, Haiku 4.5 — but call it through the OpenAI-compatible shape, next to GPT and Gemini on one prepaid key.
api.anthropic.comwww.ninjachat.ai/api/v1system="…"→{ role: "system" }content[0].text→message.contenttool_use→tool_callsimport anthropic
client = anthropic.Anthropic() # ANTHROPIC_API_KEY
msg = client.messages.create(
model="claude-opus-4-6",
max_tokens=1024, # required
system="You are a terse assistant.", # top-level param
messages=[{"role": "user", "content": "Hello!"}],
)
print(msg.content[0].text) # content block list# pip install ninjachat
import os
from ninjachat import NinjaChat
client = NinjaChat(api_key=os.environ["NINJACHAT_API_KEY"])
resp = client.chat.completions.create(
model="claude-opus-4.6",
max_completion_tokens=1024, # optional (default 2048)
messages=[
{"role": "system", "content": "You are a terse assistant."},
{"role": "user", "content": "Hello!"},
],
)
print(resp["choices"][0]["message"]["content"])Get an nj_sk_ key at /developers and add a credit pack. Balances are prepaid, and failed generations are refunded.
# Sanity check: list every model your key can call (no auth needed)
curl https://www.ninjachat.ai/api/v1/models
# First authenticated request
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer nj_sk_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-5", "messages": [{"role": "user", "content": "Say hi"}]}'Anthropic interleaves tool_use and tool_result content blocks inside messages. Chat Completions puts calls on the assistant message (tool_calls, with JSON-string arguments) and results in a dedicated role:"tool" message bound by tool_call_id. Tool schemas also move from input_schema to function.parameters.
tools = [{
"name": "get_weather",
"description": "Get weather for a city",
"input_schema": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
}]
msg = client.messages.create(
model="claude-opus-4-6", max_tokens=1024,
tools=tools,
messages=[{"role": "user", "content": "Weather in Tokyo?"}],
)
tool_use = next(b for b in msg.content if b.type == "tool_use")
# send the result back as a tool_result content block
followup = client.messages.create(
model="claude-opus-4-6", max_tokens=1024, tools=tools,
messages=[
{"role": "user", "content": "Weather in Tokyo?"},
{"role": "assistant", "content": msg.content},
{"role": "user", "content": [{
"type": "tool_result",
"tool_use_id": tool_use.id,
"content": "22°C, clear",
}]},
],
)tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get weather for a city",
"parameters": { # was: input_schema
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
resp = client.chat.completions.create(
model="claude-opus-4.6",
tools=tools,
messages=[{"role": "user", "content": "Weather in Tokyo?"}],
)
message = resp["choices"][0]["message"]
call = message["tool_calls"][0] # arguments is a JSON string
followup = client.chat.completions.create(
model="claude-opus-4.6", tools=tools,
messages=[
{"role": "user", "content": "Weather in Tokyo?"},
message, # assistant msg w/ tool_calls
{"role": "tool", # was: tool_result block
"tool_call_id": call["id"],
"content": "22°C, clear"},
],
)Anthropic streams typed events (content_block_delta, message_delta). Chat Completions streams uniform chunks whose choices[0].delta.content carries the text — most code gets simpler.
stream = client.chat.completions.create(
model="claude-sonnet-4.6",
messages=[{"role": "user", "content": "Explain WebSockets in 3 lines"}],
stream=True,
)
for chunk in stream:
delta = chunk.get("choices", [{}])[0].get("delta", {}).get("content")
if delta:
print(delta, end="", flush=True)x-api-key + anthropic-version headersAuthorization: Bearer nj_sk_... (no version header)NinjaChat's v1 API covers chat, images, video, and search. If your pipeline embeds documents or calls a moderation endpoint, keep those calls on your current provider — only the completion traffic needs to move.
Every API key gets 60 requests per minute (video submissions are throttled harder because each one is a long-running job). 429 responses include Retry-After. Need more? Contact us from the console.
You buy a credit pack up front instead of getting a surprise invoice. Failed generations are automatically refunded, and you can set monthly spend limits per account, key, or project.
Chat bills per token — published $/MTok input and output rates for every model on GET /v1/models (cache-read and long-context tiers included). Requests preauthorize an estimated maximum and settle to actual usage; a typical request runs from ≈$0.002 on open models to ≈$0.05 on Claude Opus (gpt-5.4-pro ≈$0.33). No length guard — long context just needs balance.
The thinking parameter, Anthropic-side prompt caching directives, the Files/Batches APIs, and computer-use betas are not exposed. You get the Claude models through the standard Chat Completions surface — reasoning still happens, you just don't tune its budget. Prompt caching still happens automatically through routing.caching, and cache reads are billed at the discounted cache-read rate.
This is the one guide here where the client library changes (anthropic → openai). Budget an hour, not a sprint: the mapping is mechanical and the cheat-sheet table above covers every field the Messages API uses in normal operation.
claude-opus-4.6 ($0.050/request), claude-sonnet-4.6 ($0.030), claude-sonnet-4.5 ($0.030), and claude-haiku-4.5 ($0.010) — typical-request estimates at metered per-token rates; the current production line-up, with 200K context.
Send it as the first message: {role: "system", content: "..."}. It's translated to the model's native system position server-side — behavior is equivalent.
Yes — assistant tool_calls and role:"tool" results are translated to Anthropic's native tool_use/tool_result blocks upstream, including recovering tool names across the round-trip. Multi-step tool loops work.
No — the Messages API's thinking parameter isn't exposed through the Chat Completions surface. If per-request thinking control is load-bearing for you, keep those calls on Anthropic directly.
One key and one prepaid balance across chat, image, and video; published per-token rates; ordered model fallbacks; and no per-provider billing setup.