Coming from Anthropic

Keep Anthropic. Add every other lab.

Keep Claude — Opus 4.6, Sonnet 4.6, Haiku 4.5 — but call it through the OpenAI-compatible shape, next to GPT and Gemini on one prepaid key.

api.anthropic.comwww.ninjachat.ai/api/v1
Messages API → Chat Completions

Claude stays. The shape becomes OpenAI.

system="…"{ role: "system" }
content[0].textmessage.content
tool_usetool_calls
before· anthropic messages api
import anthropic

client = anthropic.Anthropic()  # ANTHROPIC_API_KEY

msg = client.messages.create(
    model="claude-opus-4-6",
    max_tokens=1024,                       # required
    system="You are a terse assistant.",  # top-level param
    messages=[{"role": "user", "content": "Hello!"}],
)
print(msg.content[0].text)                # content block list
after· ninjachat chat completions
# pip install ninjachat
import os
from ninjachat import NinjaChat

client = NinjaChat(api_key=os.environ["NINJACHAT_API_KEY"])

resp = client.chat.completions.create(
    model="claude-opus-4.6",
    max_completion_tokens=1024,                      # optional (default 2048)
    messages=[
        {"role": "system", "content": "You are a terse assistant."},
        {"role": "user", "content": "Hello!"},
    ],
)
print(resp["choices"][0]["message"]["content"])
Create API key API docs ↗
You keep

Your Anthropic models still have names.

Same key, new catalog

What joining NinjaChat actually adds.

The rest

Then these.

  1. Create a key and add credits

    Get an nj_sk_ key at /developers and add a credit pack. Balances are prepaid, and failed generations are refunded.

    verify your key works
    # Sanity check: list every model your key can call (no auth needed)
    curl https://www.ninjachat.ai/api/v1/models
    
    # First authenticated request
    curl https://www.ninjachat.ai/api/v1/chat/completions \
      -H "Authorization: Bearer nj_sk_YOUR_KEY" \
      -H "Content-Type: application/json" \
      -d '{"model": "gpt-5", "messages": [{"role": "user", "content": "Say hi"}]}'
  2. Translate tool use: tool_use / tool_result → tool_calls / role:"tool"

    Anthropic interleaves tool_use and tool_result content blocks inside messages. Chat Completions puts calls on the assistant message (tool_calls, with JSON-string arguments) and results in a dedicated role:"tool" message bound by tool_call_id. Tool schemas also move from input_schema to function.parameters.

    before· anthropic
    tools = [{
        "name": "get_weather",
        "description": "Get weather for a city",
        "input_schema": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    }]
    
    msg = client.messages.create(
        model="claude-opus-4-6", max_tokens=1024,
        tools=tools,
        messages=[{"role": "user", "content": "Weather in Tokyo?"}],
    )
    tool_use = next(b for b in msg.content if b.type == "tool_use")
    
    # send the result back as a tool_result content block
    followup = client.messages.create(
        model="claude-opus-4-6", max_tokens=1024, tools=tools,
        messages=[
            {"role": "user", "content": "Weather in Tokyo?"},
            {"role": "assistant", "content": msg.content},
            {"role": "user", "content": [{
                "type": "tool_result",
                "tool_use_id": tool_use.id,
                "content": "22°C, clear",
            }]},
        ],
    )
    after· ninjachat
    tools = [{
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get weather for a city",
            "parameters": {                       # was: input_schema
                "type": "object",
                "properties": {"city": {"type": "string"}},
                "required": ["city"],
            },
        },
    }]
    
    resp = client.chat.completions.create(
        model="claude-opus-4.6",
        tools=tools,
        messages=[{"role": "user", "content": "Weather in Tokyo?"}],
    )
    message = resp["choices"][0]["message"]
    call = message["tool_calls"][0]  # arguments is a JSON string
    
    followup = client.chat.completions.create(
        model="claude-opus-4.6", tools=tools,
        messages=[
            {"role": "user", "content": "Weather in Tokyo?"},
            message,                              # assistant msg w/ tool_calls
            {"role": "tool",                      # was: tool_result block
             "tool_call_id": call["id"],
             "content": "22°C, clear"},
        ],
    )
  3. Streaming: event stream → chat.completion.chunk deltas

    Anthropic streams typed events (content_block_delta, message_delta). Chat Completions streams uniform chunks whose choices[0].delta.content carries the text — most code gets simpler.

    streaming claude with the ninjachat sdk
    stream = client.chat.completions.create(
        model="claude-sonnet-4.6",
        messages=[{"role": "user", "content": "Explain WebSockets in 3 lines"}],
        stream=True,
    )
    for chunk in stream:
        delta = chunk.get("choices", [{}])[0].get("delta", {}).get("content")
        if delta:
            print(delta, end="", flush=True)

Request translation cheat sheet

Anthropic Messages APINinjaChat Chat Completions
tools[].input_schematools[].function.parameters
tool_result content block (user){role: "tool", tool_call_id, content} message
stop_reason: "end_turn" / "tool_use" / "max_tokens"finish_reason: "stop" / "tool_calls" / "length"
x-api-key + anthropic-version headersAuthorization: Bearer nj_sk_... (no version header)

Model mapping — Anthropic ids to NinjaChat slugs

Anthropic modelNinjaChat slugPrice / request
claude-opus-4-6claude-opus-4.6$0.050
claude-sonnet-4-6claude-sonnet-4.6$0.030
claude-sonnet-4-5claude-sonnet-4.5$0.030
claude-haiku-4-5claude-haiku-4.5$0.010
The honest part

What doesn't come along.

No /v1/embeddings and no moderations endpoint

NinjaChat's v1 API covers chat, images, video, and search. If your pipeline embeds documents or calls a moderation endpoint, keep those calls on your current provider — only the completion traffic needs to move.

60 requests/min default rate limit

Every API key gets 60 requests per minute (video submissions are throttled harder because each one is a long-running job). 429 responses include Retry-After. Need more? Contact us from the console.

Prepaid credits, not postpaid billing

You buy a credit pack up front instead of getting a surprise invoice. Failed generations are automatically refunded, and you can set monthly spend limits per account, key, or project.

Metered $/MTok chat pricing

Chat bills per token — published $/MTok input and output rates for every model on GET /v1/models (cache-read and long-context tiers included). Requests preauthorize an estimated maximum and settle to actual usage; a typical request runs from ≈$0.002 on open models to ≈$0.05 on Claude Opus (gpt-5.4-pro ≈$0.33). No length guard — long context just needs balance.

No extended-thinking controls or Anthropic-native beta features

The thinking parameter, Anthropic-side prompt caching directives, the Files/Batches APIs, and computer-use betas are not exposed. You get the Claude models through the standard Chat Completions surface — reasoning still happens, you just don't tune its budget. Prompt caching still happens automatically through routing.caching, and cache reads are billed at the discounted cache-read rate.

Different SDK, same models

This is the one guide here where the client library changes (anthropic → openai). Budget an hour, not a sprint: the mapping is mechanical and the cheat-sheet table above covers every field the Messages API uses in normal operation.

FAQ

Which Claude models are available?

claude-opus-4.6 ($0.050/request), claude-sonnet-4.6 ($0.030), claude-sonnet-4.5 ($0.030), and claude-haiku-4.5 ($0.010) — typical-request estimates at metered per-token rates; the current production line-up, with 200K context.

Where does my system prompt go without the system parameter?

Send it as the first message: {role: "system", content: "..."}. It's translated to the model's native system position server-side — behavior is equivalent.

Does Claude tool calling actually round-trip through the OpenAI shape?

Yes — assistant tool_calls and role:"tool" results are translated to Anthropic's native tool_use/tool_result blocks upstream, including recovering tool names across the round-trip. Multi-step tool loops work.

Can I control Claude's extended thinking budget?

No — the Messages API's thinking parameter isn't exposed through the Chat Completions surface. If per-request thinking control is load-bearing for you, keep those calls on Anthropic directly.

Why route Claude through NinjaChat instead of Anthropic directly?

One key and one prepaid balance across chat, image, and video; published per-token rates; ordered model fallbacks; and no per-provider billing setup.

Ready for every lab?

Start building All guides →