Coming from Fireworks AI

Keep Fireworks AI. Add every other lab.

Same OpenAI-compatible mechanics as Fireworks' serverless endpoint — shorter model ids, one prepaid per-token balance, and frontier closed models next to your open ones.

api.fireworks.ai/inference/v1www.ninjachat.ai/api/v1
Drop the deployment path

accounts/fireworks/models/… becomes the slug.

accounts/fireworks/models/deepseek-v3p2deepseek-v3.2
accounts/fireworks/models/qwq-32bqwq-32b
accounts/fireworks/models/kimi-k2-instructkimi-k2.6
before· fireworks ai
from openai import OpenAI

client = OpenAI(
    base_url="https://api.fireworks.ai/inference/v1",
    api_key="fw_...",
)

resp = client.chat.completions.create(
    model="accounts/fireworks/models/deepseek-v3p2",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
after· ninjachat
from openai import OpenAI

client = OpenAI(
    base_url="https://www.ninjachat.ai/api/v1",   # ← changed
    api_key="nj_sk_YOUR_KEY",                    # ← changed
)

resp = client.chat.completions.create(
    model="deepseek-v3.2",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
Create API key API docs ↗
You keep

Your Fireworks AI models still have names.

Same key, new catalog

What joining NinjaChat actually adds.

The rest

Then these.

  1. Create a key and add credits

    Get an nj_sk_ key at /developers and add a credit pack. Balances are prepaid, and failed generations are refunded.

    verify your key works
    # Sanity check: list every model your key can call (no auth needed)
    curl https://www.ninjachat.ai/api/v1/models
    
    # First authenticated request
    curl https://www.ninjachat.ai/api/v1/chat/completions \
      -H "Authorization: Bearer nj_sk_YOUR_KEY" \
      -H "Content-Type: application/json" \
      -d '{"model": "gpt-5", "messages": [{"role": "user", "content": "Say hi"}]}'
  2. Drop the accounts/fireworks/models/ prefix

    Fireworks ids carry the deployment path. NinjaChat slugs are just the model: accounts/fireworks/models/deepseek-v3p2 → deepseek-v3.2. Mapping table below; anything not listed can be checked live via GET /v1/models.

  3. Verify streaming and tool calling

    Same SDK code. Streaming arrives as SSE deltas, and tool_calls round-trip with role: "tool" messages.

    streaming + tools, unchanged sdk
    # Streaming
    stream = client.chat.completions.create(
        model="claude-sonnet-4.6",
        messages=[{"role": "user", "content": "Write a haiku about ninjas"}],
        stream=True,
    )
    for chunk in stream:
        if chunk.choices and chunk.choices[0].delta.content:
            print(chunk.choices[0].delta.content, end="", flush=True)
    
    # Tool calling
    resp = client.chat.completions.create(
        model="gpt-5",
        messages=[{"role": "user", "content": "Weather in Tokyo?"}],
        tools=[{
            "type": "function",
            "function": {
                "name": "get_weather",
                "parameters": {
                    "type": "object",
                    "properties": {"city": {"type": "string"}},
                    "required": ["city"],
                },
            },
        }],
    )
    call = resp.choices[0].message.tool_calls[0]
    # ...run your function, then send the result back:
    followup = client.chat.completions.create(
        model="gpt-5",
        messages=[
            {"role": "user", "content": "Weather in Tokyo?"},
            resp.choices[0].message,
            {"role": "tool", "tool_call_id": call.id, "content": "22°C, clear"},
        ],
    )

Model mapping — Fireworks ids to NinjaChat slugs

Fireworks idNinjaChat slugPrice / request
accounts/fireworks/models/deepseek-v3p2deepseek-v3.2$0.004
accounts/fireworks/models/qwq-32bqwq-32b$0.002
accounts/fireworks/models/kimi-k2-instruct (any K2)kimi-k2.6$0.009
accounts/fireworks/models/glm-4p5 (and newer)glm-5$0.008
accounts/fireworks/models/flux-1-dev-fp8flux-2-pro (POST /v1/images/generations)$0.03 / image
— (closed models)gpt-5, claude-opus-4.6, gemini-3.1-pro…$0.022–$0.05
The honest part

What doesn't come along.

No /v1/embeddings and no moderations endpoint

NinjaChat's v1 API covers chat, images, video, and search. If your pipeline embeds documents or calls a moderation endpoint, keep those calls on your current provider — only the completion traffic needs to move.

60 requests/min default rate limit

Every API key gets 60 requests per minute (video submissions are throttled harder because each one is a long-running job). 429 responses include Retry-After. Need more? Contact us from the console.

Prepaid credits, not postpaid billing

You buy a credit pack up front instead of getting a surprise invoice. Failed generations are automatically refunded, and you can set monthly spend limits per account, key, or project.

Metered $/MTok chat pricing

Chat bills per token — published $/MTok input and output rates for every model on GET /v1/models (cache-read and long-context tiers included). Requests preauthorize an estimated maximum and settle to actual usage; a typical request runs from ≈$0.002 on open models to ≈$0.05 on Claude Opus (gpt-5.4-pro ≈$0.33). No length guard — long context just needs balance.

No dedicated deployments, fine-tuning, or FireAttention tuning

Fireworks' on-demand deployments, LoRA fine-tuning, and performance-tuned dedicated instances have no equivalent here — NinjaChat is serverless multi-provider inference. Keep custom-model workloads on Fireworks and migrate the standard-model traffic.

No embeddings endpoint

Fireworks serves embeddings models; NinjaChat doesn't. Point an embeddings-only client at Fireworks and your completions client at NinjaChat — the OpenAI SDK supports multiple client instances with different base_urls.

FAQ

Is the migration really just two lines?

For the client, yes: base_url and api_key. Model ids also shrink from accounts/fireworks/models/deepseek-v3p2 to deepseek-v3.2 — a find-and-replace, with the table on this page as the reference.

Does grammar/JSON mode survive the move?

OpenAI-style response_format json_object and json_schema are enforced server-side. Fireworks' custom grammar mode (GBNF) has no equivalent — express constraints as a JSON schema instead.

What's the latency story compared to Fireworks?

NinjaChat routes to upstream providers rather than running its own FireAttention stack, so raw tokens/sec on open models can be lower than a tuned Fireworks deployment. What you gain is one key across open + closed models, image, and video. Benchmark your workload with the free playground before committing.

Can I keep fine-tuned models?

No — NinjaChat has no fine-tuning or custom-checkpoint hosting. Standard model traffic migrates; custom checkpoints stay on Fireworks.

Ready for every lab?

Start building All guides →