Coming from Together AI

Keep Together AI. Add every other lab.

Keep your open-source models — DeepSeek, Qwen, Kimi, GLM — and add the closed frontier (GPT-5, Claude, Gemini) plus image and video, without changing SDKs.

api.together.xyz/v1www.ninjachat.ai/api/v1
Shorten the Hugging Face path

Open models keep working. Closed ones join them.

deepseek-ai/DeepSeek-V3.2deepseek-v3.2
moonshotai/Kimi-K2-Instructkimi-k2.6
black-forest-labs/FLUX.1-proflux-2-pro
before· together ai
from openai import OpenAI

client = OpenAI(
    base_url="https://api.together.xyz/v1",
    api_key="tgp_...",
)

resp = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V3.2",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
after· ninjachat
from openai import OpenAI

client = OpenAI(
    base_url="https://www.ninjachat.ai/api/v1",   # ← changed
    api_key="nj_sk_YOUR_KEY",                    # ← changed
)

resp = client.chat.completions.create(
    model="deepseek-v3.2",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
Create API key API docs ↗
You keep

Your Together AI models still have names.

Same key, new catalog

What joining NinjaChat actually adds.

The rest

Then these.

  1. Create a key and add credits

    Get an nj_sk_ key at /developers and add a credit pack. Balances are prepaid, and failed generations are refunded.

    verify your key works
    # Sanity check: list every model your key can call (no auth needed)
    curl https://www.ninjachat.ai/api/v1/models
    
    # First authenticated request
    curl https://www.ninjachat.ai/api/v1/chat/completions \
      -H "Authorization: Bearer nj_sk_YOUR_KEY" \
      -H "Content-Type: application/json" \
      -d '{"model": "gpt-5", "messages": [{"role": "user", "content": "Say hi"}]}'
  2. Shorten the model ids

    Together uses full Hugging Face-style repo paths. NinjaChat slugs are short and stable — deepseek-ai/DeepSeek-V3.2 becomes deepseek-v3.2, and the slug keeps pointing at the current serving of that model, so a provider-side redeploy never breaks your code. Full table below.

  3. Verify streaming and tool calling

    Same SDK code. Streaming arrives as SSE deltas, and tool_calls round-trip with role: "tool" messages.

    streaming + tools, unchanged sdk
    # Streaming
    stream = client.chat.completions.create(
        model="claude-sonnet-4.6",
        messages=[{"role": "user", "content": "Write a haiku about ninjas"}],
        stream=True,
    )
    for chunk in stream:
        if chunk.choices and chunk.choices[0].delta.content:
            print(chunk.choices[0].delta.content, end="", flush=True)
    
    # Tool calling
    resp = client.chat.completions.create(
        model="gpt-5",
        messages=[{"role": "user", "content": "Weather in Tokyo?"}],
        tools=[{
            "type": "function",
            "function": {
                "name": "get_weather",
                "parameters": {
                    "type": "object",
                    "properties": {"city": {"type": "string"}},
                    "required": ["city"],
                },
            },
        }],
    )
    call = resp.choices[0].message.tool_calls[0]
    # ...run your function, then send the result back:
    followup = client.chat.completions.create(
        model="gpt-5",
        messages=[
            {"role": "user", "content": "Weather in Tokyo?"},
            resp.choices[0].message,
            {"role": "tool", "tool_call_id": call.id, "content": "22°C, clear"},
        ],
    )
  4. Move FLUX image calls to /v1/images/generations

    Together's images.generate endpoint maps to POST /v1/images/generations. Same FLUX family, synchronous response with hosted URLs.

    before· together ai
    # Before — Together images
    img = client.images.generate(
        model="black-forest-labs/FLUX.1-pro",
        prompt="a ninja cat, studio lighting",
    )
    after· ninjachat
    # pip install ninjachat
    import os
    from ninjachat import NinjaChat
    
    client = NinjaChat(api_key=os.environ["NINJACHAT_API_KEY"])
    image = client.images.generate(
        model="flux-2-pro",
        prompt="a ninja cat, studio lighting",
        n=1,
    )
    print(image["data"][0]["url"])

Model mapping — Together ids to NinjaChat slugs

Together AI idNinjaChat slugPrice / request
deepseek-ai/DeepSeek-V3.2deepseek-v3.2$0.004
Qwen/QwQ-32Bqwq-32b$0.002
moonshotai/Kimi-K2-Instruct (any K2)kimi-k2.6$0.009
zai-org/GLM (4.x/5 family)glm-5$0.008
mistralai/Mistral-Large…mistral-large$0.016
black-forest-labs/FLUX.1-proflux-2-pro (POST /v1/images/generations)$0.03 / image
— (closed models Together can't serve)gpt-5, claude-opus-4.6, gemini-3.1-pro…$0.022–$0.05
The honest part

What doesn't come along.

No /v1/embeddings and no moderations endpoint

NinjaChat's v1 API covers chat, images, video, and search. If your pipeline embeds documents or calls a moderation endpoint, keep those calls on your current provider — only the completion traffic needs to move.

60 requests/min default rate limit

Every API key gets 60 requests per minute (video submissions are throttled harder because each one is a long-running job). 429 responses include Retry-After. Need more? Contact us from the console.

Prepaid credits, not postpaid billing

You buy a credit pack up front instead of getting a surprise invoice. Failed generations are automatically refunded, and you can set monthly spend limits per account, key, or project.

Metered $/MTok chat pricing

Chat bills per token — published $/MTok input and output rates for every model on GET /v1/models (cache-read and long-context tiers included). Requests preauthorize an estimated maximum and settle to actual usage; a typical request runs from ≈$0.002 on open models to ≈$0.05 on Claude Opus (gpt-5.4-pro ≈$0.33). No length guard — long context just needs balance.

No dedicated endpoints, fine-tuning, or GPU rental

Together's infra products (dedicated instances, fine-tuning jobs, GPU clusters) have no NinjaChat equivalent — this is a serverless inference API only. If you fine-tune on Together, keep that workload there and route inference wherever it's cheapest.

A curated open-model roster, not every checkpoint

Together serves hundreds of open checkpoints. NinjaChat serves the popular ones (DeepSeek V3.2, Qwen, Kimi K2.6, and GLM 5). Check GET /v1/models for the live list before migrating a niche model.

FAQ

Will my Together AI code work unchanged?

The SDK code, yes — it's the same OpenAI-compatible surface (chat.completions, streaming, tools). You change base_url, api_key, and shorten model ids from repo paths to slugs (deepseek-ai/DeepSeek-V3.2 → deepseek-v3.2).

Why would I move open-source models off Together?

Consolidation: one key and one prepaid balance for open models AND GPT-5/Claude/Gemini AND image/video generation. Metered $/MTok pricing means you pay for the tokens you actually use, with published rates for every model on GET /v1/models.

Does NinjaChat host embeddings or rerankers like Together?

No — there's no /v1/embeddings or rerank endpoint. Keep those on Together (or another provider) and move the completion traffic; both SDKs coexist fine since each client instance has its own base_url.

What about FLUX image generation?

POST /v1/images/generations serves the current FLUX family (flux-2-pro $0.03, flux-2-klein $0.014, flux-kontext-max $0.08, and more) plus Imagen 4, Nano Banana, Recraft, and gpt-image-2 — synchronous, returning hosted URLs.

Ready for every lab?

Start building All guides →