Coming from Replicate

Keep Replicate. Add every other lab.

Trade the predictions lifecycle for two plain endpoints: synchronous /v1/images/generations and submit-then-poll /v1/videos — same FLUX, Imagen, Veo, Kling, and Seedance models, prepaid and auto-refunded on failure.

api.replicate.comwww.ninjachat.ai/api/v1
Predictions → two endpoints

Images are synchronous. Video is submit + poll.

replicate.run(flux-2-pro) + pollPOST /v1/images/generations → url
predictions.create(veo) + loopPOST /v1/videos + GET /videos/{id}
black-forest-labs/flux-2-proflux-2-pro
before· replicate
import replicate  # REPLICATE_API_TOKEN

output = replicate.run(
    "black-forest-labs/flux-2-pro",
    input={"prompt": "a ninja cat, studio lighting"},
)
url = str(output[0]) if isinstance(output, list) else str(output)
after· ninjachat
# pip install ninjachat
import os
from ninjachat import NinjaChat

client = NinjaChat(api_key=os.environ["NINJACHAT_API_KEY"])
image = client.images.generate(
    model="flux-2-pro",
    prompt="a ninja cat, studio lighting",
    n=1,
)
print(image["data"][0]["url"])
Create API key API docs ↗
You keep

Your Replicate models still have names.

Kling 2.6
Same key, new catalog

What joining NinjaChat actually adds.

The rest

Then these.

  1. Create a key and add credits

    Get an nj_sk_ key at /developers and add a credit pack. Balances are prepaid, and failed generations are refunded.

    verify your key works
    # Sanity check: list every model your key can call (no auth needed)
    curl https://www.ninjachat.ai/api/v1/models
    
    # First authenticated request
    curl https://www.ninjachat.ai/api/v1/chat/completions \
      -H "Authorization: Bearer nj_sk_YOUR_KEY" \
      -H "Content-Type: application/json" \
      -d '{"model": "gpt-5", "messages": [{"role": "user", "content": "Say hi"}]}'
  2. Video: predictions + polling → /v1/videos submit + GET /v1/videos/{id}

    Video stays async on both platforms — the shape just gets simpler. POST /v1/videos charges up front and returns an id; GET /v1/videos/{id} reports queued → processing → completed (with video_url) or failed (auto-refunded). duration 4–15s, aspect_ratio 16:9 or 9:16, image_url for image-to-video.

    before· replicate
    prediction = replicate.predictions.create(
        model="google/veo-3.1-fast",
        input={"prompt": "drone shot over a neon city at night"},
    )
    while prediction.status not in ("succeeded", "failed", "canceled"):
        time.sleep(5)
        prediction.reload()
    video_url = prediction.output
    after· ninjachat
    # pip install ninjachat
    import os
    from ninjachat import NinjaChat
    
    client = NinjaChat(api_key=os.environ["NINJACHAT_API_KEY"])
    job = client.videos.generate(
        model="veo-3.1-fast",
        prompt="drone shot over a neon city at night",
        duration=8,
        aspect_ratio="16:9",
    )
    video = client.videos.wait_for(job["id"])
    print(video["video_url"])
  3. Optional: webhooks instead of polling

    Register an HTTPS endpoint in the console (https://www.ninjachat.ai/developers/webhooks) or via POST /webhooks. NinjaChat POSTs signed video.completed / video.failed events, so you can drop the polling loop. The signing secret is returned once at registration.

    register a webhook endpoint
    curl https://www.ninjachat.ai/api/v1/webhooks \
      -H "Authorization: Bearer nj_sk_YOUR_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "url": "https://yourapp.com/hooks/ninjachat",
        "events": ["video.completed", "video.failed"]
      }'

Model mapping — Replicate models to NinjaChat slugs

Fixed price per generation — no GPU-second metering, no cold-start surprises.

Replicate modelNinjaChat slugPrice
black-forest-labs/flux-2-proflux-2-pro$0.03 / image
black-forest-labs/flux-kontext-maxflux-kontext-max$0.08 / image
black-forest-labs/flux-1.1-pro-ultraflux-1-pro-ultra$0.06 / image
google/imagen-4google-imagen-4temporarily unavailable$0.04 / image
recraft-ai/recraft-v3recraft-v3$0.04 / image
bytedance/seedream (3/4)seedream$0.04 / image
google/nano-banananano-banana$0.02 / image
google/veo-3.1veo-3.1 (POST /v1/videos)$3.20 / video
google/veo-3.1-fastveo-3.1-fast$1.20 / video
google/veo-2google-veo-2temporarily unavailable$4.00 / video
kwaivgi/kling-v2 (family)kling-video$1.40 / video
bytedance/seedance (2.0)seedance-2$3.64 / video
The honest part

What doesn't come along.

A fixed catalog, not 10,000 community models

Replicate's superpower is running any public model or your own Cog container. NinjaChat serves a curated set: 39 image, 12 video, and 169 text models. If your pipeline depends on a niche community model or custom weights, that part stays on Replicate.

Fixed per-generation pricing instead of GPU-second metering

You pay a listed price per image/video (e.g. flux-2-pro $0.03, veo-3.1-fast $1.20), charged up front and auto-refunded if generation fails. No hardware classes, no per-second billing, no paying for cold starts.

No /v1/embeddings and no moderations endpoint

NinjaChat's v1 API covers chat, images, video, and search. If your pipeline embeds documents or calls a moderation endpoint, keep those calls on your current provider — only the completion traffic needs to move.

60 requests/min default rate limit

Every API key gets 60 requests per minute (video submissions are throttled harder because each one is a long-running job). 429 responses include Retry-After. Need more? Contact us from the console.

Prepaid credits, not postpaid billing

You buy a credit pack up front instead of getting a surprise invoice. Failed generations are automatically refunded, and you can set monthly spend limits per account, key, or project.

No training, no custom weights, no deployments

There is no equivalent of Replicate trainings, LoRA hosting, or dedicated deployments. This is inference on standard models only.

FAQ

Do I still need a polling loop for images?

No — POST /v1/images/generations is synchronous and returns hosted URLs in the response body. Only video keeps the async submit + poll pattern, or register a webhook at https://www.ninjachat.ai/developers/webhooks to skip polling.

What happens if a generation fails after I've been charged?

It's refunded automatically — for images within the same request, for video via the status endpoint or the sweeper when the job fails asynchronously. The error responses say "You were not charged" and mean it.

Can I run my own model or custom weights like on Replicate?

No. NinjaChat serves a fixed catalog of standard models. Custom Cog containers, fine-tuned weights, and trainings are Replicate features with no equivalent here — migrate the standard-model traffic and keep custom workloads where they are.

How do image-to-image and image-to-video work?

Pass a public image URL: image on /v1/images/generations (FLUX Kontext models are built for editing), image_url on /v1/videos. Seedance 2 additionally accepts up to 4 reference_images plus reference_video. Uploaded reference audio is not currently supported; use generate_audio for a generated soundtrack instead.

Ready for every lab?

Start building All guides →