Coming from fal.ai

Keep fal.ai. Add every other lab.

Swap fal's per-model queue endpoints for two stable ones: synchronous /v1/images/generations and submit-then-poll /v1/videos — with chat models on the same key and prepaid, auto-refunding billing.

fal.runwww.ninjachat.ai/api/v1
One endpoint per model → one endpoint, model field

fal-ai/flux-pro/v2 becomes model: "flux-2-pro".

fal-ai/flux-pro/v2model: "flux-2-pro"
fal-ai/veo3.1/fastmodel: "veo-3.1-fast"
queue.subscribe + pollPOST /v1/images/generations → url
before· fal.ai
import fal_client  # FAL_KEY

result = fal_client.subscribe(
    "fal-ai/flux-pro/v2",
    arguments={"prompt": "a ninja cat, studio lighting"},
)
url = result["images"][0]["url"]
after· ninjachat
# pip install ninjachat
import os
from ninjachat import NinjaChat

client = NinjaChat(api_key=os.environ["NINJACHAT_API_KEY"])
image = client.images.generate(
    model="flux-2-pro",
    prompt="a ninja cat, studio lighting",
    n=1,
)
print(image["data"][0]["url"])
Create API key API docs ↗
You keep

Your fal.ai models still have names.

Kling 2.6
Same key, new catalog

What joining NinjaChat actually adds.

The rest

Then these.

  1. Create a key and add credits

    Get an nj_sk_ key at /developers and add a credit pack. Balances are prepaid, and failed generations are refunded.

    verify your key works
    # Sanity check: list every model your key can call (no auth needed)
    curl https://www.ninjachat.ai/api/v1/models
    
    # First authenticated request
    curl https://www.ninjachat.ai/api/v1/chat/completions \
      -H "Authorization: Bearer nj_sk_YOUR_KEY" \
      -H "Content-Type: application/json" \
      -d '{"model": "gpt-5", "messages": [{"role": "user", "content": "Say hi"}]}'
  2. Video: queue.submit + status → /v1/videos submit + GET /v1/videos/{id}

    fal's queue.submit / queue.status / queue.result triple becomes submit + one status endpoint that carries the result URL when done. Charged on submit, auto-refunded on failure.

    before· fal.ai
    handle = fal_client.submit(
        "fal-ai/veo3.1/fast",
        arguments={"prompt": "drone shot over a neon city at night"},
    )
    # poll:
    status = fal_client.status("fal-ai/veo3.1/fast", handle.request_id)
    result = fal_client.result("fal-ai/veo3.1/fast", handle.request_id)
    video_url = result["video"]["url"]
    after· ninjachat
    # pip install ninjachat
    import os
    from ninjachat import NinjaChat
    
    client = NinjaChat(api_key=os.environ["NINJACHAT_API_KEY"])
    job = client.videos.generate(
        model="veo-3.1-fast",
        prompt="drone shot over a neon city at night",
        duration=8,
        aspect_ratio="16:9",
    )
    video = client.videos.wait_for(job["id"])
    print(video["video_url"])
  3. Optional: webhooks instead of polling

    Register an HTTPS endpoint in the console (https://www.ninjachat.ai/developers/webhooks) or via POST /webhooks. NinjaChat POSTs signed video.completed / video.failed events, so you can drop the polling loop. The signing secret is returned once at registration.

    register a webhook endpoint
    curl https://www.ninjachat.ai/api/v1/webhooks \
      -H "Authorization: Bearer nj_sk_YOUR_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "url": "https://yourapp.com/hooks/ninjachat",
        "events": ["video.completed", "video.failed"]
      }'

Model mapping — fal endpoints to NinjaChat slugs

fal.ai endpointNinjaChat slugPrice
fal-ai/flux-pro/v2flux-2-pro$0.03 / image
fal-ai/flux-pro/kontext/maxflux-kontext-max$0.08 / image
fal-ai/flux-pro/kontextflux-kontext-pro$0.04 / image
fal-ai/nano-banananano-banana$0.02 / image
fal-ai/recraft/v3/text-to-imagerecraft-v3$0.04 / image
fal-ai/imagen4/previewgoogle-imagen-4temporarily unavailable$0.04 / image
fal-ai/bytedance/seedream (v3/v4)seedream$0.04 / image
fal-ai/veo3.1veo-3.1 (POST /v1/videos)$3.20 / video
fal-ai/veo3.1/fastveo-3.1-fast$1.20 / video
fal-ai/kling-video (v2 family)kling-video$1.40 / video
fal-ai/bytedance/seedance (pro)seedance-2$3.64 / video
The honest part

What doesn't come along.

One endpoint per modality, not one per model

fal routes each model through its own endpoint path with its own input schema. NinjaChat has one /v1/images/generations and one /v1/videos with a uniform schema — the model is a request field. Less flexible for exotic per-model knobs; much less code to maintain.

No realtime WebSockets, LoRA training, or private model hosting

fal's realtime inference sockets, trainers, and custom-model deployment have no equivalent. NinjaChat is standard-catalog inference (plus chat/search on the same key).

No /v1/embeddings and no moderations endpoint

NinjaChat's v1 API covers chat, images, video, and search. If your pipeline embeds documents or calls a moderation endpoint, keep those calls on your current provider — only the completion traffic needs to move.

60 requests/min default rate limit

Every API key gets 60 requests per minute (video submissions are throttled harder because each one is a long-running job). 429 responses include Retry-After. Need more? Contact us from the console.

Prepaid credits, not postpaid billing

You buy a credit pack up front instead of getting a surprise invoice. Failed generations are automatically refunded, and you can set monthly spend limits per account, key, or project.

FAQ

fal is billed per megapixel/second — how does NinjaChat bill?

Fixed price per generation, listed on the model mapping above (e.g. flux-2-pro $0.03/image, veo-3.1-fast $1.20/video), charged from a prepaid balance and auto-refunded on failure. Predictable enough to put in a spreadsheet.

Do I lose the queue's progress updates?

Images don't need them — the call is synchronous. For video, GET /v1/videos/{id} returns processing with a progress field until the job completes, or register a webhook at https://www.ninjachat.ai/developers/webhooks and receive video.completed / video.failed pushes.

Can I keep using fal's client libraries?

No — @fal-ai/client and fal_client talk fal's queue protocol. But the replacement is plain HTTP (fetch/requests as shown above), so you're removing a dependency, not adding one.

What if I use fal for a model NinjaChat doesn't serve?

Check the mapping table and GET /v1/models first. Standard models (FLUX, Veo, Kling, Seedance, Imagen, Recraft, Nano Banana) migrate; fal-exclusive community pipelines and custom LoRAs don't — keep those on fal.

Ready for every lab?

Start building All guides →