Swap fal's per-model queue endpoints for two stable ones: synchronous /v1/images/generations and submit-then-poll /v1/videos — with chat models on the same key and prepaid, auto-refunding billing.
fal.runwww.ninjachat.ai/api/v1fal-ai/flux-pro/v2→model: "flux-2-pro"fal-ai/veo3.1/fast→model: "veo-3.1-fast"queue.subscribe + poll→POST /v1/images/generations → urlimport fal_client # FAL_KEY
result = fal_client.subscribe(
"fal-ai/flux-pro/v2",
arguments={"prompt": "a ninja cat, studio lighting"},
)
url = result["images"][0]["url"]# pip install ninjachat
import os
from ninjachat import NinjaChat
client = NinjaChat(api_key=os.environ["NINJACHAT_API_KEY"])
image = client.images.generate(
model="flux-2-pro",
prompt="a ninja cat, studio lighting",
n=1,
)
print(image["data"][0]["url"])Get an nj_sk_ key at /developers and add a credit pack. Balances are prepaid, and failed generations are refunded.
# Sanity check: list every model your key can call (no auth needed)
curl https://www.ninjachat.ai/api/v1/models
# First authenticated request
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer nj_sk_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-5", "messages": [{"role": "user", "content": "Say hi"}]}'fal's queue.submit / queue.status / queue.result triple becomes submit + one status endpoint that carries the result URL when done. Charged on submit, auto-refunded on failure.
handle = fal_client.submit(
"fal-ai/veo3.1/fast",
arguments={"prompt": "drone shot over a neon city at night"},
)
# poll:
status = fal_client.status("fal-ai/veo3.1/fast", handle.request_id)
result = fal_client.result("fal-ai/veo3.1/fast", handle.request_id)
video_url = result["video"]["url"]# pip install ninjachat
import os
from ninjachat import NinjaChat
client = NinjaChat(api_key=os.environ["NINJACHAT_API_KEY"])
job = client.videos.generate(
model="veo-3.1-fast",
prompt="drone shot over a neon city at night",
duration=8,
aspect_ratio="16:9",
)
video = client.videos.wait_for(job["id"])
print(video["video_url"])Register an HTTPS endpoint in the console (https://www.ninjachat.ai/developers/webhooks) or via POST /webhooks. NinjaChat POSTs signed video.completed / video.failed events, so you can drop the polling loop. The signing secret is returned once at registration.
curl https://www.ninjachat.ai/api/v1/webhooks \
-H "Authorization: Bearer nj_sk_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://yourapp.com/hooks/ninjachat",
"events": ["video.completed", "video.failed"]
}'fal routes each model through its own endpoint path with its own input schema. NinjaChat has one /v1/images/generations and one /v1/videos with a uniform schema — the model is a request field. Less flexible for exotic per-model knobs; much less code to maintain.
fal's realtime inference sockets, trainers, and custom-model deployment have no equivalent. NinjaChat is standard-catalog inference (plus chat/search on the same key).
NinjaChat's v1 API covers chat, images, video, and search. If your pipeline embeds documents or calls a moderation endpoint, keep those calls on your current provider — only the completion traffic needs to move.
Every API key gets 60 requests per minute (video submissions are throttled harder because each one is a long-running job). 429 responses include Retry-After. Need more? Contact us from the console.
You buy a credit pack up front instead of getting a surprise invoice. Failed generations are automatically refunded, and you can set monthly spend limits per account, key, or project.
Fixed price per generation, listed on the model mapping above (e.g. flux-2-pro $0.03/image, veo-3.1-fast $1.20/video), charged from a prepaid balance and auto-refunded on failure. Predictable enough to put in a spreadsheet.
Images don't need them — the call is synchronous. For video, GET /v1/videos/{id} returns processing with a progress field until the job completes, or register a webhook at https://www.ninjachat.ai/developers/webhooks and receive video.completed / video.failed pushes.
No — @fal-ai/client and fal_client talk fal's queue protocol. But the replacement is plain HTTP (fetch/requests as shown above), so you're removing a dependency, not adding one.
Check the mapping table and GET /v1/models first. Standard models (FLUX, Veo, Kling, Seedance, Imagen, Recraft, Nano Banana) migrate; fal-exclusive community pipelines and custom LoRAs don't — keep those on fal.