Trade the predictions lifecycle for two plain endpoints: synchronous /v1/images/generations and submit-then-poll /v1/videos — same FLUX, Imagen, Veo, Kling, and Seedance models, prepaid and auto-refunded on failure.
api.replicate.comwww.ninjachat.ai/api/v1replicate.run(flux-2-pro) + poll→POST /v1/images/generations → urlpredictions.create(veo) + loop→POST /v1/videos + GET /videos/{id}black-forest-labs/flux-2-pro→flux-2-proimport replicate # REPLICATE_API_TOKEN
output = replicate.run(
"black-forest-labs/flux-2-pro",
input={"prompt": "a ninja cat, studio lighting"},
)
url = str(output[0]) if isinstance(output, list) else str(output)# pip install ninjachat
import os
from ninjachat import NinjaChat
client = NinjaChat(api_key=os.environ["NINJACHAT_API_KEY"])
image = client.images.generate(
model="flux-2-pro",
prompt="a ninja cat, studio lighting",
n=1,
)
print(image["data"][0]["url"])Get an nj_sk_ key at /developers and add a credit pack. Balances are prepaid, and failed generations are refunded.
# Sanity check: list every model your key can call (no auth needed)
curl https://www.ninjachat.ai/api/v1/models
# First authenticated request
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer nj_sk_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-5", "messages": [{"role": "user", "content": "Say hi"}]}'Video stays async on both platforms — the shape just gets simpler. POST /v1/videos charges up front and returns an id; GET /v1/videos/{id} reports queued → processing → completed (with video_url) or failed (auto-refunded). duration 4–15s, aspect_ratio 16:9 or 9:16, image_url for image-to-video.
prediction = replicate.predictions.create(
model="google/veo-3.1-fast",
input={"prompt": "drone shot over a neon city at night"},
)
while prediction.status not in ("succeeded", "failed", "canceled"):
time.sleep(5)
prediction.reload()
video_url = prediction.output# pip install ninjachat
import os
from ninjachat import NinjaChat
client = NinjaChat(api_key=os.environ["NINJACHAT_API_KEY"])
job = client.videos.generate(
model="veo-3.1-fast",
prompt="drone shot over a neon city at night",
duration=8,
aspect_ratio="16:9",
)
video = client.videos.wait_for(job["id"])
print(video["video_url"])Register an HTTPS endpoint in the console (https://www.ninjachat.ai/developers/webhooks) or via POST /webhooks. NinjaChat POSTs signed video.completed / video.failed events, so you can drop the polling loop. The signing secret is returned once at registration.
curl https://www.ninjachat.ai/api/v1/webhooks \
-H "Authorization: Bearer nj_sk_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://yourapp.com/hooks/ninjachat",
"events": ["video.completed", "video.failed"]
}'Fixed price per generation — no GPU-second metering, no cold-start surprises.
Replicate's superpower is running any public model or your own Cog container. NinjaChat serves a curated set: 39 image, 12 video, and 169 text models. If your pipeline depends on a niche community model or custom weights, that part stays on Replicate.
You pay a listed price per image/video (e.g. flux-2-pro $0.03, veo-3.1-fast $1.20), charged up front and auto-refunded if generation fails. No hardware classes, no per-second billing, no paying for cold starts.
NinjaChat's v1 API covers chat, images, video, and search. If your pipeline embeds documents or calls a moderation endpoint, keep those calls on your current provider — only the completion traffic needs to move.
Every API key gets 60 requests per minute (video submissions are throttled harder because each one is a long-running job). 429 responses include Retry-After. Need more? Contact us from the console.
You buy a credit pack up front instead of getting a surprise invoice. Failed generations are automatically refunded, and you can set monthly spend limits per account, key, or project.
There is no equivalent of Replicate trainings, LoRA hosting, or dedicated deployments. This is inference on standard models only.
No — POST /v1/images/generations is synchronous and returns hosted URLs in the response body. Only video keeps the async submit + poll pattern, or register a webhook at https://www.ninjachat.ai/developers/webhooks to skip polling.
It's refunded automatically — for images within the same request, for video via the status endpoint or the sweeper when the job fails asynchronously. The error responses say "You were not charged" and mean it.
No. NinjaChat serves a fixed catalog of standard models. Custom Cog containers, fine-tuned weights, and trainings are Replicate features with no equivalent here — migrate the standard-model traffic and keep custom workloads where they are.
Pass a public image URL: image on /v1/images/generations (FLUX Kontext models are built for editing), image_url on /v1/videos. Seedance 2 additionally accepts up to 4 reference_images plus reference_video. Uploaded reference audio is not currently supported; use generate_audio for a generated soundtrack instead.