ninjachat api · benchmarks

Live AI model benchmarks.

Which provider is fastest for a given model? Which is most reliable? We transact text, images and video across every provider — so we can measure it. These leaderboards are computed from real API request traffic, not vendor marketing.

7-day window · as of Mon, 05 Oct 2026 10:47:14 UTC

Provider performance

Speed and reliability, by provider.

Every measured provider across all models it served in the last 7 days. Success counts HTTP 5xx as failures; client errors (4xx) and rate limits are excluded, so one bad caller can't sink a score.

ProviderSuccessTypical p50Typical TTFT
Alibaba100.0%2.48 s1.66 s
Anthropic100.0%5.09 s1.33 s
Baseten100.0%8.10 s17.81 s
Black Forest Labs95.7%11.17 s—
BytePlus99.9%8.65 s11.21 s
Crusoe100.0%210.35 s1.22 s
DeepInfra99.7%9.87 s3.30 s
DigitalOcean100.0%32.87 s7.72 s
Fireworks100.0%3.66 s4.17 s
Google99.0%6.75 s2.61 s
Mistral100.0%30.64 s7.00 s
Morph100.0%31.04 s17.50 s
Novita AI100.0%25.39 s25.36 s
OpenAI99.9%6.12 s1.77 s
Parasail100.0%2.32 s—
Replicate96.9%11.29 s—
SiliconFlow100.0%7.89 s15.38 s
Tencent Cloud100.0%27.36 s7.90 s
Together AI100.0%706 ms600 ms
Voyage AI100.0%262 ms—
Wafer100.0%24.76 s25.62 s
xAI98.5%25.02 s2.05 s
Price / performance

Best value.

Output tokens per second per cent of typical request price — measured throughput against what you actually pay. Chat models with measured performance.

ModelTokens/secPriceValue (tok/s per ¢)
Qwen 3.5 Flash145.2$0.00091613.3
GLM 5.3 Flash136.3$0.00131090.4
DeepSeek V4 Flash 0731141.7$0.0018805.1
Gemini 2.5 Flash292.1$0.0040730.3
Qwen 3.8 Flash79.8$0.0012654.1
MiniMax M3170.6$0.0027631.9
Qwen 3.7 Flash114.4$0.0020564.9
DeepSeek V4 Flash40.1$0.0010409.2
GPT-4.1 Mini128.7$0.0036357.5
Gemini 3 Flash193.7$0.0055352.2

Measured from live API traffic. GET /api/v1/health · Status