Which provider is fastest for a given model? Which is most reliable? We transact text, images and video across every provider — so we can measure it. These leaderboards are computed from real API request traffic, not vendor marketing.
7-day window · as of Fri, 21 Aug 2026 13:14:58 UTC
Every provider ranked across all models it served in the last 7 days. Success counts HTTP 5xx as failures; client errors (4xx) and rate limits are excluded, so one bad caller can't sink a score.
For each model, the provider currently serving it fastest. Open a model for the full provider comparison and its latency, uptime and price history.
Ranked by request volume across the network. Each links to its own benchmark page.
| Model | Provider | Requests | Success | p50 | Price | |
|---|---|---|---|---|---|---|
| seedream | — | 99 | 100.0% | 25.89 s | $0.080 | |
| Gemini 3 Flash | 70 | 100.0% | 4.51 s | $0.0066 | ||
| Uncensored AI | NinjaChat | 37 | 100.0% | 56 ms | $0.0045 | |
| Claude Sonnet 4.6 | Anthropic | 36 | 100.0% | 6.89 s | $0.034 | |
| Gemini 3.1 Pro | 31 | 100.0% | 6.40 s | $0.025 | ||
| Llama 4 Maverickcollecting | Meta | 24 | — | — | $0.0031 | |
| GPT-5collecting | OpenAI | 22 | — | — | $0.025 | |
| Claude Sonnet 4.5collecting | Anthropic | 13 | — | — | $0.034 | |
| nano-bananacollecting | — | 10 | — | — | $0.030 | |
| HY3collecting | Tencent | 8 | — | — | $0.0015 | |
| Claude Opus 4.6collecting | Anthropic | 6 | — | — | $0.056 | |
| gpt-4.1-minicollecting | — | 6 | — | — | — | |
| gpt-5,claude-sonnet-4.6,gemini-3.1-pro,deepseek-v3,gemini-3-flashcollecting | — | 5 | — | — | — | |
| Gemini 2.5 Flashcollecting | 4 | — | — | $0.0051 | ||
| GPT-5 Minicollecting | OpenAI | 3 | — | — | $0.0042 | |
| GPT-5.4collecting | OpenAI | 3 | — | — | $0.031 | |
| batchcollecting | — | 2 | — | — | — | |
| Claude Haiku 4.5collecting | Anthropic | 2 | — | — | $0.011 | |
| Gemini 2.5 Procollecting | 2 | — | — | $0.018 | ||
| Gemini 3 Procollecting | 2 | — | — | $0.025 |
Measured from real NinjaChat API request logs — the same records that bill you. Latency percentiles are computed over successful responses; reliability counts HTTP 5xx (ours or a provider outage) as failures and excludes 4xx and rate limits. Throughput is output tokens over the post-first-token generation window. A model or provider with fewer than 30 requests in the window is marked “collecting” and never ranked. We publish what we measure — no synthetic uptime.
Programmatic access: GET /api/v1/health for live per-model status, or the status page.