Which provider is fastest for a given model? Which is most reliable? We transact text, images and video across every provider — so we can measure it. These leaderboards are computed from real API request traffic, not vendor marketing.
7-day window · as of Mon, 05 Oct 2026 10:47:14 UTC
Every measured provider across all models it served in the last 7 days. Success counts HTTP 5xx as failures; client errors (4xx) and rate limits are excluded, so one bad caller can't sink a score.
| Provider | Success | Typical p50 | Typical TTFT |
|---|---|---|---|
| Alibaba | 100.0% | 2.48 s | 1.66 s |
| Anthropic | 100.0% | 5.09 s | 1.33 s |
| Baseten | 100.0% | 8.10 s | 17.81 s |
| Black Forest Labs | 95.7% | 11.17 s | — |
| BytePlus | 99.9% | 8.65 s | 11.21 s |
| Crusoe | 100.0% | 210.35 s | 1.22 s |
| DeepInfra | 99.7% | 9.87 s | 3.30 s |
| DigitalOcean | 100.0% | 32.87 s | 7.72 s |
| Fireworks | 100.0% | 3.66 s | 4.17 s |
| 99.0% | 6.75 s | 2.61 s | |
| Mistral | 100.0% | 30.64 s | 7.00 s |
| Morph | 100.0% | 31.04 s | 17.50 s |
| Novita AI | 100.0% | 25.39 s | 25.36 s |
| OpenAI | 99.9% | 6.12 s | 1.77 s |
| Parasail | 100.0% | 2.32 s | — |
| Replicate | 96.9% | 11.29 s | — |
| SiliconFlow | 100.0% | 7.89 s | 15.38 s |
| Tencent Cloud | 100.0% | 27.36 s | 7.90 s |
| Together AI | 100.0% | 706 ms | 600 ms |
| Voyage AI | 100.0% | 262 ms | — |
| Wafer | 100.0% | 24.76 s | 25.62 s |
| xAI | 98.5% | 25.02 s | 2.05 s |
For each model, the provider currently serving it fastest. Open a model for the full provider comparison and its latency, uptime and price history.
Output tokens per second per cent of typical request price — measured throughput against what you actually pay. Chat models with measured performance.
| Model | Tokens/sec | Price | Value (tok/s per ¢) |
|---|---|---|---|
| Qwen 3.5 Flash | 145.2 | $0.0009 | 1613.3 |
| GLM 5.3 Flash | 136.3 | $0.0013 | 1090.4 |
| DeepSeek V4 Flash 0731 | 141.7 | $0.0018 | 805.1 |
| Gemini 2.5 Flash | 292.1 | $0.0040 | 730.3 |
| Qwen 3.8 Flash | 79.8 | $0.0012 | 654.1 |
| MiniMax M3 | 170.6 | $0.0027 | 631.9 |
| Qwen 3.7 Flash | 114.4 | $0.0020 | 564.9 |
| DeepSeek V4 Flash | 40.1 | $0.0010 | 409.2 |
| GPT-4.1 Mini | 128.7 | $0.0036 | 357.5 |
| Gemini 3 Flash | 193.7 | $0.0055 | 352.2 |
Measured from live API traffic. GET /api/v1/health · Status