Z.ai's full GLM 4.7 reasoning model for coding and multi-step agent work.
| Provider | Role | Context | Price · in / out | p50 | p95 | First token | Uptime | tok/s |
|---|---|---|---|---|---|---|---|---|
| ddigitaloceanNo trainingNo training on your data. Zero-retention available | Primary | 164K | $1.68 / $5.28/MToksame price, any rail | 2.9s | 4.4s | — | 100% | — |
| ddeepinfraNo trainingNo training on your data |
Collecting — charts appear after two days.
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NINJACHAT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.1",
"messages": [{ "role": "user", "content": "Hello!" }],
"stream": true
}'Z.AI's agentic GLM 5.1 model with structured output, reasoning, and function-calling support.
GLM 5.1 costs $1.68/M input tokens and $5.28/M output tokens.
GLM 5.1 accepts text and returns text.
POST /api/v1/chat/completions with `"model": "glm-5.1"`.
GLM 4.7, GLM 4.7 Flash, GLM 5.2, GLM 5.2 Fast, GLM-5.
| 203K |
| 9.4s |
| 15.2s |
| 9.1s |
| 100% |
| 264 |
| ggmicloudNo trainingNo training on your data | Emergency | 203K | 2.0s | 13.5s | 1.9s | 100% | 76 |
routing.strategyproviders.onlyproviders.excludeproviders.ordermax_cost_usddata_policydocs →Z.AI's GLM 5.2 agentic model served directly through Fireworks' zero-retention endpoint.
Baseten's speed-optimized GLM 5.2 rail for real-time agentic engineering at a full 1M-token context.
Zhipu AI's GLM-5 with strong bilingual (Chinese/English) capabilities.