Z.ai's full GLM 4.7 reasoning model for coding and multi-step agent work.
| Provider | Context | Price · in / out | p50 | p95 | Uptime |
|---|---|---|---|---|---|
| ddeepinfraNo trainingNo training on your data | 203K | $0.21 / $0.55/MTok | 19.1s | 66.5s | 100% |
routing.strategyproviders.onlyproviders.excludeproviders.ordermax_cost_usddata_policydocs →Collecting — charts appear after two days.
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NINJACHAT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-4.7-flash",
"messages": [{ "role": "user", "content": "Hello!" }],
"stream": true
}'Z.ai's extremely low-cost GLM reasoning model for fast coding and tool workflows.
GLM 4.7 Flash costs $0.21/M input tokens and $0.55/M output tokens.
GLM 4.7 Flash accepts text and returns text.
POST /api/v1/chat/completions with `"model": "glm-4.7-flash"`.
GLM 4.7, GLM 5.2, GLM-5.
Zhipu AI's GLM-5 with strong bilingual (Chinese/English) capabilities.
Anthropic's highest-capability generally available model for long-running agents.
Anthropic's fastest Claude variant — lightning-quick responses.