Z.ai's full GLM 4.7 reasoning model for coding and multi-step agent work.
| Provider | Context | Price · in / out | p50 | p95 | Uptime |
|---|---|---|---|---|---|
| bbasetenNo trainingNo training on your data. Zero-retention available | 1M | $2.52 / $7.92/MTok | 2.7s | 2.7s | 100% |
routing.strategyproviders.onlyproviders.excludeproviders.ordermax_cost_usddata_policydocs →Collecting — charts appear after two days.
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NINJACHAT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.2-fast",
"messages": [{ "role": "user", "content": "Hello!" }],
"stream": true
}'Baseten's speed-optimized GLM 5.2 rail for real-time agentic engineering at a full 1M-token context.
GLM 5.2 Fast costs $2.52/M input tokens and $7.92/M output tokens.
GLM 5.2 Fast accepts text and returns text.
POST /api/v1/chat/completions with `"model": "glm-5.2-fast"`.
GLM 4.7, GLM 4.7 Flash, GLM 5.2, GLM-5.
Z.AI's GLM 5.2 agentic model served directly through Fireworks' zero-retention endpoint.
Zhipu AI's GLM-5 with strong bilingual (Chinese/English) capabilities.
Anthropic's highest-capability generally available model for long-running agents.