Providers

SiliconFlow

Verified

An inference platform serving open text and media models at scale.

Models on SiliconFlow

Open in catalog

6 of 6

GLM 5.3 FlashChatglm-5.3-flash1M contextRouteprimary$0.15 / $0.50per M input / output
LongCat 2.0Chatlongcat-2.01M contextRouteprimary$0.75 / $2.95per M input / output
Nex N2 ProChatnex-n2-pro262K contextRouteprimary$0.50 / $2.50per M input / output
Qwen 3.5 122B A10BChatqwen-3.5-122b-a10b262K contextRoutefallback$0.29 / $2.40per M input / output
Qwen 3.5 27BChatqwen-3.5-27b262K contextRoutefallback$0.26 / $2.60per M input / output
Qwen 3.5 9BChatqwen-3.5-9b262K contextRoutefallback$0.10 / $0.15per M input / output

Performance

Request success
100%
Time to first token
10.4s
Generation speed
61.2tok/s

Response latency estimates

Time to first token

Generation speed · tok/s

Request success

Text-model timing and speed, plus API request success. Daily latency percentiles are estimates.

Direct routingCall SiliconFlow explicitly
siliconflow· routing
# pin every request to SiliconFlow
curl https://www.ninjachat.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NINJACHAT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "messages": [{ "role": "user", "content": "Hello!" }],
    "routing": { "providers": { "only": ["siliconflow"] } }
  }'