Thinking Machines Lab's 975B multimodal open-weight generalist with 1M context.
Thinking Machines Lab's efficient 276B/12B-active multimodal reasoner with a 524K context window.
| Provider | Role | Context | Price · in / out |
|---|---|---|---|
| ddeepinfraNo trainingNo training on your data | Primary | 524K | $0.60 / $1.44/MToksame price, any rail |
| bbasetenNo trainingNo training on your data. Zero-retention available | Fallback | 524K |
routing.strategyproviders.onlyCollecting — charts appear after two days.
curl https://www.ninjachat.ai/api/v1/chat/completions \
-H "Authorization: Bearer $NINJACHAT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "inkling-small",
"messages": [{ "role": "user", "content": "Hello!" }],
"stream": true
}'Thinking Machines Lab's efficient 276B/12B-active multimodal reasoner with a 524K context window.
Inkling Small costs $0.60/M input tokens and $1.44/M output tokens.
Inkling Small accepts text and image and returns text.
POST /api/v1/chat/completions with `"model": "inkling-small"`.
Inkling, Kling.
providers.excludeproviders.ordermax_cost_usddata_policyAnthropic's highest-capability generally available model for long-running agents.
Anthropic's fastest Claude variant — lightning-quick responses.
Anthropic's most intelligent model — use when quality is paramount.