StepFun
StepFun's efficiency model — Apache 2.0, fast enough to call in a loop and capable enough to trust.
262.144K tokens
कॉन्टेक्स्ट
131,072 tokens
अधिकतम आउटपुट
तेज़
गति
Loops where latency and price compound quickly.
Fast answers on real code without the frontier price.
Reading and filtering a lot of material quickly.
Clean JSON when the answer feeds a pipeline.
Tuned to this model — click any line to copy.
Step 3.7 Flash is StepFun's efficiency model, released on 28 May 2026 under an Apache 2.0 licence. It is a sparse mixture-of-experts design of about 198 billion parameters that activates only around 11 billion per token, which is the whole point: it answers quickly and costs little, while staying capable enough for real work.
StepFun built it for coding agents and search-style workflows — the kind of job where a model is called many times in a loop and both latency and price compound. Throughput runs to several hundred tokens a second on a good rail, and the model exposes selectable reasoning levels so a caller can trade depth against speed per request rather than being locked into one setting. It takes tools and tool_choice for function calling and supports a JSON response format when the output feeds a pipeline.
The value claim is the interesting part. On SWE-Bench Verified with its advisor mode enabled, StepFun reports Step 3.7 Flash reaching around 97% of a frontier model's coding score at roughly a ninth of the per-task cost. Treat any single benchmark carefully, but the shape of the claim matches what the architecture is for.
The context window is 262,144 tokens, so a substantial codebase or a long document set fits in one conversation. On NinjaChat it is included in every plan — a good default when you want a capable model to answer fast and often.
शून्य से पहले नतीजे तक, एक मिनट से भी कम में।
01
Create a NinjaChat account and choose a plan
02
Open chat and pick Step 3.7 Flash from the model list
03
Use it as your default for fast, frequent questions
04
Escalate to a frontier model when a task gets hard
ईमानदार तुलना — जहाँ {model} जीतता है, और जहाँ नहीं।
Larger context and selectable reasoning levels
Qwen 3.8 Flash is cheaper still
Cheaper per token at a similar everyday coding job
GLM 4.7 has the wider GLM ecosystem behind it
Much faster and cheaper for routine work
Kimi K2.7 Code is far stronger on long agent runs
हर NinjaChat plan में 50+ models, एक पूरा image studio, और वीडियो जनरेशन शामिल है।



एक ही सब्सक्रिप्शन में NinjaChat के सभी मॉडल शामिल हैं, Step 3.7 Flash भी।