NinjaChat LogoNinjaChat Logo
模型定价工具博客
控制台
模型定价工具博客登录Download on the App Store
We Ran the Same 3 Prompts on 12 AI Models (September 2026): Unedited Answers, Latency, Price 封面图片

We Ran the Same 3 Prompts on 12 AI Models (September 2026): Unedited Answers, Latency, Price

Claude Opus 5, GPT-5.5, the three GPT-5.6 sizes, Kimi K3, Gemini 3.7 Flash, GLM-5.2, GLM-4.7, MiniMax M3, gpt-oss-120b and Qwen 3.8 Flash, given one coding task, one founder question and one rewrite through the NinjaChat API. Every answer unedited, every response timed, every request billed at the price you would pay.

Siddharth Duggal 头像Siddharth Duggal·2026年9月4日

本页目录

  • How we measured
  • Prompt 1: merge overlappin...
  • Prompt 2: monolith or micr...
  • Prompt 3: three tones
  • The whole day in one table
  • Which model for what
  • Run it yourself
  • Frequently asked questions

Every model comparison you read is either a benchmark table nobody can reproduce or a vibes review of two models the author already liked. So on September 2, 2026 we did the boring version: the same three prompts, sent once each to twelve current models through our own API, timed on the stream, billed at the price a NinjaChat key actually pays. Nothing below has been edited, cherry-picked or re-run to look better.

The short version. All twelve wrote a correct interval-merging function whose own tests pass. GPT-5.6 Terra was the fastest model to first token on average and the fastest coder outright, under three seconds. GPT-5.6 Luna and gpt-oss-120b tied for cheapest at about a tenth of a cent for all three answers. The two most expensive single answers of the day were Kimi K3's founder advice, which reasoned for 1,034 tokens and took 27 seconds, and Claude Opus 5's code, which is also the most thorough code anyone wrote. In between, the interesting differences were not about correctness at all. They were about how long each model thinks before it speaks, which host served it, and what both of those cost you.

How we measured

  • Date and place: September 2, 2026, one afternoon, one laptop on a home connection in the US.
  • Route: every request went to the NinjaChat API, POST /api/v1/chat/completions, with a normal playground key. The latency you see includes our routing, the provider we picked, and the provider's own time. The Served by column names that provider for each answer.
  • Models: Claude Opus 5, GPT-5.5, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, Kimi K3, Gemini 3.7 Flash, GLM-5.2, GLM-4.7, MiniMax M3, gpt-oss-120b and Qwen 3.8 Flash.
  • Settings: temperature 0.7, a 2,500-token output cap, streaming on with usage reporting, every other parameter left at the model's default. Most of these models reason by default, and our API reports reasoning tokens separately from visible output, so both columns are in every table. Two hosts, DeepInfra for GLM-4.7 and Fireworks for GLM-5.2, report reasoning inside the output count instead, so those rows show a large output number and a zero in the reasoning column.
  • Timing: time to first token is when the first visible character of the answer arrived, after any hidden reasoning. Total time is until the stream closed. Requests ran one at a time, never in parallel, after a warm-up request to each model so that nothing below includes a cold start on our side.
  • Cost: the amount our API reported billing for that request, to the ten-thousandth of a dollar. Nothing is estimated from a price list.
  • One run per cell. No retries, no best-of-three. A second run will not reproduce these numbers to the decimal. It should reproduce the ranking.

The three prompts were copied verbatim from the Real Outputs sections on our model pages, so what you see here is the same task those pages show:

  1. "Write a Python function that merges overlapping intervals, plus three short tests. Keep the whole answer under 40 lines of code."
  2. "In about 120 words, tell a first-time founder when to choose a monolith over microservices, and the one signal that means it's time to split."
  3. "Rewrite this sentence three ways, each in a clearly different tone (plain, playful, executive): 'Our app makes budgeting easy.' Label each."

Prompt 1: merge overlapping intervals

Sorted by total time. Every answer below passed its own tests when we ran the code.

ModelServed byTime to first tokenTotal timeOutput tokensReasoning tokensBilled for this answer
GPT-5.6 TerraOpenAI1.5 s2.9 s1840$0.0023
Gemini 3.7 FlashGoogle2.8 s3.3 s232445$0.0052
MiniMax M3Fireworks3.3 s4.9 s285271$0.0008
gpt-oss-120bFireworks4.4 s6.3 s280434$0.0006
GPT-5.6 LunaOpenAI5.9 s7.2 s171374$0.0007
Claude Opus 5Anthropic3.7 s8.6 s62189$0.0180
Qwen 3.8 FlashNovita8.1 s9.1 s175757$0.0005
GPT-5.6 SolOpenAI8.8 s10.7 s19426$0.0046
GLM-5.2Fireworks12.9 s14.6 s9890$0.0045
Kimi K3Fireworks16.5 s21.7 s309771$0.0166
GPT-5.5OpenAI9.5 s29.4 s19660$0.0079
GLM-4.7DeepInfra60.6 s67.1 s13630$0.0031

Twelve for twelve on correctness, which tells you this is a solved problem for anything current. The differences are in taste, in how the model spends tokens, and in one case in what "test" means.

The tightest answer came from GLM-4.7: fourteen lines, no commentary, no separate empty-input check because the loop handles it.

def merge_intervals(intervals):
    intervals.sort(key=lambda x: x[0])
    merged = []
    for start, end in intervals:
        if not merged or start > merged[-1][1]:
            merged.append([start, end])
        else:
            merged[-1][1] = max(merged[-1][1], end)
    return merged

Claude Opus 5 went the other way. It wrote the only version with a docstring, three named test functions covering empty, single, touching, unsorted and nested input, and a note explaining that touching intervals merge because the comparison is closed, with the one-character change if you want half-open semantics. It is also the answer most likely to have annoyed a strict reader: the code block is exactly 40 lines, against a prompt that asked for under 40. Kimi K3 wrote the next most careful version, one test function with three commented cases and an unsorted, nested input.

The GPT family produced nearly the same function four times. Terra did it in 2.9 seconds with no reasoning tokens and no docstring, the fastest answer of the prompt. Sol added a docstring and returned tuples. Luna reasoned for 374 tokens on the way to the shortest GPT answer. GPT-5.5 wrote the same function as Sol, then its stream stalled: first token at 9.5 seconds, last token at 29.4, for 196 tokens of output. That is the kind of single-run event a second run would probably not repeat, and exactly why we say so.

MiniMax M3 is the row to read twice. Its three "tests" are print statements with the expected value in a comment beside each, so they cannot fail. The code is right, and we ran it, but the model answered a different question than the one asked. Four models, Gemini 3.7 Flash, GLM-4.7, MiniMax M3 and Qwen 3.8 Flash, sort the caller's list in place, which is fine for a snippet and a bug in a library.

The slow rows are host rows. GLM-4.7, served by DeepInfra, took a full minute to its first character; the host counts its reasoning inside the 1,363 output tokens it reported. Kimi K3 on Fireworks reasoned for 771 tokens and took 22 seconds.

Verdict: Terra for a function like this. Opus 5 if you want the answer that teaches you something, and accept the 40 lines.

Prompt 2: monolith or microservices

ModelServed byTime to first tokenTotal timeOutput tokensReasoning tokensBilled for this answer
MiniMax M3Fireworks2.5 s4.0 s17776$0.0004
GPT-5.6 LunaOpenAI2.6 s4.4 s16575$0.0003
GPT-5.6 TerraOpenAI2.4 s4.6 s1670$0.0021
Gemini 3.7 FlashGoogle4.3 s4.6 s160901$0.0081
GPT-5.5OpenAI2.7 s5.0 s16331$0.0061
gpt-oss-120bFireworks5.5 s6.5 s18070$0.0003
Claude Opus 5Anthropic2.7 s7.5 s33223$0.0092
GPT-5.6 SolOpenAI7.0 s10.0 s164238$0.0082
GLM-5.2Fireworks15.0 s16.5 s12590$0.0057
Kimi K3Fireworks23.8 s26.9 s1701034$0.0185
Qwen 3.8 FlashNovita29.9 s30.9 s1623810$0.0019
GLM-4.7DeepInfra78.9 s86.2 s10630$0.0024

Every model gave the same headline advice, start with a monolith, which is the correct headline advice. The question was really about the second half: name the one signal that means it is time to split. The answers fell into two camps.

The organisational camp said the signal is people. Kimi K3: "teams can no longer deploy independently," with a warning that splitting along the wrong seams gives you "a distributed monolith: all the operational pain, none of the independence." Gemini 3.7 Flash called it "organizational bottlenecking" and added the line "microservices are a tool to scale human organizations and deployment velocity, not early-stage codebases." GLM-5.2 called it "deployment contention." GLM-4.7 said "organizational friction," and offered a test: if no single person can understand the whole system any more, you have outgrown the monolith.

The bottleneck camp said the signal is one specific part of the system. Claude Opus 5 made it concrete, "a CPU-heavy video transcoder starving your API," and closed with the best sentence of the day: "Everything else — 'it feels messy,' 'Netflix does it' — is not a signal." GPT-5.6 Sol: "split when one well-defined part of the product repeatedly needs to be deployed or scaled independently, and the monolith measurably prevents it," then "extract only that boundary, give it a clear interface and owner, and leave everything else together." Luna added the sentence the others forgot: extract it, "measure the improvement," and stop there. MiniMax M3 got the same idea into fewer words and one memorable line: "Premature splitting kills more startups than bad code does."

The table adds what the prose hides. Opus 5 ran to 162 words against a 120-word target, the longest answer of the prompt. Qwen 3.8 Flash reasoned for 3,810 tokens, more than twenty times what it wrote, and took 31 seconds to say something sensible. Sol reasoned for 238 tokens, took ten seconds and was billed nearly four times what Terra's answer cost, for a paragraph that is better by a sentence.

Verdict: Opus 5 if you want the answer you would forward to a founder. GPT-5.5 or Terra for the best paragraph per second. MiniMax M3 if the point is that this cost four hundredths of a cent.

Prompt 3: three tones

ModelServed byTime to first tokenTotal timeOutput tokensReasoning tokensBilled for this answer
gpt-oss-120bFireworks1.6 s1.8 s6241$0.0002
GPT-5.6 LunaOpenAI1.6 s2.1 s440$0.0001
GPT-5.6 SolOpenAI2.2 s3.2 s470$0.0011
GPT-5.6 TerraOpenAI2.4 s3.3 s540$0.0008
GPT-5.5OpenAI2.2 s3.4 s470$0.0016
MiniMax M3Fireworks3.0 s3.4 s55191$0.0004
Gemini 3.7 FlashGoogle3.3 s3.4 s57429$0.0037
Claude Opus 5Anthropic2.8 s3.9 s10128$0.0035
Qwen 3.8 FlashNovita6.3 s6.7 s4448$0.0001
Kimi K3Fireworks9.2 s10.2 s69413$0.0076
GLM-5.2Fireworks10.8 s11.3 s7470$0.0034
GLM-4.7DeepInfra68.6 s69.7 s10490$0.0024

This is the prompt that exposes lazy reading. gpt-oss-120b and Qwen 3.8 Flash both returned the original sentence, unchanged, as their "plain" version. That is a defensible reading of the word plain and an obvious miss of the word rewrite. gpt-oss-120b has done it on every run of this test we have made, so it is a habit, not a fluke.

Two models wrote a playful line with an actual voice. Kimi K3: "Budgeting? Piece of cake. Our app does the number-crunching so you don't have to!" Claude Opus 5: "Budgeting used to be a chore — now it's basically a tap, a swipe, and you're done." Opus 5 also wrote the only executive line with a concrete claim in it, "reducing financial admin to a few minutes a month," where everyone else reached for abstractions. Seven of twelve used "a breeze" somewhere, and eleven of twelve used "streamlines" in the executive line. Gemini's executive version, "seamless budget optimization and fiscal clarity," is the one to show your marketing team as a warning.

The cost story is the real one here. GLM-4.7 took 70 seconds to produce 32 words, because its host reported 1,049 output tokens for them, most of it reasoning it does not itemize. Luna and Qwen 3.8 Flash each did the job for a hundredth of a cent, Luna in two seconds. If your product rewrites sentences at scale, this is a cost decision, not a quality one, unless you specifically want Opus 5's or Kimi's voice.

Verdict: Opus 5 or Kimi K3 for copy a human will read, Luna for copy a pipeline will read.

The whole day in one table

What each model was billed for all three answers, cheapest first, with the provider that served it and the list price each model bills at on NinjaChat's API as of September 2, 2026.

ModelBilled for all three answersMean time to first tokenSlowest answerServed byInput / output price per 1M tokens
GPT-5.6 Luna$0.00113.4 s7.2 sOpenAI$0.2 / $1.2
gpt-oss-120b$0.00113.8 s6.5 sFireworks$0.35 / $0.75
MiniMax M3$0.00162.9 s4.9 sFireworks$0.3 / $1.2
Qwen 3.8 Flash$0.002514.8 s30.9 sNovita$0.15 / $0.47
GPT-5.6 Terra$0.00522.1 s4.6 sOpenAI$2 / $12
GLM-4.7$0.007969.4 s86.2 sDeepInfra$0.6 / $2.2
GLM-5.2$0.013612.9 s16.5 sFireworks$1.4 / $4.4
GPT-5.6 Sol$0.01396.0 s10.7 sOpenAI$4 / $20
GPT-5.5$0.01564.8 s29.4 sOpenAI$5 / $30
Gemini 3.7 Flash$0.01703.5 s4.6 sGoogle$1.5 / $7.5
Claude Opus 5$0.03073.1 s8.6 sAnthropic$5 / $25
Kimi K3$0.042716.5 s26.9 sFireworks$3 / $15

Three patterns fall out of it.

Reasoning is the hidden price. Kimi K3 lists below the three frontier models and finished as the most expensive model of the day, because it reasoned for 400 to 1,000 tokens on every prompt, including the rewrite. Gemini 3.7 Flash cost more than GPT-5.5 for the same reason. Sol's most expensive answer was not its longest, it was the one it thought about. If you do not need chain-of-thought for a task, pick a model that does not do it by default, or turn the reasoning effort down where the API allows it.

Cheap is now genuinely good. Luna, gpt-oss-120b and MiniMax M3 wrote correct code, sensible advice and usable copy for between a tenth and a sixth of a cent each, all three prompts included. Two years ago that tier of model would have failed the coding prompt.

The host is part of the latency. GLM-4.7 is a fast model on paper and was the slowest thing in the test, because the host that served it took a minute to start every answer. Kimi K3 and GLM-5.2 on Fireworks were quick once they started writing and slow before it. The OpenAI, Google and Anthropic models answered from the labs directly. That is why our API shows uptime and latency per provider, not per model, and why the Served by column is in every table above.

Which model for what

  • You want the best answer and cost is secondary: Claude Opus 5. Most thorough code, the founder answer you would actually forward, the only rewrite with a concrete claim, and 3 seconds to first token on average.
  • You want frontier quality at a middle price: GPT-5.6 Terra. Fastest to first token in the whole set, no reasoning unless the question needs it, less than half of Sol's bill on the day.
  • You want the most capable OpenAI model regardless of price: GPT-5.5, but note its 32K output ceiling before you point it at long generations.
  • You are running a high-volume pipeline: GPT-5.6 Luna or MiniMax M3. A tenth to a sixth of a cent for the whole exercise, everything under five seconds on the two prose prompts.
  • You want a large open-weight model with a long context: Kimi K3 for the writing, GLM-5.2 for the price, and budget for reasoning time on both.
  • You want copy people will read: Opus 5, then Kimi K3. Skip the two models that returned the original sentence.

Run it yourself

Every model above is one model string away on the NinjaChat API, which is OpenAI-compatible. A phone-verified account starts with a $0.50 starter balance, no card needed until you top up. Swap the id and re-run any prompt in this post. The response reports what you were billed and which provider served you.

curl https://www.ninjachat.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $NINJACHAT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-terra",
    "messages": [{"role": "user", "content": "In about 120 words, tell a first-time founder when to choose a monolith over microservices, and the one signal that means it is time to split."}],
    "max_tokens": 2500,
    "temperature": 0.7,
    "stream": true,
    "stream_options": {"include_usage": true}
  }'
import time
from openai import OpenAI

client = OpenAI(base_url="https://www.ninjachat.ai/api/v1", api_key="YOUR_KEY")

for model in ["claude-opus-5", "gpt-5.6-terra", "minimax-m3"]:
    t0 = time.perf_counter()
    first = None
    stream = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": "Rewrite three ways (plain, playful, executive): Our app makes budgeting easy."}],
        max_tokens=2500,
        temperature=0.7,
        stream=True,
        stream_options={"include_usage": True},
    )
    for chunk in stream:
        if chunk.choices and chunk.choices[0].delta.content and first is None:
            first = time.perf_counter() - t0
        if chunk.usage:
            print(model, "reasoning tokens:", chunk.usage.completion_tokens_details.reasoning_tokens)
    print(model, "first token", round(first, 2), "s, total", round(time.perf_counter() - t0, 2), "s")

Live prices for every model, no key required, are on GET https://www.ninjachat.ai/api/v1/models. The consumer side is the same models in a chat window: pick any of them on /models with a paid plan.

Frequently asked questions

Which AI model was fastest in this test? GPT-5.6 Terra: 2.1 seconds to first token on average, and the fastest complete answer on the coding prompt at 2.9 seconds. MiniMax M3 and Claude Opus 5 were next, both around three seconds to first token on average.

Which AI model was cheapest? GPT-5.6 Luna and gpt-oss-120b tied at about a tenth of a cent for all three answers as billed by the NinjaChat API on September 2, 2026. MiniMax M3 was a hair behind.

Did any model get the coding task wrong? No. All twelve produced a merge function whose own tests pass. Claude Opus 5 hit the 40-line limit exactly, MiniMax M3 printed its results instead of asserting them, and four models sort the caller's list in place, which is worth knowing if you copy the code into a library.

Why are GLM-4.7 and Kimi K3 so slow in the tables? Reasoning and hosting, stacked. Both think at length before writing, and the host that served GLM-4.7 that afternoon took about a minute to start each answer and reported the reasoning inside the output count. On a different host, or with reasoning effort turned down, both land mid-table.

Why does the same model cost different amounts on different prompts? Because reasoning tokens are billed as output. GPT-5.6 Sol spent 26 reasoning tokens on the code prompt and 238 on the founder prompt, so the founder answer cost nearly twice the code answer despite being shorter.

Can I reproduce this? Yes. The three prompts are printed above verbatim, the settings are listed under How we measured, and the Python snippet times any model the same way we did and prints its reasoning tokens. Expect the absolute numbers to move between runs and the ranking to hold.

NinjaChatNinjaChat

所有 AI 模型,一站汇聚。

iPhone 应用网页版体验
NinjaChat LogoNinjaChat Logo

所有 AI 模型,尽在一处。

在 App Store 上下载

产品

  • 控制台
  • 定价
  • 免费 AI 工具
  • 推广合作
  • iOS 应用

开发者

  • API
  • API 模型
  • Ninja Router
  • MCP / Agents
  • API 文档
  • Status
  • Changelog

模型

  • Model Council
  • Gemini 2.5 Flash
  • Gemini 2.5 Pro
  • Gemini 3 Flash Preview
  • Gemini 3.1 Pro Preview
  • 聊天模型
  • AI 图像生成器
  • AI 视频生成器
  • 查看全部模型

公司

  • 博客
  • 社区
  • 加入我们
  • 客服支持
  • 隐私政策
  • 服务条款
  • Safety Protocol
  • Do Not Sell or Share My Personal Information
语言

版权所有 © 2026 NinjaChat AI。 旗下产品 Bloon 保留所有权利。