Anthropic
Anthropic's flagship for the work you would hand a senior colleague — thinking on by default, a million tokens of context.
1M tokens
上下文窗口
May 2026
知识截止日期
128,000 tokens
最大输出
较慢
速度
Three prompts, three unedited answers, generated on September 2, 2026 through NinjaChat. Click the prompt to copy it and run it yourself.
Write a Python function that merges overlapping intervals, plus three short tests. Keep the whole answer under 40 lines of code.
def merge_intervals(intervals):
"""Merge overlapping intervals.
Args:
intervals: iterable of (start, end) pairs with start <= end.
Returns:
New list of disjoint (start, end) tuples sorted by start.
Touching intervals like (1, 2) and (2, 3) are merged.
"""
merged = []
for start, end in sorted(tuple(iv) for iv in intervals):
if merged and start <= merged[-1][1]:
prev_start, prev_end = merged[-1]
merged[-1] = (prev_start, max(prev_end, end))
else:
merged.append((start, end))
return merged
def test_overlapping_and_touching():
assert merge_intervals([(1, 3), (2, 6), (8, 10), (15, 18)]) == [
(1, 6), (8, 10), (15, 18)]
assert merge_intervals([(1, 2), (2, 3)]) == [(1, 3)]
def test_unsorted_and_nested():
assert merge_intervals([(5, 7), (1, 10), (2, 3)]) == [(1, 10)]
def test_edge_cases():
assert merge_intervals([]) == []
assert merge_intervals([(4, 4)]) == [(4, 4)]
assert merge_intervals([(1, 2), (3, 4)]) == [(1, 2), (3, 4)]
if __name__ == "__main__":
test_overlapping_and_touching()
test_unsorted_and_nested()
test_edge_cases()
print("all tests passed")
A couple of notes on choices worth knowing about:
(1, 2) and (2, 3) become (1, 3) because the comparison is start <= prev_end. If your intervals are half-open (e.g. time slots where 2 is an exclusive end), change it to start < merged[-1][1] so they stay separate.(5, 1) will produce nonsense rather than an error; add a check if the input isn't trusted.Complexity is O(n log n) time from the sort, O(n) extra space.
Read: We ran the same 3 prompts on 12 AI models — Unedited answers, latency and price for every model, measured through the NinjaChat API in September 2026.
Served by Anthropic as claude-opus-5. Answers vary run to run; these are one take each, not the best of several.
带测试的多文件改动,一路做到完成,而不是只写个草稿。
百万 token 上下文:整个代码仓库、成套合同、研究资料库都能放进一个会话。
架构评审、取舍分析,以及经得起推敲的方案。
截图、图表、PDF 和表格都可以成为对话的一部分。
Tuned to this model — click any line to copy.
Claude Opus 5 是 Anthropic Claude 5 系列的顶级模型,于 2026 年 7 月 24 日发布,定位高于 Claude Sonnet 5 和 Claude Fable 5。这是 Anthropic 专为智能体编程、电脑操作和长周期推理打造的模型:这类任务要求模型在多个步骤中始终守住一个计划、在过程中调用工具,并能察觉自己何时出错。在 NinjaChat 上,Claude Opus 5 与其他模型一起包含在所有套餐中。
最大的变化是默认开启思考。Opus 5 会自行判断回答前需要推理多少,并在问题值得时投入更多推理——因此两行字的提问会很快得到回复,而跨代码库的重构则能获得应有的深思。配合 100 万 token 的上下文窗口和 128K token 的输出,你可以在一次对话中交给它整个代码仓库、一整套合同或一个季度的研究笔记,并要求它完成贯穿全部内容的任务。
实际使用中,当工作属于那种你会交给资深同事、之后再来验收的类型,而不是随手的助理小活,Opus 5 就是首选。带测试的多文件代码改动、需要权衡取舍的架构评审、需要被辩驳而非简单摘要的长文档,都属于此类。它能读取图片和文件、调用工具,也能在你要求 JSON 时输出结构化结果,因此既适合对话,也能嵌入工作流。
它也很审慎。Opus 5 比 Sonnet 5 更慢,API 上每 token 成本更高,所以 NinjaChat 在模型选择器里同时保留了两者:难的部分用 Opus 起步,量大的部分转到 Sonnet。如果你主要看重速度,Claude Fable 5 能以更低的开销覆盖日常对话。
Anthropic 公布 Opus 5 在 SWE-bench Verified 上取得 96.0%——这项软件工程基准已成为衡量智能体编程能力的代名词。基准不等于实际工作,但这个数字与你的使用体感一致:别的模型做到一半的任务,它能做完。
从零到第一个结果,不到一分钟。
01
创建 NinjaChat 账户并选择套餐
02
打开聊天,在模型列表中选择 Claude Opus 5
03
把完整任务连同文件一起交给它,而不是只给片段
04
让它思考:越难的问题耗时越久,这是有意为之
05
后续快速追问可切换到 Sonnet 5 或 Fable 5
客观对比——{model} 的优势所在,以及它的不足之处。
在智能体编程基准和电脑操作上表现更强
Sol 的上下文窗口更长,工具生态也不同
默认开启思考、SWE-bench 分数更高、工具使用更出色
如果你已有针对 Opus 4.8 调好的提示词,它依然可用