NinjaChat
模型API定价博客
开始使用→
模型API定价博客
  1. 首页›
  2. 模型›
  3. Nemotron 3 Ultra

NVIDIA

Chat with Nemotron 3 Ultra online

NVIDIA's open frontier model, built for agents that orchestrate other agents rather than answer once.

试用 Nemotron 3 Ultra 查看方案
790,000+ 位用户的信赖

262.144K tokens

上下文窗口

131,072 tokens

最大输出

中速

速度

你会用它做什么

Agent orchestration

Coordinating multi-step work across tools and stages.

Deep research

Long chains of reading, checking and synthesising.

Coding agents

Jobs that keep running rather than answering once.

Enterprise tasks

Work that does not fit in a single prompt.

Prompts to steal

Tuned to this model — click any line to copy.

关于 Nemotron 3 Ultra

Nemotron 3 Ultra is NVIDIA's open frontier model, released on 4 June 2026 after being announced at Computex. It is a 550-billion-parameter mixture-of-experts model with about 55 billion parameters active per token, and unusually it is not a plain transformer: NVIDIA built it on a hybrid Transformer-Mamba architecture, which is part of how it sustains high throughput on long inputs.

NVIDIA is explicit about what it is for. This is a reasoning and orchestration model, aimed at long-running agentic work — an agent that coordinates other agents, a coding agent that keeps going, deep research that spans many steps, and enterprise tasks that do not fit in a single prompt. The design includes multi-token prediction layers with shared weights, which improves the training signal and enables native speculative decoding, so it produces tokens quickly for a model of its size.

It is genuinely open. The weights ship under the OpenMDW 1.1 licence for commercial and non-commercial use, and NVIDIA published training recipes and a multi-trillion-token pre-training dataset alongside them, which is a level of disclosure most labs at this scale do not offer. NVIDIA documents a pre-training data cutoff of September 2025 with post-training data running to May 2026.

On NinjaChat it serves about 256K tokens of context with up to 128K tokens of output — the shape our providers run — so you can hand it a long research task or a multi-stage plan and let it work.

如何使用 Nemotron 3 Ultra

从零到第一个结果,不到一分钟。

01

Create a NinjaChat account and choose a plan

02

Open chat and pick Nemotron 3 Ultra from the model list

03

Give it the whole objective, not one step of it

04

Let it plan the stages before it starts executing

获得更好结果的技巧

  • Describe the end state; it is built to plan backwards from it
  • Good for jobs with many tool calls in sequence
  • Ask for the plan first on anything long
  • Use a Flash-class model when you just want a quick answer

Nemotron 3 Ultra 与同类模型对比

客观对比——{model} 的优势所在,以及它的不足之处。

对比 Inkling

查看模型

+Built specifically for orchestration and long agent runs

–Inkling is multimodal with a larger context

对比 Kimi K3

查看模型

+Open weights with published training recipes

–Kimi K3 is stronger on general chat

对比 GLM 5.2

查看模型

+Hybrid architecture built for sustained throughput

–GLM 5.2 preserves reasoning across turns and tool calls

常见问题解答

不止于对话——图像与视频

每个 NinjaChat 方案均包含 50+ 个模型、完整图像工作室和视频生成功能。

A real FLUX Pro Ultra outputA real Google Imagen 4 outputA real Seedream output
查看所有模型 →

Nemotron 3 Ultra,还有另外 50 多个模型。

一个订阅即可使用 NinjaChat 上的所有模型,包括 Nemotron 3 Ultra。

开始使用对比方案
Download on the App Store

更多模型,等你探索

Inkling

Thinking Machines Lab's open multimodal generalist

→

Kimi K3

Moonshot 最新开源权重旗舰,原生视觉能力,100 万上下文

→

GLM 5.2

Z.ai 面向编程、智能体与系统工程的旗舰模型

→
NinjaChat

所有 AI。 一个应用。

在 App Store 上下载

产品

  • 控制台
  • ninja
  • 定价
  • Enterprise
  • 免费 AI 工具
  • Cinema
  • 推广合作
  • iOS 应用

开发者

  • API
  • API 模型
  • Router
  • MCP / Agents
  • Search API
  • API Pricing
  • Migration Guides
  • API 文档

模型

  • Model Council
  • Seed 1.8
  • Gemini 2.5 Flash
  • Gemini 2.5 Pro
  • Gemini 3 Flash Preview
  • 查看全部模型

公司

  • 博客
  • 社区
  • 加入我们
  • 客服支持
  • 隐私政策
  • 服务条款
  • Safety Protocol
  • Do Not Sell or Share My Personal Information

版权所有 © 2026 NinjaChat AI。 旗下产品 Bloon 保留所有权利。