NVIDIA
NVIDIA's open frontier model, built for agents that orchestrate other agents rather than answer once.
262.144K tokens
コンテキスト
131,072 tokens
最大出力
普通
速度
Coordinating multi-step work across tools and stages.
Long chains of reading, checking and synthesising.
Jobs that keep running rather than answering once.
Work that does not fit in a single prompt.
Tuned to this model — click any line to copy.
Nemotron 3 Ultra is NVIDIA's open frontier model, released on 4 June 2026 after being announced at Computex. It is a 550-billion-parameter mixture-of-experts model with about 55 billion parameters active per token, and unusually it is not a plain transformer: NVIDIA built it on a hybrid Transformer-Mamba architecture, which is part of how it sustains high throughput on long inputs.
NVIDIA is explicit about what it is for. This is a reasoning and orchestration model, aimed at long-running agentic work — an agent that coordinates other agents, a coding agent that keeps going, deep research that spans many steps, and enterprise tasks that do not fit in a single prompt. The design includes multi-token prediction layers with shared weights, which improves the training signal and enables native speculative decoding, so it produces tokens quickly for a model of its size.
It is genuinely open. The weights ship under the OpenMDW 1.1 licence for commercial and non-commercial use, and NVIDIA published training recipes and a multi-trillion-token pre-training dataset alongside them, which is a level of disclosure most labs at this scale do not offer. NVIDIA documents a pre-training data cutoff of September 2025 with post-training data running to May 2026.
On NinjaChat it serves about 256K tokens of context with up to 128K tokens of output — the shape our providers run — so you can hand it a long research task or a multi-stage plan and let it work.
ゼロから最初の結果まで、1分もかかりません。
01
Create a NinjaChat account and choose a plan
02
Open chat and pick Nemotron 3 Ultra from the model list
03
Give it the whole objective, not one step of it
04
Let it plan the stages before it starts executing