NVIDIA
NVIDIA's open frontier model, built for agents that orchestrate other agents rather than answer once.
262.144K tokens
कॉन्टेक्स्ट
131,072 tokens
अधिकतम आउटपुट
मध्यम
गति
Coordinating multi-step work across tools and stages.
Long chains of reading, checking and synthesising.
Jobs that keep running rather than answering once.
Work that does not fit in a single prompt.
Tuned to this model — click any line to copy.
Nemotron 3 Ultra is NVIDIA's open frontier model, released on 4 June 2026 after being announced at Computex. It is a 550-billion-parameter mixture-of-experts model with about 55 billion parameters active per token, and unusually it is not a plain transformer: NVIDIA built it on a hybrid Transformer-Mamba architecture, which is part of how it sustains high throughput on long inputs.
NVIDIA is explicit about what it is for. This is a reasoning and orchestration model, aimed at long-running agentic work — an agent that coordinates other agents, a coding agent that keeps going, deep research that spans many steps, and enterprise tasks that do not fit in a single prompt. The design includes multi-token prediction layers with shared weights, which improves the training signal and enables native speculative decoding, so it produces tokens quickly for a model of its size.
It is genuinely open. The weights ship under the OpenMDW 1.1 licence for commercial and non-commercial use, and NVIDIA published training recipes and a multi-trillion-token pre-training dataset alongside them, which is a level of disclosure most labs at this scale do not offer. NVIDIA documents a pre-training data cutoff of September 2025 with post-training data running to May 2026.
On NinjaChat it serves about 256K tokens of context with up to 128K tokens of output — the shape our providers run — so you can hand it a long research task or a multi-stage plan and let it work.
शून्य से पहले नतीजे तक, एक मिनट से भी कम में।
01
Create a NinjaChat account and choose a plan
02
Open chat and pick Nemotron 3 Ultra from the model list
03
Give it the whole objective, not one step of it
04
Let it plan the stages before it starts executing
ईमानदार तुलना — जहाँ {model} जीतता है, और जहाँ नहीं।
Built specifically for orchestration and long agent runs
Inkling is multimodal with a larger context
Open weights with published training recipes
Kimi K3 is stronger on general chat
Hybrid architecture built for sustained throughput
GLM 5.2 preserves reasoning across turns and tool calls
हर NinjaChat plan में 50+ models, एक पूरा image studio, और वीडियो जनरेशन शामिल है।



एक ही सब्सक्रिप्शन में NinjaChat के सभी मॉडल शामिल हैं, Nemotron 3 Ultra भी।