On June 4, NVIDIA released Nemotron 3 Ultra, a new open model built for long-running AI agents (AI that carries out tasks autonomously, step by step)[1][2]. It uses a configuration with 550 billion total parameters, of which only 55 billion are active at any time, and NVIDIA says it balances strong reasoning with high processing speed[2]. The model also became available on Ollama's cloud the same day, so it can be tried with a single command[1][3].

What Nemotron 3 Ultra Is

Nemotron 3 Ultra is a Mixture-of-Experts (MoE, a design that uses only a subset of specialized parameters for each input) model with 550 billion total parameters and 55 billion active parameters[2]. It is an open model with published weights, training recipes, and licensing, so developers can fine-tune it for their own use cases[2].

What NVIDIA is targeting with this model is the role of orchestrator for agents that run for a long time across many steps[2]. Unlike a chatbot that answers a single question, an agent plans, calls tools, delegates work to other agents, receives results, and then reasons again, repeating this loop many times[2]. The tokens (the smallest units of text an AI processes) exchanged in this process grow quickly, which tends to increase cost and the risk of drifting from the goal[2]. NVIDIA envisions combining a high-performance model that handles complex decisions with lightweight models that handle high-volume execution, and explains that Nemotron 3 Ultra is meant to take on the "hard calls"[2].

Design Choices for Efficiency

Nemotron 3 Ultra incorporates several design choices aimed at balancing performance and efficiency[2]. It adopts a hybrid configuration that combines "Mamba" layers, which handle long context efficiently, with "Transformer" layers, which accurately recall the information that is needed, so it can keep context across long-running tasks[2]. It also supports quantization (a compression technique that makes a model lighter) using NVFP4, NVIDIA's own 4-bit floating-point format, which the company says allows it to run across a range of GPUs[2].

On performance, NVIDIA states that throughput (the amount processed per unit of time) reaches 5x that of other open models in its class[2]. The company also says that in evaluations on "SWE-bench" and "Terminal-bench 2.0," which measure coding ability, the model completed tasks using fewer tokens, lowering the cost of agentic work by up to 30%[2]. As for long-context ability, NVIDIA says it scored 95% on "Ruler," an evaluation at the 1-million-token scale[2]. It should be noted, however, that these are all figures published by NVIDIA itself and are not third-party verified.

Easy to Try on Ollama Cloud

At the same time as its release, Nemotron 3 Ultra became usable through Ollama's cloud[1]. After installing Ollama, running the following command lets you start a conversation right away[1].

ollama run nemotron-3-ultra:cloud

You can also build it into coding-assistance tools. For example, to run it in Claude Code, you launch it like this[1].

ollama launch claude --model nemotron-3-ultra:cloud

In addition, it supports integration with tools such as Codex App, OpenClaw, Hermes Agent, and OpenCode[1]. Calls from the API, Python, and JavaScript are also provided[3]. According to Ollama's model page, the cloud version supports tool calling and the display of the thinking process, and the length of context it can handle is listed as 256,000 tokens[3]. Compared with the 1-million-token evaluation figure cited by NVIDIA, the context ceiling currently offered on Ollama's cloud is a more modest setting[2][3].

Summary

Nemotron 3 Ultra is an NVIDIA open model aimed at serving as the orchestrator for long-running agents, and its Mixture-of-Experts configuration and NVFP4 quantization are claimed to balance performance and efficiency[2]. It is also available on Ollama's cloud, and a major appeal is that it can be tried from your own machine with a single command[1][3]. The performance numbers are, for now, on a vendor-published basis, but the expansion of openly released large-model options is no small matter.

Source: https://ollama.com/blog/nemotron-3-ultra

Source: https://developer.nvidia.com/blog/nvidia-nemotron-3-ultra-powers-faster-more-efficient-reasoning-for-long-running-agents/

Source: https://ollama.com/library/nemotron-3-ultra