NVIDIA says its open large language model, Nemotron 3 Ultra, achieved the highest accuracy among open models in testing by LangChain, the widely used agent development platform[1]. After LangChain tuned its Deep Agents harness specifically for Nemotron 3 Ultra, the model matched the work of leading closed models while running at roughly one-tenth the inference cost per run. This article breaks down what the result actually shows and the idea of an "open AI stack" that NVIDIA is promoting alongside it.
Optimizing How the Model Is Used, Not Retraining It
The most interesting part of this result is that Nemotron 3 Ultra itself was left untouched. Instead of retraining the model, LangChain tuned the environment around it — the so-called harness that gives a model its supporting machinery. In practical terms, the team revised system prompts, tool descriptions, and the middleware sitting between processing steps[1].
LangChain first ran Nemotron 3 Ultra against its public Deep Agents benchmark, then traced execution logs to pinpoint exactly where the model was losing points. From there, it raised accuracy by adjusting the harness rather than retraining the model. Co-founder and CEO Harrison Chase said the way to build better agents is to keep improving the system around the model, noting that memory, tool use, evaluation, and model behavior compound when teams can tune them together[1].
The tuned profile is available through LangChain, so developers already pairing Deep Agents with Nemotron 3 Ultra can put it to work right away. That said, these figures were published by NVIDIA and LangChain, and real-world performance varies by task and conditions, so it is worth reading them alongside third-party verification.
What Kind of Model Is Nemotron 3 Ultra
Nemotron 3 Ultra is a large model with 550 billion total parameters, but it uses a Mixture-of-Experts (MoE, a design that selectively activates specialized sub-models) approach that keeps only 55 billion parameters active per pass[2]. That lets a very large model hold down its compute load, and it is designed to act as the command center for long-running agents.
Several internal techniques support this. The model combines Mamba layers, which handle long context efficiently, with conventional Transformer layers that retrieve specific facts precisely — a hybrid structure. It also uses a precision format called NVFP4, which lets a single model run across different generations of NVIDIA GPUs, including Hopper, Blackwell, and Ampere[2]. According to NVIDIA, throughput reaches up to 5x that of other open models in its class.
Its openness is clear as well. Nemotron 3 Ultra is released with its weights (parameters), training data, and training recipes, under the permissive OpenMDW-1.1 license maintained by the Linux Foundation. It can be obtained and tried through Hugging Face, OpenRouter, and build.nvidia.com, and it runs on your own infrastructure or in the cloud[2].
An Open Stack Built Around "AI You Can Own"
What NVIDIA emphasizes this time is not the model on its own, but the idea of opening up the full set of pieces needed to run agents safely. At the center is an open reference blueprint, NVIDIA NemoClaw for LangChain Deep Agents[1]. It combines LangChain's Deep Agents, tuned for Nemotron 3 Ultra, with NVIDIA OpenShell, a runtime environment for executing agent actions safely.
By assembling an open model, an open harness, and an open, secure runtime, enterprises can own the entire mechanism, adapt it to their own work, and run it anywhere, NVIDIA explains. As agents move from simply answering questions to actually taking action inside core business systems, the company argues, the value of that kind of ownership grows further[1].
Companies are already embedding specialized agents into their platforms, including Abridge in health care, telecom software maker Amdocs, and content management firm Box. Global consulting giant EY is also expanding its NVIDIA implementation support using the NemoClaw blueprint. LangChain's agent development platform sees more than 200 million monthly downloads, so tuning its underpinnings for Nemotron 3 Ultra is no small step[1].
Summary
The story around Nemotron 3 Ultra is a case study showing that performance and cost hinge not only on a model's raw intelligence but on the design of the machinery that runs it. How far the claim — that an open model paired with an open stack can approach leading closed models while holding down cost — actually holds up in day-to-day work remains to be seen. Attention now turns to how developers and enterprises put this open setup to use.
出典:https://blogs.nvidia.com/blog/nemotron-langchain-agents-open-stack/
