At Microsoft Build, NVIDIA and Microsoft announced an expanded partnership that brings together everything needed to run agentic AI end to end[1]. NVIDIA founder and CEO Jensen Huang joined Satya Nadella's keynote via livestream from Taipei, presenting a unified stack that spans Windows devices, the Azure cloud, and local environments[1]. The centerpieces are NVIDIA's open models on Microsoft Foundry and the integration of a secure runtime into the developer tool GitHub Copilot[1].
The Big Picture From Microsoft Build
NVIDIA argues that agentic AI takes more than good models. It also needs fast hardware, a secure runtime, a responsive data layer, and models tuned for long-running reasoning (the process by which a trained AI produces answers)[1]. The expanded partnership aims to make all 4 elements available to developers on Windows devices, in the Azure cloud, or locally[1].
The themes highlighted in the keynote include RTX Spark and DGX Station for Windows, GPU-accelerated Microsoft Fabric, NVIDIA open models on Microsoft Foundry, the NVIDIA OpenShell integration into GitHub Copilot, and the next generation of AI factories[1]. Among these, the cloud-side groundwork that enterprises and developers can use right away advanced the most[1].
NVIDIA Open Models Come to Microsoft Foundry
In Microsoft Foundry's agent service, models from NVIDIA, Anthropic, and OpenAI are now available alongside Hermes special agents for specific tasks, letting enterprises build agents on Azure with built-in identity and governance[1]. Anthropic's Claude now runs natively on NVIDIA GB300 Blackwell Ultra systems on Azure, with customer availability expected in the weeks ahead[1].
This month, NVIDIA will offer Nemotron 3 Ultra, a new open reasoning model for long-running agents across coding, research, and enterprise workflows, on Foundry managed compute[1]. It is joined by Nemotron 3.5 ASR for speech recognition and Nemotron 3.5 Content Safety for safety screening[1]. Developers can combine Nemotron with frontier and local models to optimize cost and quality for each workload[1].
The open-model lineup extends to physical and scientific AI as well[1]. Cosmos 3, a fully open model for physical AI, and the Earth-2 AI weather models for enterprise forecasting and risk analysis are offered through Foundry and other channels[1]. NVIDIA also provides the open-source NVIDIA Agent Toolkit and NemoClaw blueprints for building production agents, and CUDA-X libraries such as cuDF, cuOpt, AI-Q, and NeMo can be called by agents as domain-specific skills[1].
A Secure Runtime, OpenShell, Comes to GitHub Copilot
As agents move from coding assistance to autonomous execution, they need a way to do real work without being handed real credentials[1]. NVIDIA has integrated its runtime, NVIDIA OpenShell, into GitHub Copilot to address this[1].
With OpenShell, each agent runs inside an isolated container (a self-contained execution space), and every outbound call is checked against policy before it can reach files, networks, or credentials[1]. Policies are written as code, versioned in the repository, and can be updated while running[1]. OpenShell is open source under Apache 2.0, is model-agnostic, and works across on-premises, hybrid, and cloud environments[1].
GPU Acceleration for the Microsoft Fabric Data Warehouse
Because agentic AI runs on data, fast access to it is essential[1]. NVIDIA accelerated computing is now built into the Microsoft Fabric data warehouse, and in Microsoft's internal benchmarking, SQL execution for high-concurrency workloads ran up to 6x faster than a CPU-based baseline and up to 7x faster than 3 other leading cloud data warehouse providers[1].
This lets a company's data layer keep pace with agents that continuously query and reason over data[1]. NVIDIA describes it as the result of years of collaboration spanning research to production[1].
From the Cloud to Local and Windows Devices
The partnership reaches beyond the cloud as well[1]. Microsoft is bringing Foundry Local on its on-premises platform, Azure Local, to the NVIDIA RTX PRO 6000 Blackwell Server Edition[1]. Combined with the Nemotron open models, enterprises can run high-performance AI where their data resides, whether on-premises, hybrid, or in sovereign (self-contained within a country) environments[1]. Foundry Local on Azure Local now also supports multinode deployments and the vLLM inference runtime[1].
On Windows devices, the personal-agent PC RTX Spark and the deskside AI supercomputer DGX Station for Windows both run the OpenShell described above, so agents can run safely on the local machine[1]. In physical AI, Microsoft is integrating NVIDIA's open-source physical AI tools into Azure and its Physical AI Toolchain for developing robots, autonomous vehicles, and more[1].
The large-scale foundation is coming online too[1]. Microsoft's Fairwater Wisconsin AI factory has gone live ahead of schedule, and the next-generation NVIDIA Vera Rubin platform has entered production[1]. According to NVIDIA, Vera Rubin can be added alongside Blackwell, delivering up to 10x the inference throughput per megawatt and sharply lowering the cost per token that agents process[1].
Summary
This announcement is, in effect, a partnership that frees agentic AI from any single place to run. With open models on Microsoft Foundry and OpenShell integrated into GitHub Copilot, developers can more easily build and operate securely managed agents in the cloud or on their own machines. Add the Fabric acceleration and the Azure Local rollout, and NVIDIA and Microsoft have made their direction clear: lining up the model, runtime, data, and infrastructure layers into one continuous stack from Windows to the cloud.
Source: https://blogs.nvidia.com/blog/microsoft-build-windows-local-cloud-devices/
