NVIDIA has released NVIDIA XR AI, a developer platform for running AI agents (AI that autonomously judges a situation and helps with tasks) on AR glasses and XR (an umbrella term covering VR, AR, and MR) devices[1]. It takes in real-world information such as video, audio, and depth, links it with enterprise data and a range of tools, and assists the wearer's work on the spot. Deployments are already underway in manufacturing, scientific research, healthcare, design, and other fields[1].

What NVIDIA XR AI Is

AI is beginning to move beyond the stage of chatbots and copilots (AI that assists with work) toward working alongside people in physical space. In factories, labs, and hospitals, AI agents that understand their surroundings, access knowledge, and act in the moment are starting to appear[1].

Building such agents into something genuinely useful on-site, however, is not easy. Beyond generating responses, they must perceive their surroundings from video, audio, and sensor data, read fast-changing conditions and spatial context, retrieve the information they need from enterprise systems, reason about the next action, and use software tools to finish the task. All of this has to happen with low latency, without distracting the wearer[1].

NVIDIA XR AI is a developer library that takes on this difficulty. It connects input from AR glasses and XR devices with AI models, enterprise data, tools, and accelerated computing, providing the foundation for agents that can perceive, reason, and act within the flow of real work[1]. NVIDIA also provides developer materials such as tutorials and samples[2].

Four Pillars That Connect On-Site Information to AI

NVIDIA XR AI brings together four broad capabilities[1]. The first is the input layer, which ingests real-world signals from AR glasses and XR devices, including video, audio, depth, pose, and sensor data.

The second connects agents to specialized tools and services. For understanding video, it can use NVIDIA Metropolis and NVIDIA Metropolis for VSS, which handles video search and summarization; for retrieving in-house knowledge, it can use NVIDIA NeMo Retriever, built for RAG (retrieval-augmented generation), which pulls relevant information and uses it in responses[1].

The third is support for a wide range of AI models. Alongside NVIDIA Nemotron for reasoning and NVIDIA Cosmos Reason for interpreting space and conditions, it can combine various compatible foundation models[1]. The fourth is the orchestration that coordinates multiple agents, together with the runtime that drives it at speed. With these in place, it becomes easier to build multimodal (a method that handles multiple types of input together) AI agents that understand space[1].

The Foundation Supporting Development

To assemble agents, developers can use NVIDIA NeMo Agent Toolkit, which handles tool calls, reasoning workflows, and coordination among multiple agents[1]. On the hardware side, NVIDIA DGX Spark, NVIDIA DGX Station, and systems with NVIDIA RTX PRO are available, so inference can run across cloud, data center, and the edge near the work site, depending on the setup[1]. This makes possible agents that understand their surroundings, draw on in-house knowledge, reason about complex tasks, and return context-appropriate support in real time[1].

From Factories to Operating Rooms to the Titanic

NVIDIA XR AI is already being tried across a variety of settings[1].

In manufacturing, Siemens, as a research effort, is exploring how NVIDIA XR AI and DGX Spark can help factory engineers look up maintenance information, isolate faults, and verify their work. An engineer wearing lightweight glasses can ask an AI agent about a problem with a PLC (programmable logic controller, a device that controls industrial equipment) and receive step-by-step guidance on the spot[1].

In scientific research, Rana, a company building AI systems, has introduced LabOS, which runs on NVIDIA XR AI, to bring spatially aware intelligence directly into experimental workflows. LabOS is first being used at the Cong Lab at Stanford University School of Medicine and the Wang Lab at Princeton University for stem cell therapy and gene-editing research. It helps researchers identify the correct sample and CRISPR (a gene-editing method) editing tool, guides each step, and keeps a reproducible record, acting as a kind of co-pilot[1]. LabOS works with smart glasses from Meta, Rokid, and VITURE. VITURE has built NVIDIA XR AI into a wearable interface that quickly surfaces the context needed at the work site and guides the next move[1].

In healthcare, the Surreality Lab at University of Pittsburgh Medical Center (UPMC) demonstrated a system that supports surgical teams as the situation requires, running on NVIDIA XR AI and DGX Station. It is designed to understand what must not be occluded in the surgeon's field of view, surfacing useful information while preserving focus on the patient and the procedure[1].

Elsewhere, Innoactive in automotive design uses DGX Spark to carry over context from design reviews, product showrooms, and digital twins (virtual recreations of real-world equipment) without losing it. Atlantic Studios in media production turned a precise scan of the Titanic as it rests today into an experience you can tour by voice command to find points of interest, transforming a complex underwater model into an interactive story[1].

Summary

NVIDIA XR AI is a developer platform for building agents that perceive, reason, and act within real work by linking information from AR glasses and XR devices with AI models, enterprise data, and accelerated computing. It assembles video understanding, in-house search, reasoning models, and orchestration in one place, and it has shown real-world examples across factories, labs, operating rooms, and immersive exhibits. It looks like a new foothold for NVIDIA toward AI that works in physical space, beyond the confines of chat.

Source: https://blogs.nvidia.com/blog/nvidia-xr-ai/

Source: https://developer.nvidia.com/xr/xr-ai