NVIDIA has released "Metropolis agent skills" and blueprints, a set of reusable workflows that let developers build vision AI agents faster, turning footage from factories, cities, and warehouses into actionable operational intelligence[1]. By pairing them with synthetic data generation through OpenUSD and NVIDIA Omniverse, the offering aims to close three common gaps: missing training data, a shortage of fine-tuning expertise, and complex assembly work. Partners such as Roboflow, Linker Vision, and Foxconn are already reporting concrete results.

Three Hurdles Facing Vision AI Agents

Vision AI agents are systems that understand footage captured by cameras and sensors and translate it into judgments and alerts based on conditions in the field. The push to run them at the edge, close to where data is generated, is gaining momentum. Research firm Gartner projects that more than two-thirds of enterprise-managed data will be created and processed outside the data center or cloud by 2028, and that the share of enterprises deploying edge AI will grow from 10 percent in 2025 to more than two-thirds globally by 2029[1]. At the same time, as much as 90 percent of existing edge data is said to go unprocessed[1].

NVIDIA highlights three points where building these agents tends to stall[1]. The first is an accuracy plateau caused by a lack of training data. An inspection model on a production line, for example, may spot common scratches and dents but miss a new type of hairline crack that is not represented in the training data. The second is a shortage of fine-tuning know-how. Even after a performance gap is identified, steps such as preparing labeled data, configuring training, and running evaluations are required, and many organizations do not have large in-house machine learning teams. The third is the complexity of assembly work. Developers must stitch together video pipelines, models, search, alerts, and reporting, and customizing all of this for each environment takes significant time and specialized expertise.

"Agent Skills" That Make Synthetic Data and Fine-Tuning Reusable

The new agent skills and blueprints give developers reusable starting points so they do not have to rebuild these steps from scratch[1]. On the simulation and synthetic data side, the foundation is OpenUSD (Universal Scene Description), which describes and reuses 3D spaces within a common framework, along with the NVIDIA Omniverse libraries built on top of it. Teams can recreate a wide range of conditions in digital twins, including lighting, weather, traffic patterns, camera angles, occlusion, and rare events, to expand the range of scenarios available for training[1].

The skills are organized by purpose[1]. They include the Defect Image Generation skill for creating synthetic defect images, the Video Data Augmentation skill for improving scenario coverage, the NVIDIA TAO skills for model fine-tuning, and the VSS (video search and summarization) skills that turn video search and summarization into operational workflows such as alerts, reporting, and stream management. Used alongside NVIDIA Omniverse and Metropolis, they are intended to streamline the entire process from data generation to model improvement and agent deployment.

Results in Manufacturing, Smart Cities, and Industrial Sites

In manufacturing, the more successfully a factory prevents defects, the harder it becomes to collect the defect samples needed to train the next inspection model[1]. Roboflow has integrated NVIDIA's Defect Image Generation skill and NVIDIA Cosmos (world foundation models) into its vision AI platform to generate synthetic defect images when real data is scarce. In a benchmark conducted with the engineering team at Corning, an optical fiber maker, a model trained on just 8 real defect images supplemented with synthetic data achieved 95 percent average precision and perfect recall on the most challenging defect class, surpassing a baseline model trained on real data alone. NVIDIA says this compressed an inspection project that would take several quarters into just a few days[1].

In smart cities, Linker Vision uses the NVIDIA Metropolis Blueprint for VSS to deploy agents that interpret video across city infrastructure[1]. OpenUSD-based Omniverse digital twins recreate the urban environment so teams can test how the system responds to traffic, weather, and emergencies. In Kaohsiung, Taiwan, using this VSS blueprint cut development effort by 85 percent and reduced incident response times by up to 80 percent. The successor AI-GRID expansion adopts NVIDIA NemoClaw blueprints for secure agentic AI[1].

In industrial settings, what matters is not only detecting what appears in a video frame but understanding whether work is being performed correctly[1]. At Foxconn, DeepHow's Live SOP (standard operating procedure) Verification agent uses the NVIDIA Metropolis VSS blueprint as the layer for search, summarization, and analysis, and relies on NVIDIA Cosmos for the reasoning that interprets human work sequences in context. Deployed on NVIDIA GB300 server production lines, it improved first-pass yield by 3 percent, achieved 99 percent accuracy in micro-action understanding of critical SOP steps, and reduced rework by catching problems earlier[1].

Summary

NVIDIA's agent skills and blueprints package the stages of vision AI work — synthetic data generation, model fine-tuning, and agent deployment — into reusable building blocks so that even companies without specialized teams can accelerate development. The standout point this time is that field results are shown with concrete numbers, from Roboflow's high accuracy starting with 8 real images, to an 85 percent reduction in development effort in Kaohsiung, to improved yield at Foxconn. As edge-running AI moves into practical use, the groundwork for lowering the development barrier is taking shape.

出典:https://blogs.nvidia.com/blog/vision-ai-agent-skills-omniverse-metropolis/