Skild AI has shared results from S1, its flagship robot foundation model. Show it a single video of the job you want done, and it performs that job without any retraining. It reportedly works even on long, 10-minute sequences that never appeared in its training data, removing the assumption that has long defined industrial robotics: every new task means fresh data collection and another training run[1][2].

One video in, no weight updates

Factory floors and warehouses rarely stay still. Products change, layouts shift, new processes arrive. Most industrial robots are built around fixed jobs, so each change triggers another round of data collection, retraining and validation[1].

S1 takes a different route. An operator records a video of the desired task and hands it to the model as a prompt. S1 reads the intent, the objects and the order of steps from that footage, then maps it onto the robot standing in front of it. Crucially, the model weights are never updated. This is in-context learning, the same mechanism that separates modern large language models from their predecessors, applied to manipulation[1][2].

Robot learning has so far leaned on language instructions, the approach behind conventional VLA policies. Simple atomic actions like "hand me the mug" are easy enough to describe, but folding a fitted sheet, whisking egg whites or tying a knot are things people teach by showing, not telling. S1 builds that distinction into its design[2].

More than seven times the success rate on unseen tasks

Skild AI demonstrated four tasks absent from pre-training: potting a plant, cooking pancakes, brewing pour-over coffee and assembling a kit. Each runs up to 10 minutes and spans dozens of manipulation steps[1][2].

The numbers are stark. On unseen tasks, S1 reached a 66 percent cumulative per-step success rate, while a language-prompted policy trained on identical data managed 9 percent — a gap of more than seven times. The gap also widened as pre-training data grew, which Skild AI reads as a promising scaling law for this approach[2].

On tasks that were part of pre-training, S1 reaches roughly 96 percent. What makes the comparison interesting is that at a smaller scale of 1,000 hours the ordering flipped: the language-prompted policy hit 53 percent against 43 percent for S1. Instructions compressed into language hold an advantage while data is scarce; the richer information in video pays off once scale arrives[2].

S1 also completed tasks when objects were slid away mid-execution, swapped for different ones, or the lighting changed. In one case the demonstration watered a plant with a watering can, but only a cup was available, so S1 used the cup. It treats the demonstration as a goal to reach rather than a trajectory to copy[2].

One video is worth about 380 teleoperation runs

For anyone weighing deployment effort, this is the comparison that matters. Reaching the same level with conventional methods is estimated to require roughly 380 teleoperated demonstrations. For tasks running 4 to 10 minutes, collecting those 380 runs takes 50 to 100 hours of hands-on work. A single video replaces all of it[1][2].

In the plant-potting test, Skild AI logged 11 minutes between finishing the demonstration recording and the robot starting autonomous execution. Counting from the moment soil and a pot arrived at the office, most of the elapsed time went into moving furniture and setting up the scene[1][2].

That said, the conventional route does not stay behind forever. Post-training on 2,000 demonstrations pushes success to 86 percent, ahead of S1's 66 percent[2]. Whether 66 percent from a single video is good enough depends on the failure rate a given site can absorb and the cost of gathering demonstrations.

Trained on NVIDIA infrastructure, deployed alongside Foxconn

S1 was trained and validated on NVIDIA AI infrastructure. The NVIDIA Cosmos family of open world foundation models helps diversify training data and turn video into structured descriptions, while Cosmos Curator annotates, filters and organizes data at scale. Before hardware deployment, NVIDIA Omniverse libraries and the Isaac Sim framework provide physically based virtual environments for generating data and testing edge cases[1].

Reinforcement learning runs in Isaac Lab, powered by the Newton physics engine, which lets Skild AI model forces, contact, collision and pressure accurately enough to narrow the simulation-to-reality gap. The two companies are also jointly building GPU-accelerated simulation solvers for how robots touch, grip and manipulate solid objects, to be released to developers as part of Newton[1].

The factory side is equally concrete. Skild AI, NVIDIA and Foxconn are running Skild Brain on dual-arm manipulators for high-precision assembly of NVIDIA Blackwell systems. In one published workflow, a robot installs a busbar and limit block, fastens 16 screws, and adapts when the scene diverges from the plan[1].

On the business side, Skild AI reports reaching a 100 million USD annual revenue run rate (about 15.3 billion yen) 10 months after its first commercial deployment, with more than 60 deployment partnerships spanning manufacturing, logistics, inspection, security and food preparation[1].

※1 USD = 153 JPY (as of September 12, 2026)

Summary

S1 is a robot foundation model that executes untrained, 10-minute-class tasks from a single demonstration video held in context. On unseen work it hit a 66 percent success rate against 9 percent for language-prompted policies, a gap of more than seven times. The estimate that one video equals roughly 380 teleoperation runs is the kind of number that reshapes the cost structure of robot deployment. Weights and an API are not generally available yet, and commercial partner sites are running ahead of any public release — but the shift in how robots are taught, from writing programs to simply showing them, is now clearly visible.

Source: https://blogs.nvidia.com/blog/skild-ai-s1-physical-ai/

Source: https://skild.ai/blogs/s1