Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video
| Source: NVIDIA Blog
Tags: Skild AI, S1, robotics, NVIDIA, in-context learning, physical AI, foundation model
Skild AI's S1 robot foundation model learns previously unseen, long-horizon tasks from a single video demo using in-context learning — no retraining required. It achieves 66% step-level success vs. 9% for competing systems on novel multistep tasks, and the company hit $100M ARR 10 months after first commercial deployment.
Details
Skild AI launched S1, a robotic foundation model that uses in-context learning to execute tasks from a single video demonstration without updating its weights or undergoing task-specific post-training. An operator records a video of the desired task, and S1 interprets the intent, objects, and sequence, then maps them to robot actions on new hardware it hasn't seen before. Benchmarks show S1 succeeding at 66% of steps on novel multistep tasks, compared with 9% for a comparable AI system — a 7x improvement in generalization. The model handles tasks lasting up to 10 minutes with dozens of manipulation steps, including plant potting, pancake making, and kit assembly. In one test, the team moved from recording a video demonstration to autonomous robot execution in 11 minutes. The model was built on NVIDIA Isaac Lab and Cosmos, covering synthetic data generation, training, simulation, and real-world deployment. The NVIDIA collaboration is positioned as the foundation for scaling adaptable robot intelligence from lab settings into manufacturing, logistics, and other dynamic industrial environments. Commercially, Skild AI reached $100M annual revenue run rate 10 months after its first deployment and has built 60+ partnerships across manufacturing, logistics, inspection, security, and food preparation — signaling rapid market penetration for an approach that traditionally requires substantial per-task training.