Robots can now copy unfamiliar tasks from a single demonstration
Skild AI introduced S1, a robotics foundation model that it says can learn an unfamiliar task from a single video demonstration without fine-tuning. If the result generalizes, the important breakthrough is not one impressive robot trick. It is replacing task-specific retraining with in-context learning for physical work.
Introducing S1, our new foundation model that learns from one example.
— Skild AI (@SkildAI) August 25, 2026
It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning.
Watch S1 operate in real-time via in-context learning: pic.twitter.com/wmF3Byv179
Q1What was actually announced?
Skild AI announced S1 directly, saying the model can learn tasks lasting up to ten minutes from one video prompt and then operate in real time without fine-tuning. The demonstrations and benchmark numbers are company-reported, so independent replication remains important.
Q2How big is the signal?
Skild has described training on about 100,000 hours of robot and interaction data. In internal evaluations, video prompting reportedly reached around 66% success on unseen tasks versus about 9% for language prompting. The company also shows multi-minute tasks such as plant repotting after recording a single example.
Q3Why does it matter now?
Programming robots one task at a time is too expensive for general-purpose deployment. Humans instead watch someone perform a new job and imitate it. A foundation model that can do the same turns a camera demonstration into a temporary program, making warehouse, home, and field robots far easier to adapt.
Q4What is the catch?
One-shot success can hide narrow task selection, favorable environments, or retries. Long-horizon manipulation compounds small perception and control errors, and a 66% success rate is not sufficient for many commercial settings. The model also needs to transfer across robot bodies, viewpoints, and object variations.
Q5What should we watch next?
Watch independent benchmarks, success after genuinely user-recorded demonstrations, cross-embodiment transfer, and performance over hundreds of repetitions. The real threshold is when a non-expert can teach a deployed robot a useful task in minutes and trust it to repeat the job safely.
