Learning an action from video, not generating a video
NVIDIA described Skild AI's S1 robot foundation model on September 10. An operator records the desired process once; S1 interprets the intent, objects and sequence and maps them to actions for the robot in front of it. The companies describe this as in-context learning for unfamiliar work without building a new dataset or performing task-specific post-training. It is a physical-AI use of video as a prompt, not a video-generation model.
Skild demonstrates tasks such as potting a plant, making pancakes and pour-over coffee, and assembling kits. They can span dozens of steps and last as long as ten minutes. In one plant-potting example, the company says it moved from recording to autonomous hardware execution in 11 minutes. It reports about 66% success per step on new multi-step tasks versus 9% for a comparison system and estimates one short video can carry information similar to roughly 380 hands-on examples.
Per-step success is not end-to-end success
The figures are vendor and partner reports from selected tasks and environments, not independent replications. A 66% result for each step is not the probability that every step in a long task succeeds. The public article also does not provide enough detail about the comparator, failure definition and trial counts to generalize the ratio. A factory demonstration is not proof of safety across every disturbance.
Deployments should separately test full-task completion, recovery time, unseen objects, lighting and camera angles, force control and human proximity. Rights to demonstration video and operational data, emergency stops and human approval remain part of the system evaluation.