Nvidia-backed Skild AI teaches robots new tasks from a single video

The development adds to the rising momentum toward general-purpose robots and the real-world deployment of physical AI.

Skild, a startup developing general-purpose physical AI, is using Nvidia technology to create a robot foundation model that can learn new tasks from a single video.

The project is part of a broader industry push toward general-purpose robots, with researchers and developers seeking ways to train physical AI systems more quickly and move them from laboratory demonstrations to real-world deployments.

S1 foundation model

The Pittsburgh-based company said its robot foundation model is designed to learn previously unseen, multistep tasks from a single video demonstration.

The technology is already being deployed at an Nvidia factory in Houston, where Skild, Nvidia and contract electronics manufacturing giant Foxconn recently began using the Skild Brain omni-bodied foundation model on dual-arm robots to assemble Nvidia Blackwell GPU systems.

S1 was built on Nvidia AI infrastructure, using Cosmos, Isaac Lab, Isaac Sim and Omniverse across data generation, training, simulation and deployment. The model can also adjust when objects move, recover from errors and combine skills in sequences it was not explicitly programmed to perform.

Skild reached a $100 million annual revenue run rate earlier this month, just 10 months after its first commercial deployment. This, according to the company, is evidence that robot foundation models are moving beyond demos and into real work.

“S1 shows how quickly an operator could teach a robot new work,” Amit Goel, director of product management for autonomous machines at Nvidia, told TechTarget.com. “For manufacturers, the opportunity is to make automation more adaptable as products, components and processes change.”

In one experiment, Skild reported just 11 minutes elapsed between recording a plant-potting demonstration and the robot performing the task autonomously.

For the deployment building Blackwell systems, Goel said the project demonstrates the value of adaptable robot intelligence in a concrete manufacturing application.

“S1 is an encouraging step toward general-purpose physical AI,” he said.

The data dilemma

Despite physical AI’s promise, scaling the tech has its own challenges. According to Goel, data and evaluation remain the biggest bottlenecks.

“Collecting that experience on real robots is slow and expensive,” he said. “A robot needs to handle variations in objects, lighting, placement and physical contact, including situations that rarely occur during normal operation.”

Simulation and synthetic data can, he said, help developers expand that experience more quickly. Skild is using Nvidia's Isaac Sim and Isaac Lab, alongside its Omniverse and Cosmos technologies, across simulation, training and data generation.

But generating more training data is only part of the challenge. Developers also need to establish that models can perform reliably when they encounter the variability of physical environments.

“The goal is to connect training and testing to real operating conditions so that improvements translate into dependable performance on the factory floor,” Goel said.

Toward shared robot intelligence 

Looking ahead, Goel said Nvidia sees the industry moving toward an inflection point, when shared foundation models can provide intelligence for different types of robots.

"Skild AI’s work across robotic arms, quadrupeds and humanoids shows progress toward shared intelligence that can adapt to different robot bodies," he said. “The opportunity is to reuse what a model has learned across more machines and applications, reducing the development effort required for each new deployment.”

The next challenge is making that capability consistently useful in production. Different robots come with different physical capabilities, and each application has its own performance requirements.

A shared foundation model can provide a starting point, with integration, validation and, where needed, more training to prepare it for the specific job.

“From a user’s perspective, the meaningful milestone is dependable work across more tasks and more types of robots,” Goel added.

Scarlett Evans is a freelance writer with a focus on robotics and emerging technologies. Previously, she was assistant editor at IoT World Today, where she specialized in robotics and smart city technologies. Scarlett also has a background in the mining and resources sector, with experience at Mine Australia, Mine Technology and Power Technology. 

Dig Deeper on AI Infrastructure