Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video



Manufacturing floors, warehouses and production lines rarely stay fixed — tasks change, layouts shift and new products arrive, and most robots can’t keep up without significant reprogramming.

Skild AI’s new S1 robot foundation model helps address this, designed to learn previously unseen, long-horizon tasks from a single video demonstration. The model, launched last week, uses video as input to understand and execute the task without updating its weights or undergoing task-specific post-training — a technique called in-context learning.  

Skild built S1 and conducted the research on NVIDIA AI infrastructure, part of a broader collaboration spanning synthetic data generation, model training, simulation and real-world physical AI deployment. The companies are working together to move adaptable robot intelligence from the lab into factories and other dynamic operating environments.

“Learning by experience, and not preprogramming, is the step change that has happened in robotics,” said Deepak Pathak, cofounder and CEO of Skild AI. “NVIDIA Isaac Lab and NVIDIA Cosmos technologies help Skild create the scalable, diverse experience its robots need to learn across many scenarios and embodiments.”

The launch comes as the company reached a $100 million annual revenue run rate 10 months after its first commercial deployment. In that time, Skild has built more than 60 deployment partnerships with work spanning manufacturing, logistics, inspection, security, food preparation and other applications. 

Learning New Work From One Video

Most industrial robots are built for fixed jobs, so each new product, process or layout requires more data, retraining and validation.

S1 takes a different approach: An operator records a video of the desired task and provides it to the model as a prompt. It interprets the demonstrated intent, objects and sequence, then maps them into actions for the robot in front of it — with no retraining — and often for a task not covered by its pretraining dataset. 

S1 can perform unfamiliar tasks lasting up to 10 minutes, including plant potting, pancake making, pour-over coffee brewing and kit assembly. These tasks can span dozens of manipulation steps and require the robot to compose skills in sequences it hasn’t previously performed. 

 

In one plant-potting test, the Skild AI team moved from recording the demonstration to autonomous execution on hardware in just 11 minutes. The model can also adjust when objects move, recover from errors and combine skills in sequences that weren’t explicitly programmed.

In Skild’s tests on new, multistep tasks, its S1 robot succeeded about 66% of the time at each step, compared with 9% for a similar AI system — a more than sevenfold improvement. Skild also estimates that showing the robot one short video example can be as useful as giving it roughly 380 hands-on training examples. A person collecting those examples manually could take 50-100 hours.

From Research to Factory Work

S1 breaks the cycle of needing to constantly retrain robots for new factors by letting operators demonstrate new tasks directly without requiring a new dataset or training run for every change. Where customer agreements permit, experience from Skild’s commercial deployments can inform the broader model and help accelerate future deployments.

That work is already in action on the factory floor. Skild, NVIDIA and Foxconn are deploying the Skild Brain on dual-arm manipulators for high-precision assembly of NVIDIA Blackwell systems. In one demonstrated workflow, a robot installs a busbar and limit block, fastens 16 screws and adapts to disturbances across a multistep task. The work requires precise motion, contact-aware control, sequence tracking and recovery when the scene differs from the plan.

 

NVIDIA Technology Across the Development Cycle

NVIDIA accelerated computing gives Skild the scale to train its shared robot brain using simulation, human video, teleoperation and, where permitted, deployment data. NVIDIA Cosmos open world foundation models help diversify training data and turn video into structured descriptions, while Cosmos Curator helps annotate, filter and organize data at scale.

Skild is extensively using NVIDIA’s open simulation frameworks to train and validate its robot brain before real-world deployment. NVIDIA Omniverse libraries and the NVIDIA Isaac Sim framework provide physically based virtual environments for generating data, testing edge cases and validating behaviors.

Skild further strengthens the skills of its brain through reinforcement learning in Isaac Lab, an open modular robot learning framework. Powered by the Newton physics engine, Isaac Lab helps Skild’s engineers accurately model various physical parameters, such as forces, contact, collision and pressure, and reduce the simulation-to-reality gap.

Skild and NVIDIA are also jointly developing new GPU-accelerated simulation solvers that quickly and accurately model how robots physically touch, grip and manipulate solid objects. They’ll soon be made available to all developers as part of Newton. 

As models move toward production, NVIDIA Nsight tools help engineers find performance bottlenecks during training, and the NVIDIA TensorRT software development kit optimizes inference so robots can respond quickly in the physical world. Together, these technologies connect the data, simulation, training and deployment stages instead of treating them as separate systems.

Read Skild AI’s S1 research and explore the NVIDIA Isaac robotics platform.

NVIDIA Blackwell Leads on First Agentic AI Infrastructure Benchmark


AgentPerf from Artificial Analysis, the industry’s first agentic AI benchmark, gives developers, enterprises and infrastructure providers a clear way to compare systems for agentic AI. In the first round of published results, the NVIDIA Blackwell Ultra NVL72 platform delivers leading performance across the agentic AI workloads tested, running 20x more agents per megawatt than NVIDIA Hopper.

Agentic AI is a fundamentally different workload than conversational AI. A single chat completion is a sprint: one large language model (LLM) call, one response. An agent functions more like a relay: It breaks a goal into many steps and keeps going until the task is done. 

Agents chain together multiple LLM calls and tool calls to gather context, observe, reason and act.

That results in dozens to hundreds of LLM calls chained together, each passing growing context to the next, with tool calls like code compile and execution, database search and web browsing at every handoff. The complexity isn’t additive; it’s multiplicative. 

The distinction matters enormously for performance measurement. Existing AI inference benchmarks measure one LLM call: how fast an LLM responds to a single request and how many simultaneous requests a system can handle. They weren’t designed for agentic workloads, where chained LLM calls, tool call delays and growing context stress accelerated computing systems in fundamentally different ways than a single LLM call ever could. 

For companies building and deploying agents at scale, it’s important to understand how responsive agents are, how many can be deployed simultaneously and how much useful work AI infrastructure can deliver for every dollar and watt invested.

NVIDIA GB300 NVL72 Runs 20x More Agents per Megawatt

In this first round, AgentPerf measures agentic performance with DeepSeek V4 Pro, a large mixture-of-experts (MoE) model that represents the class of frontier models powering today’s most capable agents. On this workload, NVIDIA GB300 NVL72 delivers the highest performance in the benchmark, running up to 20x more agents per megawatt than the NVIDIA HGX H200 system.

NVIDIA GB300 NVL72 supports far more concurrent agents per megawatt than NVIDIA H200 at both service-level objectives of 20 and 60 tokens per second per agent.

The performance advantage comes from extreme codesign across the full stack. GB300 NVL72 connects 72 GPUs into a single rack-scale system, enabling large MoE models like DeepSeek V4 Pro to distribute model execution efficiently at scale. 

CUDA kernels accelerate this further by overlapping communication and compute, so the cost of coordinating across experts is absorbed rather than added to latency. 

NVIDIA TensorRT LLM sustains efficiency as concurrent agent sessions scale. For example, it separates the processing of inputs from the generation of outputs so each can be optimized independently. 

These results are grounded in a benchmark methodology built from the ground up to reflect how agentic AI actually works in production.

Artificial Analysis AgentPerf: Built on Real-World Agentic Workloads

AgentPerf is built based on real coding agent trajectories: an agent receives a task, reads files, writes and edits code, executes commands and iterates based on the results — all drawn from real public code repositories across 12+ programming languages. The long sequence lengths, tool call patterns and delays are all representative of real-world coding workflows. 

AgentPerf then measures how many of these agentic tasks a platform can support simultaneously while meeting defined performance thresholds for responsiveness and output token rate. Tool calls are not executed but simulated using representative CPU processing time, so differences in results reflect accelerated computing performance only. 

The results translate directly into infrastructure decisions: how many concurrent agentic tasks can be run per accelerator and per megawatt of power. For enterprises deploying AI agents at scale, those numbers determine how much productive work a given infrastructure investment can actually deliver.

NVIDIA Ecosystem Partners Harness Blackwell’s Leading Performance

Leading inference providers including Baseten, DeepInfra and Together AI are already serving agentic workloads on frontier models such as DeepSeek V4 Pro on NVIDIA Blackwell and powering production agentic applications today. 

Together AI powers real-time inference for Cursor, an AI-powered agentic coding platform, on NVIDIA Blackwell. Cursor’s agents debug issues, generate features and execute refactors while developers continue working.  

DeepInfra powers Pam.ai, an AI workforce platform for car dealerships, which deploys agents to book service appointments, handle calls and run outbound sales campaigns, entirely on NVIDIA Blackwell. 

As NVIDIA and the open source ecosystem continue to optimize inference software, performance and efficiency on agentic workloads will only improve. The NVIDIA Vera Rubin architecture is now in full production, bringing the next generation of infrastructure capacity to meet the growing demands of agentic AI at scale. 

Dive deeper into AgentPerf’s methodology and NVIDIA’s full-stack optimizations for agentic AI in this technical blog.