
To operate reliably in dynamic real-world settings, robots should be able to acquire new skills quickly without undergoing extensive additional training. Most existing robotic systems, however, primarily perform well on the tasks that they were trained to complete.
Teaching robots to tackle new tasks can be both time-consuming and costly. Moreover, additional training sometimes hinders their performance on previously learned tasks.
Researchers at Beijing Institute of Technology, X SQUARE ROBOT and Tsinghua University recently developed HOST (Human-to-robot One-Shot Skill AcquisiTion), a new framework that could allow robots to acquire new skills faster and more efficiently. Their proposed learning approach, introduced in a paper posted to the arXiv preprint server, allows a robot to acquire a new skill from a single video showing a human demonstration without compromising previously acquired abilities.
“The ability to acquire skills rapidly and effortlessly while retaining those already mastered is essential for robots,” wrote Guangyan Chen, Meiling Wang and their colleagues in their paper.
“However, current methods still rely on a cumbersome training-time loop that is costly and slow, while eroding skills already mastered. We introduce HOST, a framework that enables a robot to acquire skills in seconds from a single human video while retaining previously mastered skills.”
How HOST turns human demonstrations into robot actions
The framework developed by Chen, Wang and their colleagues allows robots to acquire a new skill simply by analyzing a video showing a human demonstration. This is achieved in three main stages.
First, HOST uses images captured by cameras to determine how far the robot has progressed through the procedure shown in the human video. In other words, it identifies the part of the demonstration that corresponds to the robot’s current stage.
The framework then examines the upcoming part of the human demonstration and predicts what the robot itself should expect to observe as it advances through the task. This intermediate step helps HOST account for differences between the human demonstrator’s body and the robot’s physical structure.
Finally, HOST derives the actions the robot should perform from these predicted future observations. By repeatedly determining its current progress, anticipating what it should observe next and selecting corresponding actions, the robot can follow the demonstrated procedure while adapting it to its own body structure.
Although HOST acquires a new skill from a single video without updating its parameters, the underlying framework was trained beforehand using 193,462 robot trajectories spanning 229 tasks. It was subsequently adapted using 5,847 human demonstration videos, each paired with robot trajectories for the demonstrated tasks.
“HOST resolves skill acquisition through a cascade of self-grounded prediction. It first estimates the robot’s progress within the demonstrated task, then translates the upcoming progression into the robot’s own future observations and finally derives actions from these predicted observations,” wrote the authors.
“This cascade is trained on targets coupled to the video demonstration, obtained by mapping the robot trajectory and the video demonstration onto a shared task progress manifold, then redefining each target to align with the future progression of the video. HOST thereby enables the robot to actively follow the demonstrated procedure and adapt it to the robot’s embodiment.”
A possible route to faster robot skill acquisition
The researchers evaluated their framework in real-world tests with a two-armed robotic system. This system consisted of two ARX R5 six-axis robotic arms, each equipped with a parallel-jaw gripper, and three RGB cameras.
The robot was tested on 50 previously unseen physical manipulation tasks that involved different objects and tools and required distinct movements. It attempted each task 20 times, with objects placed in different starting positions or at varying orientations each time. Human evaluators determined whether each attempt could be deemed successful.
“HOST acquires novel skills at inference time from a single human video in an average of 29 seconds and achieves a 62% average success rate,” wrote Chen, Wang and their colleagues. “It exceeds the zero-shot baseline by 45% while retaining previously mastered skills. HOST even exceeds the baseline fine-tuned on 50 robot demonstrations per task while requiring 50 times fewer demonstrations and acquiring each skill 507 times faster.”
The results of initial tests highlight HOST’s potential for teaching robots new skills faster and without extensive task-specific demonstrations. In the future, the framework could be refined further to improve its success rate and tested on other robots or in a wider range of scenarios. https://techxplore.com/news/2026-08-robots-skills-video-seconds.html





Recent Comments