Skip to Main Content

Humans are difficult to simulate. We are quick, complex thinkers, and our exact actions aren’t predictable. So how can we sufficiently train robots to collaborate with us?

Traditionally, humans are treated as a fixed input in simulation-based learning environments. This means that they perform consistent, predictable actions that aren’t truly representative of human behavior. Hao Zhang, a post-doctoral research fellow in Ding Zhao’s Safe AI lab, is working toward better human-robot collaboration through HALO, or heterogeneous agent Lyapunov policy optimization. HALO is a framework that uses multi-agent reinforcement learning to help robots independently learn how to interact and collaborate with humans.

\

During simulation, the methodology represents the human partner with a learning-capable robot, so both parties are learning how to interact with each other from scratch, with no fixed input or scripted behavior. “Effective human-robot collaboration should reduce, rather than create, additional training requirements for people,” says Zhao, and associate professor of mechanical engineering. “Human partners naturally differ in how they move, communicate, and make decisions, so asking everyone to follow a prescribed interaction pattern would limit the usefulness of the robot.” Gradually, they learn how to collaborate with each other. Now, when a human exhibits an unusual or unexpected behavior, the robot has already learned how to adapt and react to unfamiliar movements.

Rather than training a person to behave in a particular way for the robot, we train the robot in an open-ended interaction setting where different partner behaviors can emerge.

Hao Zhang, Postdoctoral Research Associate, Mechanical Engineering

In a rescue mission, or in a hospital room, it’s crucial that transportation and other manual labor can be carried out safely and reliably. If robots are transferring a patient in a hospital, they need to be able to accurately respond to each other and everyone else involved to ensure safe and efficient care. “When you transport something, you are physically connected to your partner by the object,” says Zhang. “Even small mismatches in timing or motion can cause the agents’ learning processes to interfere with each other, making effective collaboration difficult to learn. Stability is critical during the learning process.”

The researchers have demonstrated that using HALO, these robots can navigate around various furniture and walls while carrying large, unwieldy objects. They can respond quickly and accurately to shifts in weight or unexpected movements from a human counterpart, adjusting their stance to better distribute the weight it’s carrying or pivoting more or less than they originally predicted to account for the human’s path.

Related video

Watch the video about the related APEX humanoid robotic system.

A problem that arises during this learning process, however, is a rationality gap. Both parties share the same goal, but because they learn independently, each may develop a different idea of how best to collaborate. A strategy that appears reasonable from one party’s individual perspective may not align with what would work best for the team as a whole. There are multiple ways to move a couch around a corner, for example. One party might try to turn the couch onto one of its arms so that it stands vertically, while the other is working to keep the couch horizontal and move it forward. Though they are both trying to eventually move in the same direction with the same object, the task is being prolonged as their method interferes with their counterpart’s.

To combat this, Zhang is applying the Lyapunov stability theory. The theory is a safety guarantee that ensures the learning processes of each party can converge and reroute to a trajectory that is optimal for the team as a whole. In other words, each party would observe and analyze the other’s movements to be able to adjust their behavior to align with each other.

Implementation in the lab began on computers with two robots, like most other simulations. This was to eliminate the problem of improbable scripted human behavior. Once they could smoothly collaborate with each other, one robot was substituted for a real human. “In HALO, human variability becomes part of the robot’s learning process, shifting the burden of adaptation from the human to the robot,” Zhao explains.

As they continue to demonstrate the scalability of their methodology, they plan to begin introducing increasing numbers of robots and humans that interact with each other, with up to four robots collaborating with each other by September.

This work from the Safe AI lab is conducted in collaboration with Professor H. Eric Tseng at the University of Texas at Arlington.