A glass wall can look like an open path to a robot. Reflections and transparency can fool its sensors, leaving it with a faulty map—and a route that could end in a collision.
University of Maryland computer science doctoral student Naitri Rajyaguru and collaborators are helping robots recognize those barriers using information ordinary cameras miss: the polarization of light.
Their work earned two honors at workshops associated with the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), held Sept. 27–Oct. 1 in Pittsburgh. In all, UMD’s Perception and Robotics Group received five honors across five workshops at the flagship robotics conference.
The honors recognized research led by Rajyaguru and fellow computer science doctoral students Zhi “Leo” Wang and Amir-Hossein Shahidzadeh. Their three projects address how robots perceive their surroundings, learn from people and track an action’s progress.
The students are advised by the group’s co-directors, Yiannis Aloimonos, a professor of computer science, and Cornelia Fermüller, a research scientist in the University of Maryland Institute for Advanced Computer Studies (UMIACS). Aloimonos also holds an appointment in UMIACS, which provides administrative and technical support for the group.
“Our group has been at the forefront of research connecting language, perception and robotic action for many years,” said Aloimonos. “What makes this latest recognition especially meaningful is that it highlights the emerging researchers in our lab. These students are bringing fresh ideas to longstanding challenges and helping shape the next generation of intelligent robots.”
Rajyaguru led “PolarDepth: Polarization-Guided Monocular Depth for Visual Odometry,” which earned the Best Poster Award at the Insect-scale Autonomy Workshop and Best Poster Runner-Up recognition at the Bridging Perspectives in Navigation: Lessons Learned and Future Directions Workshop.
PolarDepth improves distance estimates and movement tracking in spaces lined with glass. Rajyaguru and co-authors explain that polarization—the orientation of light’s oscillations—provides clues to surface geometry that conventional images can miss.
Using a single polarization camera, the system combines color imagery with polarization cues. It learns which is more reliable in different parts of an image, drawing on polarization around reflective surfaces and color imagery elsewhere.
Tests on real and synthetic scenes showed better depth estimates and movement tracking in glass-walled corridors than approaches using color images alone. That could help robots navigate buildings filled with glass doors, walls and partitions.
The Perception and Robotics Group’s other recognized projects address learning and interacting with objects.
Wang led “HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos,” which received the Best Paper Award at the WORLDS: World Models and Spatial Intelligence for Physical AI Workshop and the Best Demo Paper Award at the Building Scalable Infrastructure for Robot Learning: From Data Scaling to Real-World Deployment Workshop.
HumanEgo helps robots learn tasks from first-person videos of human demonstrations, overcoming differences in how human hands and robot grippers look and move. With just 30 minutes of video per task, it achieved an average success rate of 92.5% across four tasks and transferred to different robots and environments without retraining.
Previously featured by UMIACS, the approach could make training data easier to collect by letting people demonstrate tasks themselves instead of remotely controlling a robot.
Shahidzadeh led “What to Attend, What to Keep: Skill-Conditioned Visuotactile Representation with Progress-Guided Event Memory,” which received the Best Paper Award at the Touch-to-Action: Enhancing Robot Manipulation through Tactile Perception Workshop.
The research examines how robots combine sight and touch to recognize a task’s progress. A robot screwing a light bulb into a socket, for example, must know when to stop. Successive turns can look almost identical, but increasing resistance signals that the bulb is fully tightened.
Different stages call for different sensory information: vision can guide reaching toward an object, while touch can reveal whether a grasp is secure or a movement is complete.
The team developed SkillFormer, a model that uses a language description of a skill to identify relevant sensory information and retains key observations from earlier steps. They tested it using demonstrations of lamp replacement, bottle capping and a search for a cube among boxes containing different shapes.
The cube search shows why memory matters. Touch helps identify a hidden object’s shape, while memory preserves which boxes have been searched and where the cube was found.
SkillFormer estimated task progress more accurately and detected key contact events sooner than comparison methods. During bottle capping, it reduced slip-detection delay by 87% compared with the progress-estimation models used as benchmarks.
A slip may last only a moment, then disappear from a system’s recent observations. SkillFormer retains that evidence as the task continues—a step toward helping robots recognize when to keep moving, when to adjust and when the job is done.
—Story by UMIACS communications group