Nabugu - stock.adobe.com
Physical AI's challenge: Making humanoid robots work in the real world
The next stage of physical AI will depend on improving human-robot interactions.
CAMBRIDGE, ENGLAND -- Cambridge Consultants, the deep tech arm of French IT consulting firm Capgemini, opened its Cambridge Science Park labs to the media this week, showcasing the technologies it sees as underpinning the next generation of industrial robots.
Among the demos on Monday were two humanoid robots that the company has modified with its own hardware and software. One had been fitted with a new head and backpack housing additional compute, connectivity, and a safety kill switch. The other had been given new hands, designed to improve its ability to pick up and place down boxes for factory and warehouse applications.
The centerpiece, however, was a new robotic awareness platform built on top of foundation models from Nvidia and Jeff Bezos and OpenAI-backed Physical Intelligence, demonstrating a less obvious challenge for physical AI: teaching robots to understand the people and environments around them.
Teaching robots to read a room
A key feature of the new system is gesture recognition, which enables a robot to understand and respond to human gestures and movements. For example, a user can point to the object they want moved.
There’s also a speech version where a user can say “move this here,” and the robot infers the reference.
"Matching together speech, gesture, pose, spatial understanding and object recognition allows a much more human-like interaction between a human and robot," Tim Ensor, head of AI at Cambridge Consultants, told TechTarget AI and Emerging Tech. "You can refer to objects in the room through pointing and through ambiguous references."
That kind of interaction, Ensor argued, is becoming increasingly necessary as robots move out of controlled demonstrations and into environments in which people are unavoidably part of the system.
"The utility of your robot is limited if it can't interact with humans who are unpredictable," he said.
Factories offer a relatively structured setting for early deployments (more so than domestic settings), but even there, workers are the unpredictable variable robots will have to contend with.
Robot-human interaction is also important for training machines to perform novel tasks. As robots become more general-purpose, they may need to perform multiple jobs rather than being tied to a single workflow, meaning humans will need ways to instruct and train them.
“In the course of getting a robot to a general-purpose level, somebody's got to show them how to do those tasks,” Ensor said. “Even though it may not have to interact with humans constantly during a workflow, there still has to be a way for these robots to interact effectively with humans.”
The deployment problem
Vision-language-action models, which connect language and visual understanding to robot movements, have become a leading approach to making robots more adaptable.
Ensor’s team has been experimenting with systems that can learn from video and combine that understanding with language, enabling a robot to effectively anticipate what a described movement should look like before attempting it.
But the technology is not yet ready to go straight from the lab floor to an industrial setting.
"You can do that for a robot, and it will be able to perform the task at a certain speed and with a certain percentage reliability," Ensor said. "But we tend to find that they're not actually fast or reliable enough to go straight into an industrial deployment."
That has pushed the company toward what Ensor called “non-core functional components”: the less glamorous engineering required to turn increasingly capable AI models into usable products.
That includes running large models efficiently on smaller compute platforms, building self-learning loops that enable robots to improve through repeated attempts with some human intervention, and developing validation and verification processes to demonstrate that systems are reliable and safe.
The result is a full-stack system, encompassing the computing architecture fast enough to run the models in real time, a knowledge layer that gives a system a genuine understanding of its environment, and an intelligence layer capable of acting usefully on that information.
While the industry is moving rapidly (with Ensor projecting commercial deployments of real robots doing “useful work” within the next year), it is still in its early stages.
“This is an exciting area that we’re really only at the beginning of,” Ensor said. “The real promise of physical AI is a transformation of manual labor to a similar extent that we're seeing digital work being transformed by agentic AI. We’re really only going to see growth towards this goal from here.”
Scarlett Evans is a freelance writer with a focus on emerging technologies and the minerals industry. Previously, she served as assistant editor at IoT World Today, where she specialized in robotics and smart city technologies. Scarlett also has a background in the mining and resources sector, with experience at Mine Australia, Mine Technology and Power Technology.