A new robotic system links speech, gestures, motion and safety controls, helping humanoid robots respond more naturally during real-time face-to-face conversations for expressive, context-aware interactions.

Galbot has developed RoboGesture, a system designed to help humanoid robots generate natural gestures during conversations by connecting spoken language with physical movement. The technology combines speech processing, artificial intelligence, motion generation and safety controls to help robots respond with movements that better reflect the meaning, rhythm and tone of speech.
Unlike systems that rely mainly on written speech transcripts, RoboGesture analyses audio directly. Changes in pitch, rhythm and intonation can provide important contextual information, allowing the robot to determine when a gesture is appropriate and what type of movement should accompany the response.
The system first converts incoming speech into audio features using the Mimi audio codec. It then processes these signals at different levels. Lower-level features capture rapid changes in sound and rhythm, while deeper features provide broader semantic information. Researchers trained the system across 300 gesture categories to give it a wider range of movements.
RoboGesture sends these signals to a motion-generation model based on a diffusion transformer. Instead of producing isolated actions, the model operates in a continuous motion space, allowing gestures to flow from one movement to another while remaining connected to the conversation.
Safety has also been incorporated into the system. An Anti-Inertia classifier helps prevent the motion model from becoming overly dependent on previous movements and repeating similar actions. Generated movements are then passed through a model predictive control safety filter, which checks motion constraints and adjusts movements before execution. Testing reduced the self-collision frame rate from 4.16% to 0.13%.
The system has been tested on a Unitree G1 humanoid robot equipped with dexterous hands. Its interaction pipeline combines speech recognition, language processing, text-to-speech, audio analysis, gesture generation and safety filtering.
Developed by researchers from Galbot and several Chinese universities and institutes, RoboGesture could make humanoid robots more expressive in face-to-face settings. The work is particularly relevant to service robots, where natural communication depends on both spoken language and physical cues.







