Who spoke, when they spoke, and what they said — in real multi-speaker human-robot interaction, perceived through the sensors of a real mobile robot.
As service robots are increasingly deployed in applications such as companionship, elderly care, domestic assistance, navigation, and shopping guidance, natural and robust human-robot interaction (HRI) is becoming increasingly important. RoboHEARD 2027 focuses on multimodal perception for mobile service robots in two representative scenarios: home companionship and navigation / shopping assistance.
The core task is multi-speaker audio-visual speech recognition in real robot-centered interaction environments. During robot movement and service delivery, participants use data collected from the robot's onboard multimodal sensors — including microphone arrays and multiple cameras with different viewing angles — to determine who spoke, when they spoke, and what they said in multi-speaker conversational scenes.
Data are collected using real sensors mounted on actual mobile robots in real-world service scenarios, rather than simulated data or wearable devices.
Inputs come from heterogeneous robot-mounted sensors: microphone arrays and multi-view cameras — richer but more complex perception cues.
Realistic multi-speaker human-robot interaction — home companionship and navigation/shopping — where conversations are spontaneous, overlapping, and context-dependent.
Both speakers and the robot continuously move, changing positions, orientations, and spatial relationships. Robot motion introduces ego-noise — motor noise and motion-related acoustic interference.
The dataset is collected from the QUANTA X2 robot (or its newer model) by X Square Robot. Over 100 conversations were recorded in two scenarios — home companionship and navigation / shopping guidance. Each session lasts approximately 5–8 minutes and involves 3–6 human participants who move along with the robot following predefined scripts.
In total, the speech corpus amounts to about 60 hours of audio, divided into strict speaker- and environment-disjoint splits:
| Modality | Configuration |
|---|---|
| Audio | Multi-channel audio (2-mic or 6-mic configurations) |
| Video | Multi-view cameras |
Recognize who spoke, when they spoke, and what they said in multi-speaker conversational scenes from the robot's onboard sensors, under the offline mode with full-utterance access.
Recognize who spoke, when they spoke, and what they said in multi-speaker conversational scenes from the robot's onboard sensors, under the online (streaming) mode.
The organizers will provide an open-source baseline system for each track, released together with the development data on October 15, 2026.
Time-constrained minimum permutation Character Error Rate. The primary metric for both tracks, measuring how accurately the system recognizes who spoke, when they spoke, and what they said as time-stamped multi-speaker transcriptions.
A latency metric for streaming ASR measuring system responsiveness at the end of each utterance. The 95th percentile of the difference between the system's predicted final-word timestamp and the reference end timestamp (via forced alignment).
Latency is validated on an in-house NVIDIA RTX 5090 server from December 7–22, 2026, after the leaderboard freeze. By default, teams submit a Docker container exposing the standardized streaming API, which the organizers run on the 5090 server.
We will use Codabench as the main host for results submission and leaderboard demonstrations. The leaderboard will open after the release of the test sets on December 1, 2026.
All times follow the official challenge announcement unless stated otherwise.
Please register via the Google Form below. We will send you a successful registration email, along with dataset access, shortly after registration.
Participants must declare the datasets and models/tools they intend to use to the organizers before November 15, 2026 (the registration deadline). All data sources and checkpoints must be clearly documented in the system description with appropriate citations or links.
After collecting this information, we will compile and release a full list to all participants to ensure transparency and a fair, reproducible evaluation environment.
| Track | Mode | Key Requirement |
|---|---|---|
| Track 1 | Offline | Full-context inference; no latency constraint |
| Track 2 | Online | Streaming; Endpointing Delay (P95) ≤ 2 seconds on a single RTX 5090 GPU |
Verification of the latency requirement for the online track is mandatory for the top-5 performing teams. By default, teams submit a Docker container exposing the standardized streaming API, which the organizers run on a provided password-enabled GPU server; alternatively, teams may provide a hosted API endpoint with per-request timing and a hardware declaration.
Each track awards cash prizes to the top three teams, with official certificates presented during the ICASSP 2027 Challenge Session.
The RoboHEARD 2027 Challenge is sponsored by X Square Robot, a robotics company dedicated to building embodied intelligence. As our sponsor, they open the QUANTA X2 robot platform used to collect the RoboHEARD 2027 dataset, support the cash awards for the winning teams, and provide NVIDIA 5090 GPU servers for the online latency testing.
The RoboHEARD 2027 Challenge is supported by the following volunteers.
For any questions or inquiries, please feel free to reach out to us.