A Satellite Challenge of IEEE ICASSP 2027

RoboHEARD 2027

Robot-centric Embodied Hearing and Dialogue for
Multi-Speaker Audio-Visual Speech Recognition

Who spoke, when they spoke, and what they said — in real multi-speaker human-robot interaction, perceived through the sensors of a real mobile robot.

News

Introduction

As service robots are increasingly deployed in applications such as companionship, elderly care, domestic assistance, navigation, and shopping guidance, natural and robust human-robot interaction (HRI) is becoming increasingly important. RoboHEARD 2027 focuses on multimodal perception for mobile service robots in two representative scenarios: home companionship and navigation / shopping assistance.

The core task is multi-speaker audio-visual speech recognition in real robot-centered interaction environments. During robot movement and service delivery, participants use data collected from the robot's onboard multimodal sensors — including microphone arrays and multiple cameras with different viewing angles — to determine who spoke, when they spoke, and what they said in multi-speaker conversational scenes.

What makes this challenge new

Real Sensor Data

Data are collected using real sensors mounted on actual mobile robots in real-world service scenarios, rather than simulated data or wearable devices.

Heterogeneous Multimodal Fusion

Inputs come from heterogeneous robot-mounted sensors: microphone arrays and multi-view cameras — richer but more complex perception cues.

Spontaneous, Overlapping Speech

Realistic multi-speaker human-robot interaction — home companionship and navigation/shopping — where conversations are spontaneous, overlapping, and context-dependent.

Dynamic Positions & Ego-Noise

Both speakers and the robot continuously move, changing positions, orientations, and spatial relationships. Robot motion introduces ego-noise — motor noise and motion-related acoustic interference.

Data

The dataset is collected from the QUANTA X2 robot (or its newer model) by X Square Robot. Over 100 conversations were recorded in two scenarios — home companionship and navigation / shopping guidance. Each session lasts approximately 5–8 minutes and involves 3–6 human participants who move along with the robot following predefined scripts.

The QUANTA X2 robot
Fig 1. The QUANTA X2 robot.

Corpus Overview

In total, the speech corpus amounts to about 60 hours of audio, divided into strict speaker- and environment-disjoint splits:

30h
Training
10h
Development
20h
Test

Sensor Specifications

ModalityConfiguration
Audio Multi-channel audio (2-mic or 6-mic configurations)
Video Multi-view cameras

Tracks

Track 1 Offline

Multi-speaker time-stamped AVSR (offline)

Recognize who spoke, when they spoke, and what they said in multi-speaker conversational scenes from the robot's onboard sensors, under the offline mode with full-utterance access.

  • Full-context inference, no latency constraint
  • Metric: time-constrained minimum permutation CER (tcpCER)
MetrictcpCER

Track 2 Online

Multi-speaker time-stamped AVSR (online)

Recognize who spoke, when they spoke, and what they said in multi-speaker conversational scenes from the robot's onboard sensors, under the online (streaming) mode.

  • Endpointing Delay (P95) limited to 2 seconds on a single RTX 5090 GPU
  • Metric: tcpCER (only latency-compliant solutions are ranked)
MetrictcpCER + ED (P95) ≤ 2s

Baselines

The organizers will provide an open-source baseline system for each track, released together with the development data on October 15, 2026.

Evaluation

tcpCER

Time-constrained minimum permutation Character Error Rate. The primary metric for both tracks, measuring how accurately the system recognizes who spoke, when they spoke, and what they said as time-stamped multi-speaker transcriptions.

Endpointing Delay (P95) (Track 2)

A latency metric for streaming ASR measuring system responsiveness at the end of each utterance. The 95th percentile of the difference between the system's predicted final-word timestamp and the reference end timestamp (via forced alignment).

Track 2 solutions must achieve an Endpointing Delay (P95) ≤ 2 seconds to be eligible for ranking.

Latency Validation (Track 2)

Latency is validated on an in-house NVIDIA RTX 5090 server from December 7–22, 2026, after the leaderboard freeze. By default, teams submit a Docker container exposing the standardized streaming API, which the organizers run on the 5090 server.

  • Validation is mandatory for the top-5 performing teams; other teams are encouraged to participate.
  • Teams may alternatively provide a hosted API endpoint (on our server or their own). For ranking eligibility, the endpoint must report per-request timing so network overhead can be separated from model latency, and must declare its hardware.
  • Inference models/codes are removed after validation and used only for this purpose.

Leaderboard

We will use Codabench as the main host for results submission and leaderboard demonstrations. The leaderboard will open after the release of the test sets on December 1, 2026.

Leaderboard (opening soon)

Timeline

All times follow the official challenge announcement unless stated otherwise.

  • July 27, 2026
    Proposal Submission Deadline Completed
  • August 21, 2026
    Proposal Acceptance Notification Completed
  • September 15, 2026
    Launch of the Challenge Completed
  • October 15, 2026
    Development Data & Baseline Systems Release
  • November 15, 2026
    Registration Deadline · List of Allowed Data & Pre-trained Models Released
  • December 1, 2026
    Test Sets & Leaderboard Release
  • December 7, 2026
    Leaderboard Freeze · Latency Test Opens
  • December 15, 2026
    Technical Report Submission
  • December 22, 2026
    Latency Test Submission (Track 2)
  • December 30, 2026
    Official Ranking Released
  • January 07, 2027
    2-page Papers Due (by invitation only)
  • January 21, 2027
    2-page Paper Acceptance Notification
  • January 28, 2027
    Camera-ready 2-page Papers Due

Registration

Please register via the Google Form below. We will send you a successful registration email, along with dataset access, shortly after registration.

Registration Form

Data & Model Declaration

Participants must declare the datasets and models/tools they intend to use to the organizers before November 15, 2026 (the registration deadline). All data sources and checkpoints must be clearly documented in the system description with appropriate citations or links.

After collecting this information, we will compile and release a full list to all participants to ensure transparency and a fair, reproducible evaluation environment.

Rules

Training Data & Pre-trained Models

Submission

TrackModeKey Requirement
Track 1OfflineFull-context inference; no latency constraint
Track 2OnlineStreaming; Endpointing Delay (P95) ≤ 2 seconds on a single RTX 5090 GPU

Verification of the latency requirement for the online track is mandatory for the top-5 performing teams. By default, teams submit a Docker container exposing the standardized streaming API, which the organizers run on a provided password-enabled GPU server; alternatively, teams may provide a hosted API endpoint with per-request timing and a hardware declaration.

FAQ

Participants may use any openly available datasets for model training, as well as pre-trained models, provided they are openly and freely accessible and are reported to the organizers before November 15, 2026. All data sources and checkpoints must be documented in the system description.

Yes. The challenge offers two tracks (offline and online). Participants may submit results for one or both tracks.

Latency is validated on an in-house NVIDIA RTX 5090 server using the Endpointing Delay (P95) metric from December 7–22, 2026, and is mandatory for the top-5 teams. By default, teams submit a Docker container exposing the standardized streaming API, which the organizers run on the server; alternatively, teams may provide a hosted API endpoint (on our server or their own). For ranking eligibility, hosted endpoints must report per-request timing so network overhead can be separated from model latency, and must declare their hardware. Submitted code is removed after validation.

We use Codabench as the main host for results submission and leaderboard demonstrations.

Awards

Each track awards cash prizes to the top three teams, with official certificates presented during the ICASSP 2027 Challenge Session.

1st Place
$1,000
2nd Place
$500
3rd Place
$300

Sponsor

The RoboHEARD 2027 Challenge is sponsored by X Square Robot, a robotics company dedicated to building embodied intelligence. As our sponsor, they open the QUANTA X2 robot platform used to collect the RoboHEARD 2027 dataset, support the cash awards for the winning teams, and provide NVIDIA 5090 GPU servers for the online latency testing.

Organizers

Chinese University of Hong Kong, Shenzhen
Chinese Academy of Sciences
The Hong Kong Polytechnic University
National Taiwan University
Academia Sinica
AISHELL
Duke Kunshan University
Duke Kunshan University
Tiantian Feng (starting from 2027)
Chinese University of Hong Kong, Shenzhen
X Square Robot
X Square Robot
X Square Robot

Volunteers

The RoboHEARD 2027 Challenge is supported by the following volunteers.

Wang Xiang
Fei Su
Peijun Yang
Dong Liu
Xin Xu

Contact

For any questions or inquiries, please feel free to reach out to us.