Gering, P. orcid.org/0009-0009-5582-2718, Tian, R. orcid.org/0009-0009-5534-4832, Chardiwall, S. orcid.org/0009-0003-5741-6154 et al. (5 more authors) (2026) A multimodal cross-attention framework for robot error detection from bystander reactions. In: ICMI '26: Proceedings of the 28th International Conference On Multimodal Interaction. ICMI '26: International Conference on Multimodal Interaction, 05-09 Oct 2026, Napoli, Italy. ACM, pp. 1386-1390. ISBN: 9798400723186.
Abstract
To detect human-robot interaction errors, existing approaches have primarily focused on behavioural responses from individuals directly involved in the interaction, whereas reactions from external observers have received comparatively little attention. Similarly, these approaches have been limited to interactions in controlled environments with little variability in lighting, camera angles, and background noise. Addressing these gaps, the ERR@HRI 3.0 challenge focuses on bystander error detection in unconstrained, real-world environments. Our primary submission is a time-aware multimodal fusion framework that combines facial and acoustic representations via bidirectional cross-attention. Two versions of this framework, with different encoding approaches, are evaluated against a bidirectional LSTM (BiLSTM) classifier with simple concatenation. Our results demonstrate the effectiveness of this cross-fusion approach for capturing errors and generalising to unseen real-world data. Our cross-fusion and BiLSTM models achieved the top-three positions on Track 1, outperforming the challenge baseline and other competing submissions. Feature importance analysis further revealed that smile-related action units and their temporal dynamics serve as primary discriminative cues for error detection. Our code is available at https://github.com/chardiwall/SHEF-HRI.
Metadata
| Item Type: | Proceedings Paper |
|---|---|
| Authors/Creators: |
|
| Copyright, Publisher and Additional Information: | © 2026 Copyright held by the owner/author(s). This work is licensed under a Creative Commons Attribution 4.0 International License. https://creativecommons.org/licenses/by/4.0/ |
| Keywords: | human-robot interaction; bystander error detection; cross-attention; multimodal fusion |
| Dates: |
|
| Institution: | The University of Sheffield |
| Academic Units: | The University of Sheffield > Faculty of Engineering (Sheffield) > Department of Computer Science (Sheffield) |
| Date Deposited: | 07 Oct 2026 10:47 |
| Last Modified: | 07 Oct 2026 10:47 |
| Published Version: | https://doi.org/10.1145/3776574.3832482 |
| Status: | Published |
| Publisher: | ACM |
| Refereed: | Yes |
| Identification Number: | 10.1145/3776574.3832482 |
| Related URLs: | |
| Open Archives Initiative ID (OAI ID): | oai:eprints.whiterose.ac.uk:246391 |
Download
Filename: 3776574.3832482.pdf
Licence: CC-BY 4.0

CORE (COnnecting REpositories)
CORE (COnnecting REpositories)