Alluqmani, A., Harvey, M. orcid.org/0000-0001-5504-2089 and Paramita, M. (2026) Generating clothing graphic captions for visually impaired users. Multimedia Tools and Applications, 15. 621. ISSN: 1380-7501
Abstract
Purchasing clothing online requires the identification of many different visual features such as colour, style, material and size. While this is simple task for most people, this is no so for Visually Impaired (VI) people, who are limited to relying on textual information, which is often lacking. Although current research has proposed image captioning models to enrich clothing descriptions, the models used thus far have been unable to generate descriptions of essential fine-grained elements such as graphics. We conducted a manual analysis of Ground Truth (GT) clothing image-text datasets and determined that several limitations disqualify them from consideration as gold standard captions. We collected clothing graphic images creating a new Clothing Graphic Dataset (CGD) and developed a novel zero-shot VI-friendly clothing graphic captioning model which adopts a Region Of Interest-based approach to generate a focused, detailed description of graphic elements, leveraging the power of Large Language Models to consider VI people's needs. The results of our evaluation revealed that, compared with BLIP-2 and GT captions, our model generates the most informative captions for VI people. The model's ability to produce detailed, relevant descriptions demonstrates its potential to improve online clothing accessibility for VI users without the need for costly and time-consuming pre-training data.
Metadata
| Item Type: | Article |
|---|---|
| Authors/Creators: |
|
| Copyright, Publisher and Additional Information: | © 2026 The Authors. Except as otherwise noted, this author-accepted version of a journal article published in Multimedia Tools and Applications is made available via the University of Sheffield Research Publications and Copyright Policy under the terms of the Creative Commons Attribution 4.0 International License (CC-BY 4.0), which permits unrestricted use, distribution and reproduction in any medium, provided the original work is properly cited. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/ |
| Keywords: | GenAI; Image captioning; Clothing graphic descriptions; Clothing datasets; Visually impaired |
| Dates: |
|
| Institution: | The University of Sheffield |
| Academic Units: | The University of Sheffield > Faculty of Social Sciences (Sheffield) > School of Information, Journalism and Communication |
| Date Deposited: | 16 Jul 2026 10:54 |
| Last Modified: | 16 Jul 2026 10:54 |
| Status: | Published |
| Publisher: | Springer |
| Refereed: | Yes |
| Identification Number: | 10.1007/s11042-026-21774-w |
| Related URLs: | |
| Open Archives Initiative ID (OAI ID): | oai:eprints.whiterose.ac.uk:243106 |

CORE (COnnecting REpositories)
CORE (COnnecting REpositories)