Niu, P. orcid.org/0000-0002-0668-6950, Cai, T. orcid.org/0000-0002-3624-6120, Zhang, Y. orcid.org/0000-0001-8177-0610 et al. (4 more authors) (2025) LG-Umer: UNet-like network integrate local–global feature with novel attention for road extraction from remote sensing images. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 18. pp. 21755-21768. ISSN: 1939-1404
Abstract
Road extraction from remote sensing images is a key research area in smart city development. While deep learning techniques have demonstrated remarkable effectiveness in this domain, existing approaches exhibit limitations: convolutional neural network (CNN)-based methods struggle to capture global contextual information for long-range road networks, vision transformer (ViT)-based methods fail to adequately extract multiscale local features, and hybrid CNN-ViT architectures overlook the synergistic guidance between local and global features. To address these challenges, we propose LG-Umer, a UNet-like network that integrates Local-Global features with a novel attention mechanism, combining the complementary strengths of CNNs and ViTs within an encoder-decoder framework. Specifically, the encoder employs a multiscale strip deformational module, which utilizes deformable convolutions to adaptively extract topological structures and variable-shaped local road features. In the decoder, a multistage gate unit module is introduced, incorporating a novel attention mechanism to model long-range dependencies by leveraging local features as attention operators for global feature refinement. Extensive experiments on three public benchmarks demonstrate the superiority of LG-Umer. It achieves IoU scores of 70.4%, 71.2%, and 68.7% on the Massachusetts Road, DeepGlobe Road, and CHN6-CUG datasets, respectively, surpassing recent state-of-the-art methods by 1.2%, 0.9%, and 1.1%. These results validate the effectiveness of our approach in balancing local detail preservation and global contextual modeling for road extraction tasks.
Metadata
| Item Type: | Article |
|---|---|
| Authors/Creators: |
|
| Copyright, Publisher and Additional Information: | © 2025 The Authors. This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ |
| Keywords: | Building extraction; deep learning (DL); global attention; multiscale direction context-aware |
| Dates: |
|
| Institution: | The University of Sheffield |
| Academic Units: | The University of Sheffield > Faculty of Engineering (Sheffield) > Department of Computer Science (Sheffield) |
| Date Deposited: | 18 Aug 2026 10:10 |
| Last Modified: | 18 Aug 2026 10:10 |
| Status: | Published |
| Publisher: | Institute of Electrical and Electronics Engineers (IEEE) |
| Refereed: | Yes |
| Identification Number: | 10.1109/jstars.2025.3573735 |
| Related URLs: | |
| Sustainable Development Goals: | |
| Open Archives Initiative ID (OAI ID): | oai:eprints.whiterose.ac.uk:244466 |


CORE (COnnecting REpositories)
CORE (COnnecting REpositories)