LA-layer: General local attention layer for full attention networks

Publication date

2023-08-25

Authors

Lu, HuiISNI 0000000524098097
Poppe, RonaldISNI 0000000389426288
Salah, Albert AliORCID 0000-0001-6342-428XISNI 0000000091147032

Editors

Advisors

Supervisors

Document Type

Part of book
Open Access logo

License

taverne

Abstract

Attention layers have contributed to state-of-the-art results on vision tasks. Still, they leave room for improvement because position information is used in a fixed manner, and the computation cost is typically high. To mitigate both issues, we propose a convolution-style local attention layer (LA-layer) as a replacement for traditional attention layers. LA-layers not only encode the position information of pixels in a convolutional manner, but also produce position offsets following a novel constrained rule so that keys will deform and result in larger receptive fields. Query and keys are processed by a novel aggregation function that outputs attention weights for the values. In our experiments with different types of ResNets, we replace convolutional layers with LA-layers and address image recognition, object detection and instance segmentation tasks. We consistently demonstrate performance gains, despite having fewer FLOPs and training parameters. Our code is available at: https://github.com/hotfinda/LA-layer.

Keywords

Local attention, CNN, Deformable Kernel, Convolutional neural network, Taverne

Citation

Lu, H, Poppe, R & Salah, A 2023, LA-layer: General local attention layer for full attention networks. in Proceedings - 2023 IEEE International Conference on Multimedia and Expo, ICME 2023. Proceedings - IEEE International Conference on Multimedia and Expo, vol. 2023-July, pp. 2057-2062. https://doi.org/10.1109/ICME55011.2023.00352