Geo-Aware Image Caption Generation
Publication date
2020
Editors
Advisors
Supervisors
DOI
Document Type
Contribution to conference
Metadata
Show full item recordCollections
License
Abstract
Standard image caption generation systems produce generic descriptions of images and do not utilize any contextual information or world knowledge. In particular, they are unable to generate captions that contain references to the geographic context of an image, for example, the location where a photograph is taken or relevant geographic objects around an image location. In this paper, we develop a geo-aware image caption generation system, which incorporates geographic contextual information into a standard image captioning pipeline. We propose a way to build an image-specific representation of the geographic context and adapt the caption generation network to produce appropriate geographic names in the image descriptions. We evaluate our system on a novel captioning dataset that contains contextualized captions and geographic metadata and achieve substantial improvements in BLEU, ROUGE, METEOR and CIDEr scores. We also introduce a new metric to assess generated geographic references directly and empirically demonstrate our system's ability to produce captions with relevant and factually accurate geographic referencing.
Keywords
image captioning, caption generation, knowledge integration, geographic information, contextualized language generation, contextualization, geographic context
Citation
Nikiforova, S, Deoskar, T, Paperno, D & Vinter Seggev, Y S 2020, 'Geo-Aware Image Caption Generation', Paper presented at The 28th International Conference on Computational Linguistics (COLING), 8/12/20 - 13/12/20 pp. 3143-3156. < https://www.aclweb.org/anthology/2020.coling-main.280/ >, conference