VAQUUM: Are vague quantifiers grounded in visual data

Publication date

2025-07

Authors

Wong, Hugh Mee
Nouwen, R.W.F.ORCID 0000-0001-9571-4644ISNI 0000000398065728
Gatt, AlbertORCID 0000-0001-6388-8244ISNI 0000000048277966

Editors

Che, Wanxiang
Nabende, Joyce
Shutova, Ekaterina
Pilehvar, Mohammad Taher

Advisors

Supervisors

Document Type

Part of book
Open Access logo

License

No license information available

Abstract

Vague quantifiers such as “a few” and “many” are influenced by various contextual factors, including the number of objects present in a given context. In this work, we evaluate the extent to which vision-and-language models (VLMs) are compatible with humans when producing or judging the appropriateness of vague quantifiers in visual contexts. We release a novel dataset, VAQUUM, containing 20,300 human ratings on quantified statements across a total of 1089 images. Using this dataset, we compare human judgments and VLM predictions using three different evaluation methods. Our findings show that VLMs, like humans, are influenced by object counts in vague quantifier use. However, we find significant inconsistencies across models in different evaluation settings, suggesting that judging and producing vague quantifiers rely on two different processes. We release our dataset and code at https://github.com/hughmee/vaquum.

Keywords

Citation

Wong, H M, Nouwen, R & Gatt, A 2025, VAQUUM: Are vague quantifiers grounded in visual data. in W Che, J Nabende, E Shutova & M T Pilehvar (eds), Findings of the Association for Computational Linguistics: ACL. Association for Computational Linguistics, Vienna, pp. 11966-11982. https://doi.org/10.18653/v1/2025.findings-acl.619