Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?

Publication date

2025-07

Authors

Song, YingjinISNI 0000000527553948
Du, YupeiISNI 0000000493058809
Paperno, DenisISNI 000000037085651X
Gatt, AlbertORCID 0000-0001-6388-8244ISNI 0000000048277966

Editors

Advisors

Supervisors

Document Type

Contribution to conference
Open Access logo

License

cc_by

Abstract

This paper introduces the TempVS benchmark, which focuses on temporal grounding and reasoning capabilities of Multimodal Large Language Models (MLLMs) in image sequences. TempVS consists of three main tests (i.e., event relation inference, sentence ordering and image ordering), each accompanied with a basic grounding test. TempVS requires MLLMs to rely on both visual and linguistic modalities to understand the temporal order of events. We evaluate 38 state-of-the-art MLLMs, demonstrating that models struggle to solve TempVS, with a substantial performance gap compared to human capabilities. We also provide fine-grained insights that suggest promising directions for future research. Our TempVS benchmark data and code are available at https://github.com/yjsong22/TempVS.

Keywords

Citation

Song, Y, Du, Y, Paperno, D & Gatt, A 2025, 'Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?', Paper presented at 63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, 27/07/25 - 1/08/25 pp. 24316-24342. https://doi.org/10.18653/v1/2025.findings-acl.1248, conference