VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues
Journal:
arXiv
Published Date:
Feb 17, 2025
Abstract
Visually linking matching cues is a crucial ability in daily life, such as
identifying the same person in multiple photos based on their cues, even
without knowing who they are. Despite the extensive knowledge that
vision-language models (VLMs) possess, it remains largely unexplored whether
they are capable of performing this fundamental task. To address this, we
introduce VLM2-Bench, a benchmark designed to assess whether VLMs can Visually
Link Matching cues, with 9 subtasks and over 3,000 test cases. Comprehensive
evaluation across eight open-source VLMs and GPT-4o, along with further
analysis of various language-side and vision-side prompting methods, leads to a
total of eight key findings. We identify critical challenges in models' ability
to link visual cues, highlighting a significant performance gap where even
GPT-4o lags 34.80% behind humans. Based on these insights, we advocate for (i)
enhancing core visual capabilities to improve adaptability and reduce reliance
on prior knowledge, (ii) establishing clearer principles for integrating
language-based reasoning in vision-centric tasks to prevent unnecessary biases,
and (iii) shifting vision-text training paradigms toward fostering models'
ability to independently structure and infer relationships among visual cues.