Vision LLMs Are Bad at Hierarchical Visual Understanding, and LLMs Are the Bottleneck
Journal:
arXiv
Published Date:
May 30, 2025
Abstract
This paper reveals that many state-of-the-art large language models (LLMs)
lack hierarchical knowledge about our visual world, unaware of even
well-established biology taxonomies. This shortcoming makes LLMs a bottleneck
for vision LLMs' hierarchical visual understanding (e.g., recognizing Anemone
Fish but not Vertebrate). We arrive at these findings using about one million
four-choice visual question answering (VQA) tasks constructed from six
taxonomies and four image datasets. Interestingly, finetuning a vision LLM
using our VQA tasks reaffirms LLMs' bottleneck effect to some extent because
the VQA tasks improve the LLM's hierarchical consistency more than the vision
LLM's. We conjecture that one cannot make vision LLMs understand visual
concepts fully hierarchical until LLMs possess corresponding taxonomy
knowledge.