MMVP

Introduced by Tong et al. in Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

The MMVP (Multimodal Visual Patterns) Benchmark focuses on identifying "CLIP-blind pairs" – images that appear similar to the CLIP model despite having clear visual differences. These patterns highlight the challenges these systems face in answering straightforward questions, often leading to incorrect responses and hallucinated explanations.

Homepage