Xi Xiaoyao Technology Talk | Stop Saying GPT-4V is Amazing! It Can't Even Recognize Peking Duck, Can You Believe It??
7.5K views
Translation:
Regarding the recently highly-discussed visual language model GPT-4V, researchers have constructed a new benchmark test called HallusionBench to evaluate its image reasoning capabilities. The results indicate that models like GPT-4V perform poorly on HallusionBench, often succumbing to language hallucinations influenced by their parametric memories, with error rates as high as 90%. Additionally, GPT-4V's performance on visual tasks involving geometry is also unsatisfactory, highlighting its current limitations in visual abilities. Simple image manipulations can easily mislead GPT-4V, exposing its vulnerabilities. In contrast, LLaVA-1.5, while not as richly knowledgeable as GPT-4V, has fewer common sense errors. This study reveals the limitations of current visual language models in image reasoning and provides insights for future improvements.
Related
OpenAI Talent Mobility: Former Researcher Tian Yonglong Joins Tencent, Focused on Visual Language Model Development
16.3K
Unlocking PB-Level Video Assets! InfiniMind, Founded by a Former Google Employee, Helps Enterprises Mine Video Dark Data
15.0K

Moondream Raises $4.5 Million to Launch a 1.6 Billion Parameter Efficient AI Model with 5K GitHub Stars
15.7K

Small but Powerful! H2O.ai Launches New AI Visual Models to Outperform Tech Giants in Document Analysis
16.7K
Alibaba Cloud's Tongyi Qwen Responds to Github Page 404: In Contact with Officials
17.0K
Tongyi Qwen Open Source Visual Language Model Qwen2-VL API Available in 2B and 7B Sizes
27.9K
