On August 25, Alibaba Cloud introduced Qwen-VL, a large-scale visual language model that supports multiple languages including Chinese and English, and possesses the ability to jointly understand text and images. Based on Alibaba Cloud's previously open-sourced general-purpose language model Qwen-7B, Qwen-VL enhances its capabilities compared to other visual language models by adding features such as visual positioning and understanding of text within images. Qwen-VL has garnered over 3,400 stars on GitHub and has been downloaded more than 400,000 times. Visual language models are considered a significant evolution direction for general AI. The industry believes that models supporting multimodal inputs can enhance the understanding of the world and expand the range of applications. Through the open-sourcing of Qwen-VL, Alibaba Cloud is further advancing the development of general AI technology.
Aliyun Tongyi Qianwen Open Sources Again: Multimodal Large Model Qwen-VL
9.9K views
Related
Aliyun QoderWork Launches Peak-Valley Token: Use Qwen3.7-Max During Off-Peak Hours at Up to 20% Discount
13.2K
Aliyun Optimizes the BaiLian Multimodal Development Kit API Call Rate Limiting
15.8K

Alibaba Launches AI Development Tool Meoo to Empower Zero-Barrier Creative Monetization!
63.9K
QM Releases 2025 AI Application Ranking: Doubao, DeepSeek, Yuanbao, Afu, and Qianwen Rank in Top 5
18.4K

Ali Tengtouge Self-Developed AI Chip Zhenwu 810E Released
17.7K
Quest Mobile Releases Weekly Active Ranking of AI Applications: ByteDance-related Apps Top 3, Ant Tops 2
19.0K
