Zhejiang University Alumni Collaborate with Microsoft to Launch Multimodal Model LLaVA, Challenging GPT-4V
11.1K views
Translated data:
A Zhejiang University Chu Kochen Honors College alumnus has collaborated with Microsoft Research to launch the multimodal model LLaVA, challenging GPT-4V. LLaVA has performed exceptionally well on 11 test datasets, garnering over 6,000 stars. The model's comprehensive capabilities exceed 85% of GPT-4V's level. The open-source code, model, and training data for LLaVA are now available for use.
Related

SenseNova U1.5-Lite-Preview by SenseTime: 8B Model Supports Native 4K Image Generation
19.2K

16GB Memory, Local Instant Response! Google Releases Gemma 4 12B Revolutionary Encoder-Free Architecture Ignites Open Source Community
19.7K

Google Launches New Gemma 4 12B Model: Easily Handle Visual and Audio Data Without an Encoder
16.4K
New Developments in Domestic Large Models: MiniMax Launches the '10x Team' Program to Reward Global Top Experts
13.3K

NVIDIA Releases the Nemotron 3 Series Open-Source Models: Inference Efficiency Surges 5 Times
14.6K

Apple Releases New Multimodal Model Manzano: Breaking the Boundaries of Image Viewing and Drawing
13.9K
