The article examines the phenomenon of "benchmarking chaos" in the current evaluation systems for large models, noting that there is a widespread occurrence of "everyone being number one" in the rankings. The existing open-source benchmarking datasets can lead to a "problem-solving" mentality, while closed proprietary datasets can affect fairness. Additionally, some rankings lack scientific and comprehensive evaluation dimensions. The article suggests establishing an authoritative evaluation system, open-sourcing the evaluation tools and processes to ensure fairness, but adopting a model of open historical + closed formal datasets for evaluation. Moreover, the commercialization of large models is far more important than the parameters of the models or their rankings on the leaderboards.
"Baimao Battle" Family's First, When Will Cheating in Large Model 'Scoring' Stop?
10.9K views
Related
ByteDance DouBao Beta Test Ride-Hailing Service AI Agent Accelerates Service Reconstruction Entrance
14.2K
Behind the Hype of DeepSeek-V4: How Does the Open-Source Framework One-Eval End the AI Evaluation Nightmare?
13.3K

AliTongyi Qianwen App Exclusively Brands Four Provincial Satellite Spring Festival Galas, AI Intelligent Entity Makes Debut on the Art Stage
15.0K
Alibaba Qwen2-72B Tops HELM Ranking: Performance Surpasses Llama3-70B
13.5K
Ant Group Releases Benchmark for Large Model Evaluation in the DevOps Field
10.9K
Investigation into the Chaos of Large Model Evaluation: Parameter Scale Does Not Represent Everything
10.9K
