Scientists Innovate Technology to Successfully Train Trillion-Parameter Models at ChatGPT Level
7.1K views
Data to be translated:
Scientists have successfully trained a ChatGPT-level model using only 8% of the computational power of the world's most powerful supercomputer. This breakthrough comes from Oak Ridge National Laboratory, where the research team employed innovative techniques to train a trillion-parameter language model on the Frontier supercomputer. By utilizing distributed training and parallel technologies, they achieved 100% weak scaling efficiency. However, training large-scale language models still presents challenges, particularly in addressing memory issues. This research provides valuable experience for future training of enormous language models, highlighting the critical role of distributed training and parallel computing.
Related
ChatGPT Launches Custom Sticker Feature, Perfectly Integrates with Two Major Social Media Giants
15.2K
Say Goodbye to Fluctuations! OpenAI Takes Immediate Action: Codex Will Reset Quotas Tomorrow and Fix Usage Bug
14.9K

Pew Research: After the Release of ChatGPT, Over 30% of Web Pages Show Traces of AI-Generated Content
18.7K

ChatGPT Global Outage - User Login and Chat History Loading Function Abnormal
18.8K

ChatGPT Suffers Mass Outage, Millions of Users Face Login Crisis
16.8K

ChatGPT macOS New Computer History Feature Records User Work Trajectory
11.6K
