Computational Bottlenecks of Training Small-Scale Large Language Models

AuthorsSaleh Ashkboos, Iman Mirzadeh, Keivan Alizadeh, Mohammad Hossein Sekhavat, Moin Nabi, Mehrdad Farajtabar, Fartash Faghri

View publication

This paper was accepted at the Efficient Natural Language and Speech Processing (ENLSP) Workshop at NeurIPS 2024.

While large language models (LLMs) dominate the AI landscape, Small-scale large Language Models (SLMs) are gaining attention due to cost and efficiency demands from consumers. However, there is limited research on the training behavior and computational requirements of SLMs. In this study, we explore the computational bottlenecks of training SLMs (up to 2B parameters) by examining the effects of various hyperparameters and configurations, including GPU type, batch size, model size, communication protocol, attention type, and the number of GPUs. We assess these factors on popular cloud services using metrics such as loss per dollar and tokens per second. Our findings aim to support the broader adoption and optimization of language model training for low-resource AI research institutes.

Related readings and updates.

Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization

November 12, 2024research area Methods and Algorithms, research area Speech and Natural Language ProcessingWorkshop at NeurIPS

This paper was accepted at the Efficient Natural Language and Speech Processing (ENLSP) Workshop at NeurIPS 2024.

The pre-training phase of language models often begins with randomly initialized parameters. With the current trends in scaling models, training their large number of parameters can be extremely slow and costly. In contrast, small language models are less expensive to train, but they often cannot achieve the accuracy of large…

CAMPHOR: Collaborative Agents for Multi-Input Planning and High-Order Reasoning On Device

October 15, 2024research area Methods and Algorithms, research area Speech and Natural Language Processing

While server-side Large Language Models (LLMs) demonstrate proficiency in tool integration and complex reasoning, deploying Small Language Models (SLMs) directly on devices brings opportunities to improve latency and privacy but also introduces unique challenges for accuracy and memory. We introduce CAMPHOR, an innovative on-device SLM multi-agent framework designed to handle multiple user inputs and reason over personal context locally, ensuring…

Computational Bottlenecks of Training Small-Scale Large Language Models

Related readings and updates.

Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization

CAMPHOR: Collaborative Agents for Multi-Input Planning and High-Order Reasoning On Device

Discover opportunities in Machine Learning.