Optimizing Distributed Training on Frontier for Large Language Models

"We enable and investigate various model and data parallel training techniques, such as tensor parallelism, pipeline parallelism, and sharded data parallelism, to facilitate training a trillion-parameter model on Frontier. 

"We empirically assess these techniques and their associated parameters to determine their impact on memory footprint, communication latency, and GPU's computational efficiency. 

"We analyze the complex interplay among these techniques and find a strategy to combine them to achieve high throughput through hyperparameter tuning. 

"We have identified efficient strategies for training large LLMs of varying sizes through empirical analysis and hyperparameter tuning."

Comments

Popular posts from this blog

Supporting Artistes (SAs)

Hamza Chaudhry

Injection