TY - GEN
T1 - T3P
T2 - 2nd Workshop on Networks for AI Computing, NAIC 2025, Part of SIGCOMM 2025
AU - Yochana, Saar Ben
AU - Avin, Chen
AU - Scalosub, Gabriel
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s).
PY - 2025/9/8
Y1 - 2025/9/8
N2 - As deep learning models continue to grow in scale and complexity, methods of distributed machine learning training, and particularly those used for large language models (LLMs), have become a critical ingredient in making such computations efficient and feasible. In such contexts, tensor parallelism (TP) is widely employed to distribute computations across multiple accelerators. However, since TP mandates frequent and high-volume communication between devices, the underlying network characteristics significantly influence performance. Previous work was mostly either model-agnostic or topology-agnostic and did not pick provably optimal configurations. This study presents Topology-Tailored Tensor Parallelism, T3P, an efficient algorithm that identifies the communication-optimal TP sharding configuration (within the considered search space) based on both the model architecture and the network topology. In particular, we show that T3P is optimal for any given resharding cost model.
AB - As deep learning models continue to grow in scale and complexity, methods of distributed machine learning training, and particularly those used for large language models (LLMs), have become a critical ingredient in making such computations efficient and feasible. In such contexts, tensor parallelism (TP) is widely employed to distribute computations across multiple accelerators. However, since TP mandates frequent and high-volume communication between devices, the underlying network characteristics significantly influence performance. Previous work was mostly either model-agnostic or topology-agnostic and did not pick provably optimal configurations. This study presents Topology-Tailored Tensor Parallelism, T3P, an efficient algorithm that identifies the communication-optimal TP sharding configuration (within the considered search space) based on both the model architecture and the network topology. In particular, we show that T3P is optimal for any given resharding cost model.
KW - Data Center Networks
KW - Distributed Machine Learning
KW - Tensor Parallelism
UR - https://www.scopus.com/pages/publications/105018054985
U2 - 10.1145/3748273.3749209
DO - 10.1145/3748273.3749209
M3 - Conference contribution
AN - SCOPUS:105018054985
T3 - NAIC 2025 - Proceedings of the 2nd Workshop on Networks for AI Computing, Part of SIGCOMM 2025
SP - 81
EP - 88
BT - NAIC 2025 - Proceedings of the 2nd Workshop on Networks for AI Computing, Part of SIGCOMM 2025
PB - Association for Computing Machinery, Inc
Y2 - 8 September 2025 through 11 September 2025
ER -