Distributed Expand #18126

wschin · 2023-10-27T03:14:46Z

This PR implements DistributedExpand for llama 2.

Representative Examples of DistributedExpand:

[shard on non-expanded axis] input tensor (shape=[8, 1], spec=S[0]R, device_mesh=[0,1]) -> Expand(target_shape=[8, 2] -> output tensor (shape=[8, 2], spec=S[0]R, device_mesh=[0,1])
[sharding expanded axis is invalid since it must have dim=1 and axis with dim=1 cannot be sharded] input tensor (shape=[1, 8], spec=S[0]R, device_mesh=[0,1]) -> Expand(target_shape=[2, 8] -> output tensor (shape=[2, 8], spec=S[0]R, device_mesh=[0,1])

From those examples, we observe a few important behaviors.

The output sharding spec is always the same to the input sharding spec.
Expanding always happen on axis with dimension=1. Otherwise, it will violate the broadcasting rule.
No communication is needed since all computation can happen locally. Let's consider the first example again. If you put the first half tensor (shape: [4, 1]) on device 0 and the second half (shape: [4, 1]) on device 1, then Expand it with target shape [4, 2] , these two local tensors (shape: [4, 2]) are exactly the same as the one described by output sharding spec.

Algorithm:

Compute logical (i.e., unsharded) shapes of input and output.
Compute sharded output shape from logical output.
Call Expand to broadcast local input to sharded output shape.

How to review?

Start with changes in onnxruntime_test_distributed.py. Those tests are good examples for using this op.
Read expand.h/expand.cc. Theose changes are for exposing functionalities in Expand to DistributedExpand.
Read distributed_expand.h/distributed_expand.cc. It follows the algorithm described above. The commit 68ac301 first sketches the definition of DistributedExpand. The next commit 0eb9330 adds real implementation.

onnxruntime/contrib_ops/cuda/collective/distributed_expand.cc

Fix a function call

This PR implements DistributedExpand for llama 2. Representative Examples of DistributedExpand: - [shard on non-expanded axis] `input tensor (shape=[8, 1], spec=S[0]R, device_mesh=[0,1]) -> Expand(target_shape=[8, 2] -> output tensor (shape=[8, 2], spec=S[0]R, device_mesh=[0,1])` - [sharding expanded axis is invalid since it must have dim=1 and axis with dim=1 cannot be sharded] `input tensor (shape=[1, 8], spec=S[0]R, device_mesh=[0,1]) -> Expand(target_shape=[2, 8] -> output tensor (shape=[2, 8], spec=S[0]R, device_mesh=[0,1])` From those examples, we observe a few important behaviors. - The output sharding spec is always the same to the input sharding spec. - Expanding always happen on axis with dimension=1. Otherwise, it will violate the broadcasting rule. - No communication is needed since all computation can happen locally. Let's consider the first example again. If you put the first half tensor (shape: [4, 1]) on device 0 and the second half (shape: [4, 1]) on device 1, then `Expand` it with target shape [4, 2] , these two local tensors (shape: [4, 2]) are exactly the same as the one described by output sharding spec. Algorithm: - Compute logical (i.e., unsharded) shapes of input and output. - Compute sharded output shape from logical output. - Call Expand to broadcast local input to sharded output shape. How to review? - Start with [changes in onnxruntime_test_distributed.py](microsoft@ea33392). Those tests are good examples for using this op. - [Read expand.h/expand.cc](microsoft@e4c4998). Theose changes are for exposing functionalities in Expand to DistributedExpand. - Read distributed_expand.h/distributed_expand.cc. It follows the algorithm described above. The commit microsoft@68ac301 first sketches the definition of DistributedExpand. The next commit microsoft@0eb9330 adds real implementation.

github-advanced-security bot found potential problems Oct 27, 2023

View reviewed changes

onnxruntime/contrib_ops/cuda/collective/distributed_expand.cc Fixed Show fixed Hide fixed

wschin force-pushed the wechi/d-expand branch from 50ed22c to d92dca2 Compare October 27, 2023 18:08

wschin marked this pull request as ready for review October 27, 2023 19:26

wschin requested a review from souptc October 27, 2023 19:31

souptc approved these changes Oct 27, 2023

View reviewed changes

wschin added 5 commits October 27, 2023 14:52

Expose Expand

e4c4998

Add tests

ea33392

Skeleton of DistributedExpand

68ac301

Implement details for d-expand

0eb9330

Fix a function call

lint

128f87c

wschin force-pushed the wechi/d-expand branch from b06ffd5 to 128f87c Compare October 27, 2023 21:53

wschin closed this Oct 28, 2023

wschin reopened this Oct 28, 2023

wschin merged commit 24f9c1a into main Oct 28, 2023
97 of 101 checks passed

wschin deleted the wechi/d-expand branch October 28, 2023 07:44

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Distributed Expand #18126

Distributed Expand #18126

wschin commented Oct 27, 2023 •

edited

Loading

Distributed Expand #18126

Distributed Expand #18126

Conversation

wschin commented Oct 27, 2023 • edited Loading

wschin commented Oct 27, 2023 •

edited

Loading