[dataset] Pipeline task submission during reduce stage in push-based shuffle #25795

stephanie-wang · 2022-06-15T04:21:37Z

Why are these changes needed?

Reduce stage in push-based shuffle fails to complete at 100k output partitions or more. This is likely because of driver or raylet load from having too many tasks in flight at once.

We can fix this from ray core too, but for now, this PR adds pipelining for the reduce stage, to limit the total number of reduce tasks in flight at the same time. This is currently set to 2 * available parallelism in the cluster. We have to pick which reduce tasks to submit carefully since these are pinned to specific nodes. The PR does this by assigning tasks round-robin according to the corresponding merge task (which get spread throughout the cluster).

In addition, this PR refactors the map, merge, and reduce stages to use a common pipelined iterator pattern, since they all have a similar pattern of submitting a round of tasks at a time, then waiting for a previous round to finish before submitting more.

Related issue number

Closes #25412.

Checks

I've run scripts/format.sh to lint the changes in this PR.
I've included any doc changes needed for https://docs.ray.io/en/master/.
I've made sure the tests are passing. Note that there might be a few flaky tests, see the recent failures at https://flakey-tests.ray.io/
Testing Strategy
- Unit tests
- Release tests
- This PR is not tested :(

Adds a multinode test, but it would be nice to add a test that hooks into ray.remote and ray.get to check that we are actually submitting the right tasks at the right time.

… by tasks

stephanie-wang · 2022-06-15T04:21:56Z

(also includes changes from #25734)

ericl

The key config here is num_merge_tasks_per_round, is that right? Given this is a temporary workaround, I think we should make this feature flagged, and add a TODO to remove this once we fix the core issue. 100k tasks should fit comfortably within our scalability envelope, so this is a bit odd.

ericl · 2022-06-15T19:08:07Z

It would also be great to have a unit test that things work with pipelining enabled/disabled, so we can easily remove it in the future.

stephanie-wang · 2022-06-15T19:12:03Z

The key config here is num_merge_tasks_per_round, is that right? Given this is a temporary workaround, I think we should make this feature flagged, and add a TODO to remove this once we fix the core issue. 100k tasks should fit comfortably within our scalability envelope, so this is a bit odd.

Actually it's the reduce stage that is failing, so I think the numbers that matter are num rounds (= num map tasks / parallelism) * num reducers. Yeah, I also thought it was strange but I think it may have something to do with the fact that each reduce task also has a lot of plasma args, which we don't test in the scalability envelope. Fewer than in simple shuffle, but it's still on the order of 100s of args.

ericl

Feature flag lgtm... rubber stamping on the rest.

stephanie-wang added 9 commits June 13, 2022 13:59

Split push based schedule into all metadata and cheap metadata needed…

d64e9e5

… by tasks

tmp

129683b

tmp

b1d5c66

fix

eacabb5

fixes

e671f70

Merge branch 'fix-push-based-meta' into pipelined-reduce

779c26b

fix test

ac00bc3

pipelined map-merge stages

6d77122

pipelined reduce

3d3d152

stephanie-wang requested review from ericl, scv119, clarkzinzow, jjyao and jianoaix as code owners June 15, 2022 04:21

stephanie-wang assigned ericl, clarkzinzow and mwtian Jun 15, 2022

ericl reviewed Jun 15, 2022

View reviewed changes

stephanie-wang added 2 commits June 15, 2022 12:55

Merge remote-tracking branch 'upstream/master' into pipelined-reduce

ac42957

Fixes and unit tests

cf6e1ce

ericl approved these changes Jun 16, 2022

View reviewed changes

ericl added the @author-action-required The PR author is responsible for the next step. Remove tag to send back to the reviewer. label Jun 16, 2022

stephanie-wang added 2 commits June 17, 2022 10:21

Merge remote-tracking branch 'upstream/master' into pipelined-reduce

f47fa0b

fix

00d9b56

stephanie-wang merged commit 93aae48 into ray-project:master Jun 18, 2022

stephanie-wang deleted the pipelined-reduce branch June 18, 2022 00:33

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[dataset] Pipeline task submission during reduce stage in push-based shuffle #25795

[dataset] Pipeline task submission during reduce stage in push-based shuffle #25795

stephanie-wang commented Jun 15, 2022

stephanie-wang commented Jun 15, 2022

ericl left a comment

ericl commented Jun 15, 2022

stephanie-wang commented Jun 15, 2022

ericl left a comment

[dataset] Pipeline task submission during reduce stage in push-based shuffle #25795

[dataset] Pipeline task submission during reduce stage in push-based shuffle #25795

Conversation

stephanie-wang commented Jun 15, 2022

Why are these changes needed?

Related issue number

Checks

stephanie-wang commented Jun 15, 2022

ericl left a comment

Choose a reason for hiding this comment

ericl commented Jun 15, 2022

stephanie-wang commented Jun 15, 2022

ericl left a comment

Choose a reason for hiding this comment