Skip to content

Understand why jobs are still round to drained fleets #918

Description

@huydhn

As part of pytorch/pytorch#190347 deployment, we drained ue2 https://github.com/pytorch/ci-infra/actions/runs/29603818377/job/87962059939. Previously, draining a fleet correctly blocks jobs going there. However, the behaviour seems to have been changed in which jobs are still being routed to the drained cluster as shown in the dashboard

Per @jeanschmidt, LF fleet draining still works correctly.

Q:

  1. Could something in OSDC cause this? Or is this something that GitHub changes recently?
  2. Shoud we implement runner group routing decision ourselves on CI? Similar to mt <-> lf routing

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

Status
In Progress

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions