Repository navigation
[SPARK-60025] Support spec.suspend for a running SparkApplication - #948
dongjoon-hyun wants to merge 2 commits into
Conversation
|
Could you review this PR, @peter-toth ? FYI, for now, this is a pure Kueue integration. There is no relationship with Spark's |
There was a problem hiding this comment.
Thanks for the PR, @dongjoon-hyun!
This adds AppSuspendStep, which moves a running SparkApplication to Suspended before it releases the driver, the executors and the Kueue Workload. Clearing spec.suspend then starts a new attempt from Submitted through ApplicationStatus#resume, without the restart backoff. I found nothing wrong with the step order, the API server check for a driver which ended meanwhile, or the per-pod wait. The docs also say that a resumed attempt runs again from scratch, which matches your note that this is unrelated to Spark's hold-and-resume.
Non-blocking
- 1. Pod wait polls the API server every 2s: While pods remain, the step lists them on the API server every
terminationRequeuePeriodMillis. The pod informer already reconciles on each deletion. Requeueing at the default interval or at the next pod deadline, likeClusterSuspendStep, would need far fewer requests. inline
| log.debug("Waiting for the driver and executor pods of the suspended app to be deleted."); | ||
| return Optional.of( | ||
| completeAndRequeueAfter( | ||
| Duration.ofMillis(timeoutConfig.getTerminationRequeuePeriodMillis()))); |
There was a problem hiding this comment.
Finding 1. While pods remain, this requeues every terminationRequeuePeriodMillis, which is 2s by default. Each pass lists the driver and executor pods on the API server.
The pod informer already reconciles on each pod deletion, as the javadoc of releaseResources says. So the requeue only has to catch the end of the wait for a pod which sends no more events, e.g. one on a lost node. That wait is the grace period of the pod plus forceTerminationGracePeriodMillis. With the defaults (30s and 5 minutes), it takes about 165 lists per application. ClusterSuspendStep needs 4 for the same wait, since it requeues at the default interval or at the next pod deadline, whichever comes first.
The same shape here would be:
Instant nextDeadline = Instant.MAX;
for (Pod pod : pods) {
Instant deadline = getPodReleaseDeadline(pod, timeoutConfig);
boolean holding = now.isBefore(deadline);
// ... the deletions as they are ...
if (holding && deadline.isBefore(nextDeadline)) {
nextDeadline = deadline;
}
waiting |= holding;
}
if (waiting) {
// The deletion of each pod is observed by the pod informer, while the end of the wait for
// a pod which may send no more events is observed by the requeue
ReconcileProgress defaultRequeue = completeAndDefaultRequeue();
Duration remaining = Duration.between(now, nextDeadline);
return Optional.of(
remaining.compareTo(defaultRequeue.getRequeueAfterDuration()) < 0
? completeAndRequeueAfter(remaining)
: defaultRequeue);
}WAITING_FOR_PODS_PROGRESS in AppSuspendStepTest would then be completeAndDefaultRequeue().
There was a problem hiding this comment.
Thank you, @peter-toth. Right, the pod informer already reconciles on each pod deletion, including the executors, so the requeue only has to catch the deadline of a pod which sends no more events. I applied your suggestion, so that the wait is requeued at the next pod deadline or at the default interval, whichever comes first, like ClusterSuspendStep. I also added two test cases: the requeue at the next pod deadline, and no earlier requeue for pods which are already past their deadline.
|
Thank you always, @peter-toth . |
|
Merged to main |
What changes were proposed in this pull request?
This PR aims to support
spec.suspendfor a runningSparkApplication.AppSuspendStep, which moves a running application toSuspendedand releases its driver, executors, and KueueWorkload. Whenspec.suspendis set back tofalse, it starts a new attempt fromSubmittedwithApplicationStatus#resume.SuspendedtoAppSuspendStepinSparkAppReconciler, and skip the restart backoff for a resumed attempt inAppInitStep.Why are the changes needed?
Currently,
spec.suspendtakes effect only before the driver is requested. So a running attempt keeps its resources, including its Kueue quota, until it ends. A runningSparkClustercan already be suspended.Does this PR introduce any user-facing change?
Yes, within the unreleased
spec.suspend. Settingspec.suspendtotruenow stops the running attempt of aSparkApplication, and setting it back tofalsestarts a new attempt. There is no change compared to 1.0.0 (2026-07-23), which has nospec.suspend.How was this patch tested?
Pass the CIs with the newly added test cases.
Was this patch authored or co-authored using generative AI tooling?
Generated-by: Claude Opus 5.5