Skip to content

[SPARK-59819][CORE] Fix Resubmitted task end accounting in AppStatusListener - #59092

Open
dongjoon-hyun wants to merge 1 commit into
apache:masterfrom
dongjoon-hyun:SPARK-59819
Open

dongjoon-hyun wants to merge 1 commit into
apache:masterfrom
dongjoon-hyun:SPARK-59819

Conversation

@dongjoon-hyun

@dongjoon-hyun dongjoon-hyun commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

What changes were proposed in this pull request?

This PR aims to fix two Resubmitted accounting bugs in AppStatusListener.onTaskEnd.

When an executor is lost, TaskSetManager.executorLost posts SparkListenerTaskEnd with Resubmitted reason for the already-successful shuffle map tasks of that executor (without external shuffle service or on decommission) if the task set is not zombie. This event reuses the finished TaskInfo of the successful attempt and has no matching SparkListenerTaskStart. Since SPARK-41187 (#38702), onTaskEnd uses activeDelta = 0 for Resubmitted, but the following places still handle it like a normal task end.

  1. The speculation summary decrements numActiveTasks unconditionally. This PR uses activeDelta like the other active task counters.
  2. ExecutorStageSummary.taskTime and ExecutorSummary.totalDuration add event.taskInfo.duration again although the duration of the successful attempt was already counted. This PR skips them for Resubmitted like the executor-level task metrics, which are already skipped for Resubmitted.

Why are the changes needed?

To show correct values.

  • When a successful speculative task is resubmitted, numActiveTasks of the speculation summary becomes negative (e.g. -1), and nothing corrects it later because only onTaskStart and onTaskEnd update it. This value is stored in SpeculationStageSummaryWrapper and shown via the REST API speculationSummary field and the Speculation Summary table of the stage page.
  • For every Resubmitted task end, the duration of the successful attempt is counted twice in ExecutorStageSummary.taskTime (Task Time of Aggregated Metrics by Executor table in the stage page). ExecutorSummary.totalDuration (Task Time (GC Time) in the executors page and totalDuration_seconds_total in the Prometheus metrics) is also counted twice if the Resubmitted task end arrives before SparkListenerExecutorRemoved of the lost executor. This is timing-dependent because CoarseGrainedSchedulerBackend posts SparkListenerExecutorRemoved right after TaskSchedulerImpl.executorLost, while the Resubmitted task end is posted asynchronously by the DAGScheduler event loop. Note that the old ExecutorsListener of Apache Spark 2.2 ignored Resubmitted task ends including the duration, but SPARK-20646 ([SPARK-20646][core] Port executors page to new UI backend. #19678) kept only the task metrics part when porting it to AppStatusListener.

These were found during the review of #58968.

Does this PR introduce any user-facing change?

Yes, this fixes the incorrect values in the Spark UI, REST API, and Prometheus metrics. The negative numActiveTasks of the speculation summary exists since Apache Spark v3.3.0 (2022-06-09), and the double-counted task duration exists since Apache Spark v2.3.0 (2018-02-22).

How was this patch tested?

Pass the CIs with the newly added test case. I also manually verified that the new test case fails without this fix.

Was this patch authored or co-authored using generative AI tooling?

Generated-by: Claude Opus 5.5

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants