Spring Batch 在集群环境中正确重启未完成的作业

ale*_*oid 5 high-availability spring-batch spring-boot

我使用以下逻辑在单节点 Spring Batch 应用程序上重新启动未完成的作业:

public void restartUncompletedJobs() {

    try {
        jobRegistry.register(new ReferenceJobFactory(documetPipelineJob));

        List<String> jobs = jobExplorer.getJobNames();
        for (String job : jobs) {
            Set<JobExecution> runningJobs = jobExplorer.findRunningJobExecutions(job);

            for (JobExecution runningJob : runningJobs) {
                runningJob.setStatus(BatchStatus.FAILED);
                runningJob.setEndTime(new Date());
                jobRepository.update(runningJob);
                jobOperator.restart(runningJob.getId());
            }
        }
    } catch (Exception e) {
        LOGGER.error(e.getMessage(), e);
    }
}
Run Code Online (Sandbox Code Playgroud)

现在我正试图让它在双节点集群上工作。每个节点上的两个应用程序都将指向共享的 PostgreSQL 数据库。

让我们考虑以下示例:我有 2 个作业实例 -jobInstance1正在运行node1,jobInstance2正在运行node2。Node1在jobInstance1执行过程中由于某种原因重新启动。后node1重新启动春季批处理应用程序尝试重新启动与上面给出逻辑未完成任务-它看到有2个未完成的作业实例-jobInstance1和jobInstance2(这是正常运行的node2),并尝试重新启动它们。这种方式改为重新启动 only jobInstance1- 它将重新启动jobInstance1和jobInstance2.. 但jobInstance2不应重新启动,因为它现在正在正确执行node2。

如何在应用程序启动期间正确重启未完成的作业(在前一个应用程序终止之前)并防止类似作业也jobInstance2将重新启动的情况?

更新

这是以下答案中提供的解决方案:

Get the job instances of your job with JobOperator#getJobInstances

For each instance, check if there is a running execution using JobOperator#getExecutions.

2.1 If there is a running execution, move to next instance (in order to let the execution finish either successfully or with a failure)

2.2 If there is no currently running execution, check the status of the last execution and restart it if failed using JobOperator#restart.
Run Code Online (Sandbox Code Playgroud)

我有一个关于 #2.1 的问题 - Spring Batch 会在应用程序重新启动后自动重新启动未完成的作业,还是我需要手动执行此操作?