You are viewing a plain text version of this content. The canonical link for it is here.
Posted to issues@flink.apache.org by "Roman Khachatryan (Jira)" <ji...@apache.org> on 2020/12/18 18:24:00 UTC

[jira] [Updated] (FLINK-20672) CheckpointAborted RPC failure can fail JM

     [ https://issues.apache.org/jira/browse/FLINK-20672?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ]

Roman Khachatryan updated FLINK-20672:
--------------------------------------
    Fix Version/s: 1.13.0

> CheckpointAborted RPC failure can fail JM
> -----------------------------------------
>
>                 Key: FLINK-20672
>                 URL: https://issues.apache.org/jira/browse/FLINK-20672
>             Project: Flink
>          Issue Type: Bug
>          Components: Runtime / Checkpointing
>    Affects Versions: 1.12.0, 1.11.3
>            Reporter: Roman Khachatryan
>            Priority: Major
>             Fix For: 1.13.0
>
>
> Introduced in FLINK-8871, aborted RPC notifications are done asynchonously:
>  
> {code}
> 	private void sendAbortedMessages(long checkpointId, long timeStamp) {
> 		// send notification of aborted checkpoints asynchronously.
> 		executor.execute(() -> {
> 			// send the "abort checkpoint" messages to necessary vertices.
>                         // ..
> 		});
> 	}
> {code}
> However, the executor that eventually executes this request is created as follows
> {code}
> 		final ScheduledExecutorService futureExecutor = Executors.newScheduledThreadPool(
> 				Hardware.getNumberCPUCores(),
> 				new ExecutorThreadFactory("jobmanager-future"));
> {code}
> ExecutorThreadFactory uses UncaughtExceptionHandler that exits JVM on error.
> cc: [~yunta]



--
This message was sent by Atlassian Jira
(v8.3.4#803005)