Repository navigation
SAMZA-2277: Semantics for cluster-manager.container.retry.window.ms not reflected in code - #1108
dnishimura wants to merge 5 commits into
Conversation
…ot reflected in code
| * If there are too many failed container failures (configured by job.container.retry.count) for a | ||
| * processor, the job exits. | ||
| */ | ||
| volatile boolean tooManyFailedContainers = false; |
There was a problem hiding this comment.
Would it be better to expose a getter instead, for unit-testing than make package private?
| processorFailures.put(processorId, new ProcessorFailure(1, System.currentTimeMillis())); | ||
| } | ||
|
|
||
| long lastFailureMsDiff = Instant.now().toEpochMilli() - lastFailureTime; |
There was a problem hiding this comment.
Minor: Might be preferable to set lastFailureMsDiff in the if-condition above, to avoid time-skew in measurements.
There was a problem hiding this comment.
I imagine it would be a sub-millisecond skew? Would that matter? Will change anyways to make the code cleaner.
There was a problem hiding this comment.
Typically, yes. In some cases, CPU preemption can cause the time to be relatively large, i've only seen it once when the host was so busy it was spending most of the time context-switching
rmatharu-zz
left a comment
There was a problem hiding this comment.
Thanks for this fix. Couple of minor things.
dnishimura
left a comment
There was a problem hiding this comment.
@rmatharu addressed your comments. Please merge to master after you approve the latest changes.
| processorFailures.put(processorId, new ProcessorFailure(1, System.currentTimeMillis())); | ||
| } | ||
|
|
||
| long lastFailureMsDiff = Instant.now().toEpochMilli() - lastFailureTime; |
There was a problem hiding this comment.
I imagine it would be a sub-millisecond skew? Would that matter? Will change anyways to make the code cleaner.
| * If there are too many failed container failures (configured by job.container.retry.count) for a | ||
| * processor, the job exits. | ||
| */ | ||
| volatile boolean tooManyFailedContainers = false; |
|
@rmatharu I addressed your minor comments. Please merge if you don't have any more comments. Thanks! |
| * If there are too many failed container failures (configured by job.container.retry.count) for a | ||
| * processor, the job exits. | ||
| */ | ||
| volatile boolean tooManyFailedContainers = false; |
There was a problem hiding this comment.
Should we update the comment to "If there are more than job.container.retry.count failures of a container within a job.container.retry.window period, the CPM exits"?
and update the flag to jobFailureCriteriaMet.
There was a problem hiding this comment.
Good suggestion, it makes it clearer. Will change.
| taskManager.stop(); | ||
| } | ||
|
|
||
| @Test |
There was a problem hiding this comment.
/** Test scenario where a container fails multiple times but failures are more than retryWindow apart.*/
rmatharu-zz
left a comment
There was a problem hiding this comment.
Thanks for the changes.
Left couple of minor comments.
Feel free to checkin
|
@rmatharu I made your suggested changes. I'm not a committer so I can't check in the changes. Please merge this PR when you get a chance. Thanks! |
In the current code, the window is only applied and checked on the last retry. However, the check should be done at all retries.
This was found during #1104
@rmatharu please take a look