Kotlin coroutines: job trees & cancellation deep dive
Meta description: Learn how Kotlin coroutines Job hierarchy works at runtime, how cancellation propagates up and down the tree, and why SupervisorJob is the boundary that prevents cascade failures in production apps.
Tags: kotlin android kmp architecture mobile
TL;DR
Every coroutine lives inside a Job tree. Cancellation is bidirectional by default — a child failure cancels the parent, which cancels all siblings. SupervisorJob breaks the upward propagation contract, isolating failures at a boundary. Misuse produces three production failure modes: swallowed exceptions, leaked coroutines, and frozen scopes. Know the tree, own the outcome.
The Job tree is not an abstraction — it’s a runtime graph
Most teams get this wrong about structured concurrency: they treat it as a naming convention, not a runtime contract. Every launch or async call creates a new Job instance that registers itself as a child of the scope’s current Job. This forms a live, mutable tree in memory.
val scope = CoroutineScope(SupervisorJob() + Dispatchers.Default)
val parent = scope.launch { // Job A
val child1 = launch { ... } // Job B — child of A
val child2 = launch { ... } // Job C — child of A
}
At runtime, cancelling Job A propagates CancellationException down to B and C. That is the downward contract. The upward contract is the dangerous one.
Cancellation propagation: two directions, one default
By default, a non-cancellation exception thrown by a child propagates up to the parent. The parent treats it as a failure, cancels itself, and then cancels all remaining children. This is the cascade failure pattern.
val scope = CoroutineScope(Job() + Dispatchers.Default)
scope.launch {
launch { throw RuntimeException("child failed") } // Kills the entire scope
launch { delay(10_000) } // Also cancelled
}
In a medium-complexity Android ViewModel with 4–6 active coroutines, a single unhandled exception in any child leaves the scope permanently cancelled — silent, no crash, no log unless you instrument it.
The SupervisorJob boundary
SupervisorJob overrides the upward propagation contract. A child failure does not propagate to the parent or its siblings.
| Scope type | Child failure propagates up? | Siblings cancelled? | Parent cancelled? |
|---|---|---|---|
Job() | Yes | Yes | Yes |
SupervisorJob() | No | No | No |
supervisorScope { } | No (local block) | No | No |
viewModelScope in AndroidX is backed by SupervisorJob(). UI scopes manage independent coroutines for different UI components — a failed data-fetch coroutine should not kill an ongoing animation coroutine.
// Safe — supervisor isolates failures
class MyViewModel : ViewModel() {
fun loadData() = viewModelScope.launch {
// Failure here does not cancel other viewModelScope children
repository.fetch()
}
}
Three production failure modes
1. Swallowed exceptions
SupervisorJob does not propagate exceptions, but it also does not surface them unless you install a CoroutineExceptionHandler. Without one, the exception is sent to the thread’s uncaught exception handler — invisible in most Android crash reporters.
val scope = CoroutineScope(
SupervisorJob() + Dispatchers.Default + CoroutineExceptionHandler { _, e ->
logger.error("Coroutine failure", e)
}
)
2. Leaked coroutines
A frozen or cancelled parent scope silently rejects new launch calls — they complete immediately without executing. In practice, this surfaces as features that “just stopped working” after the first error, because the scope was killed by a Job() (not SupervisorJob()) and no one noticed.
3. async + SupervisorJob = deferred time bomb
async under a SupervisorJob does not throw on failure — it stores the exception in the Deferred. The exception only materialises when you call .await(). If you never await(), the exception is silently discarded.
val deferred = supervisorScope {
async { throw IOException("disk full") }
}
// Exception stored — NOT thrown yet
deferred.await() // Throws HERE, possibly far from the original call site
In my experience building production KMP apps, this is the single most common source of ghost failures in shared data-layer modules.
Choosing the right boundary
Application root scope → SupervisorJob (isolate feature-level failures)
ViewModel scope → SupervisorJob (already provided by viewModelScope)
Request/transaction scope → Job (one failure should abort the whole operation)
supervisorScope { } → Inline supervisor for async fan-out with mixed failure tolerance
Three things worth doing
-
Always pair
SupervisorJobwith aCoroutineExceptionHandler. The supervisor prevents cascade failures but silently swallows them without a handler. Logging is not optional — it’s part of the contract. -
Audit every
asynccall under supervisor scopes. EveryDeferredthat is neverawait-ed is a silently discarded failure. Enforceawait()calls at the call site or switch tolaunchwith explicit error handling. -
Use
Job()at transaction boundaries,SupervisorJob()at feature boundaries. A database write sequence should fail atomically — useJob(). A screen loading three independent data sources should isolate failures — useSupervisorJob()orsupervisorScope.
The Job tree is always there. The only question is whether you designed it or inherited it by accident.