Kotlin Coroutines vs Threads: Cost, Dispatchers, Structured Concurrency and Cancellation
Key takeaways
Coroutines are cheap to create and suspend because a suspended coroutine is a small heap object, not a thread with its own stack. The post compares the two models, shows where the savings disappear (blocking calls on Dispatchers.IO, CPU-bound work), and covers the cancellation and exception rules that trip people up: CancellationException swallowed by catch blocks and CoroutineExceptionHandler on child coroutines.
Introduction
“Coroutines or threads?” This article compares Kotlin coroutines and Java/OS threads and gives practical defaults.
The question is slightly misleading, because coroutines do not replace threads; they run on them. What changes is what happens while a task waits. A thread that waits for a network response is parked by the OS, holding its stack and its slot in the thread pool. A coroutine that waits at a suspension point saves its local state into a heap object and gives the thread back, so the same few threads can serve thousands of waiting tasks. Everything else in this comparison, the cost numbers, the dispatcher rules and the common mistakes, follows from that one difference and from what happens when code does not actually suspend.
Quick comparison
| Coroutines | Threads | |
|---|---|---|
| Weight | Many thousands+ | Dozens–hundreds typical |
| Memory | Small heap object (captured locals) | Reserved stack (often 1MB) |
| Create cost | Object allocation + queueing | OS thread creation |
| Context switch | User-space (cheap) | Kernel (heavier) |
| Cancellation | Structured + cooperative | Manual / interrupt |
| Default | Prefer | Legacy / special |
| Scheduling | Dispatchers (thread pools) | OS scheduler |
| Blocking | Suspends without blocking | Blocks thread |
How they work
Threads
Threads: OS schedules, large stacks (typically 1MB), kernel transitions for context switches.
// Traditional thread
Thread {
println("Running on: ${Thread.currentThread().name}")
Thread.sleep(1000) // Blocks entire thread
println("Done")
}.start()
Characteristics:
- Scheduled by OS kernel
- Pre-emptive multitasking
- Each thread has its own stack (1MB default on Linux)
- Context switch involves kernel mode transition
Coroutines
Coroutines: suspend without blocking threads; delay frees the worker; state machines resume later—often on thread pools (Dispatchers).
// Coroutine
GlobalScope.launch {
println("Running on: ${Thread.currentThread().name}")
delay(1000) // Suspends, thread is free
println("Done")
}
Characteristics:
- Cooperative multitasking
- Suspend functions release thread
- State machine transformation by compiler
- Resumed on dispatcher thread pool
Under the hood
// This coroutine code:
suspend fun example() {
val result1 = fetchData()
val result2 = processData(result1)
return result2
}
// Becomes state machine (simplified):
class ExampleStateMachine : Continuation<Unit> {
var label = 0
var result1: Data? = null
override fun resumeWith(result: Result<Any?>) {
when (label) {
0 -> {
label = 1
fetchData(this) // Pass continuation
}
1 -> {
result1 = result.getOrThrow() as Data
label = 2
processData(result1, this)
}
2 -> {
// Done
}
}
}
}
This transformation explains two practical rules. First, suspension is only possible at calls to other suspend functions, because those are the points where the compiler can save label and the live locals and return. A regular blocking call such as Thread.sleep, JDBC executeQuery or InputStream.read has no suspension point inside it, so the coroutine holds the thread for the whole duration, exactly like a thread would. Second, the suspend keyword does not make a function asynchronous by itself: a suspend fun that only does CPU work or blocking I/O never suspends at all. The benefit comes from the libraries at the bottom (Ktor, R2DBC, delay, Channel) that really do suspend.
Performance benchmarks
Creation cost
import kotlin.system.measureTimeMillis
fun benchmarkThreads() {
val time = measureTimeMillis {
repeat(10_000) {
Thread {
// Do nothing
}.start()
}
}
println("Threads: ${time}ms") // each start() creates an OS thread
}
fun benchmarkCoroutines() = runBlocking {
val time = measureTimeMillis {
repeat(10_000) {
launch {
// Do nothing
}
}
}
println("Coroutines: ${time}ms") // each launch allocates a small object
}
What the results show: the thread version is far slower and the gap grows with the count. Each Thread.start() asks the OS to create a kernel thread, reserve a stack, and register it with the scheduler. Each launch only allocates a small continuation object on the heap and queues it on an existing dispatcher thread. Exact timings depend on the OS, JVM version, and stack size settings, so run the snippet on your machine; note also that JDK 21+ virtual threads narrow this gap on the thread side.
Context switch cost
fun benchmarkThreadContextSwitch() {
val threads = List(1000) {
Thread {
repeat(1000) {
Thread.yield()
}
}
}
val time = measureTimeMillis {
threads.forEach { it.start() }
threads.forEach { it.join() }
}
println("Thread switches: ${time}ms") // each yield may go through the OS scheduler
}
fun benchmarkCoroutineContextSwitch() = runBlocking {
val time = measureTimeMillis {
repeat(1000) {
launch {
repeat(1000) {
yield()
}
}
}
}
println("Coroutine switches: ${time}ms") // each yield re-queues a continuation in user space
}
Concurrent I/O operations
// 100,000 concurrent HTTP requests
fun withThreads() {
val executor = Executors.newFixedThreadPool(1000) // Limited pool
val time = measureTimeMillis {
repeat(100_000) {
executor.submit {
// Simulate HTTP call
Thread.sleep(100)
}
}
executor.shutdown()
executor.awaitTermination(1, TimeUnit.HOURS)
}
println("Threads: ${time}ms") // ~100 rounds of 100ms: 100,000 tasks / 1,000 threads
}
fun withCoroutines() = runBlocking {
val time = measureTimeMillis {
repeat(100_000) {
launch(Dispatchers.IO) {
delay(100)
}
}
}
println("Coroutines: ${time}ms") // all delays overlap, so roughly one 100ms delay plus overhead
}
This comparison is fair only because the coroutine version uses delay, which suspends. Replace delay(100) with Thread.sleep(100), which is what a blocking HTTP client or JDBC driver effectively does, and the coroutine version is no faster than the thread pool; it is slower. Dispatchers.IO is itself a thread pool limited by default to 64 threads (or the number of cores, if larger), so 100,000 blocking tasks run 64 at a time: roughly 1,563 rounds of 100 ms, far longer than the 1,000-thread executor above. This is the most common disappointment when moving blocking code to coroutines: wrapping it in withContext(Dispatchers.IO) keeps it from freezing other coroutines, but it does not make it scale. Scaling needs a client that suspends (for example Ktor’s client, or a driver with a non-blocking API), or a larger dedicated pool via Dispatchers.IO.limitedParallelism(n).
On JDK 21 and later there is a third option. Virtual threads make blocking calls cheap in the same way suspension does, by unmounting the virtual thread from its carrier while it waits, and a dispatcher built on Executors.newVirtualThreadPerTaskExecutor().asCoroutineDispatcher() lets blocking libraries scale without rewriting them. Coroutines still add structured concurrency and cancellation on top, which virtual threads do not provide by themselves.
Memory overhead
Thread memory
fun threadMemory() {
val runtime = Runtime.getRuntime()
val before = runtime.totalMemory() - runtime.freeMemory()
val threads = List(1000) {
Thread {
Thread.sleep(10000)
}.apply { start() }
}
val after = runtime.totalMemory() - runtime.freeMemory()
println("Memory per thread: ${(after - before) / 1000 / 1024}KB")
// Caveat: thread stacks live in native memory, not the Java heap,
// so this heap-based number mostly misses them. Compare process RSS instead.
}
Coroutine memory
fun coroutineMemory() = runBlocking {
val runtime = Runtime.getRuntime()
val before = runtime.totalMemory() - runtime.freeMemory()
repeat(1000) {
launch {
delay(10000)
}
}
val after = runtime.totalMemory() - runtime.freeMemory()
println("Memory per coroutine: ${(after - before) / 1000}bytes")
// Heap-only measurement: noisy, but in the range of hundreds of bytes to a few KB
}
Summary:
- Thread: each platform thread reserves its own stack (the JVM default on 64-bit Linux is typically 1MB via
-Xss, although the OS only commits pages that are touched) - Coroutine: a suspended coroutine is just its continuation object, holding the local variables live across the suspension point
- That difference is why you can keep hundreds of thousands of suspended coroutines, but not that many platform threads
Structured concurrency
Without structured concurrency (threads)
fun fetchUserData(userId: String): User {
val thread1 = Thread {
// Fetch profile
}
val thread2 = Thread {
// Fetch posts
}
thread1.start()
thread2.start()
// What if one fails? How to cancel both?
// What if parent is cancelled?
thread1.join()
thread2.join()
}
With structured concurrency (coroutines)
suspend fun fetchUserData(userId: String): User = coroutineScope {
val profile = async { fetchProfile(userId) }
val posts = async { fetchPosts(userId) }
// If one fails, both are cancelled
// If parent is cancelled, both are cancelled
User(profile.await(), posts.await())
}
Cancellation propagation
val job = GlobalScope.launch {
val child1 = launch {
delay(1000)
println("Child 1")
}
val child2 = launch {
delay(2000)
println("Child 2")
}
delay(500)
}
delay(100)
job.cancel() // Cancels parent and all children
Cancellation in coroutines is cooperative: cancel() marks the job as cancelled, and the coroutine actually stops the next time it reaches a suspension point (delay, await, send) or checks isActive/ensureActive(). The children above stop immediately because they are inside delay. A coroutine running a tight CPU loop or a blocking call keeps running until it finishes, and a cancelled job whose body is inside Thread.sleep is still occupying its thread. Structured concurrency guarantees that the parent waits for its children to finish; it cannot force blocking code to stop.
The other half of the rule is that cancellation is delivered as a CancellationException. Code that catches Exception catches it too:
launch {
while (true) {
try {
delay(1000)
poll()
} catch (e: Exception) { // ❌ also catches CancellationException
log(e) // the loop continues after cancel()
}
}
}
This loop never ends after job.cancel(): every cancellation is logged and swallowed. It is one of the easiest coroutine bugs to write and one of the hardest to notice, because the symptom is work that quietly continues after a screen closed or a request finished. Rethrow cancellation (catch (e: CancellationException) { throw e } before the general catch), catch only the exception types you expect, or check ensureActive() in the handler. The retry helper in section 10 has the same problem.
Dispatchers explained
Dispatchers.Default
// CPU-bound work
launch(Dispatchers.Default) {
val result = heavyComputation()
}
- Thread pool size: Number of CPU cores (at least 2)
- Use for: CPU-intensive tasks
- Examples: Sorting, parsing, compression
Dispatchers.IO
// I/O-bound work
launch(Dispatchers.IO) {
val data = database.query()
val response = httpClient.get(url)
}
- Thread pool size: 64 or the number of cores, whichever is larger (configurable with the
kotlinx.coroutines.io.parallelismsystem property) - Shares threads with
Dispatchers.Default, so switching between them withwithContextoften avoids an actual thread switch - Use for: Network, disk, database
- Examples: HTTP requests, file I/O, database queries
Dispatchers.Main
// Android UI thread
launch(Dispatchers.Main) {
textView.text = "Updated"
}
- Single thread (UI thread)
- Use for: UI updates
- Platform-specific (Android, JavaFX, Swing)
Custom dispatcher
val customDispatcher = Executors.newFixedThreadPool(4).asCoroutineDispatcher()
launch(customDispatcher) {
// Custom thread pool
}
A dispatcher created from an executor owns real threads, and they are not daemon threads: forgetting customDispatcher.close() keeps the JVM from exiting and leaks the pool when the code runs repeatedly. For “at most N concurrent calls to this database”, Dispatchers.IO.limitedParallelism(4) gives the same limit as a view on the shared IO pool, with nothing to close. Using Dispatchers.Main outside Android or a desktop UI app fails at runtime with IllegalStateException: Module with the Main dispatcher is missing, because the main dispatcher comes from a platform module (kotlinx-coroutines-android, -javafx or -swing).
Real-world examples
Example 1: Parallel API calls
// With threads
fun fetchDataThreads(): Result {
val executor = Executors.newFixedThreadPool(3)
val future1 = executor.submit { api.getUsers() }
val future2 = executor.submit { api.getPosts() }
val future3 = executor.submit { api.getComments() }
val users = future1.get()
val posts = future2.get()
val comments = future3.get()
executor.shutdown()
return Result(users, posts, comments)
}
// With coroutines
suspend fun fetchDataCoroutines(): Result = coroutineScope {
val users = async { api.getUsers() }
val posts = async { api.getPosts() }
val comments = async { api.getComments() }
Result(users.await(), posts.await(), comments.await())
}
Example 2: Producer-consumer
// With threads
class ThreadProducerConsumer {
val queue = LinkedBlockingQueue<Int>()
fun start() {
Thread {
repeat(100) {
queue.put(it)
Thread.sleep(10)
}
}.start()
Thread {
while (true) {
val item = queue.take()
process(item)
}
}.start()
}
}
// With coroutines
class CoroutineProducerConsumer {
val channel = Channel<Int>()
fun start() = CoroutineScope(Dispatchers.Default).launch {
launch {
repeat(100) {
channel.send(it)
delay(10)
}
channel.close()
}
launch {
for (item in channel) {
process(item)
}
}
}
}
Example 3: Timeout handling
// With threads (complex)
fun fetchWithTimeoutThread(url: String): String? {
val future = executor.submit { httpClient.get(url) }
return try {
future.get(5, TimeUnit.SECONDS)
} catch (e: TimeoutException) {
future.cancel(true)
null
}
}
// With coroutines (simple)
suspend fun fetchWithTimeoutCoroutine(url: String): String? {
return withTimeoutOrNull(5000) {
httpClient.get(url)
}
}
Common mistakes
Mistake 1: Blocking in coroutine
// ❌ BAD: Blocks thread
launch(Dispatchers.Default) {
Thread.sleep(1000) // Blocks thread!
}
// ✅ GOOD: Suspends
launch(Dispatchers.Default) {
delay(1000) // Suspends, thread is free
}
Mistake 2: Using GlobalScope
// ❌ BAD: No lifecycle management
GlobalScope.launch {
// Runs forever, no cancellation
}
// ✅ GOOD: Scoped
class MyActivity : CoroutineScope {
override val coroutineContext = Dispatchers.Main + Job()
fun loadData() {
launch {
// Cancelled when activity is destroyed
}
}
fun onDestroy() {
coroutineContext.cancel()
}
}
Implementing CoroutineScope on the class shows the idea, but on Android the lifecycle-aware scopes (lifecycleScope, viewModelScope) do the same job without the risk of forgetting onDestroy. The problem with GlobalScope is not that it is global; it is that nothing owns the coroutines launched in it. They outlive the screen or request that started them, keep references to objects that should have been garbage-collected, and their failures are not reported to any parent. Kotlin marks GlobalScope with @DelicateCoroutinesApi for this reason.
Mistake 3: Wrong dispatcher
// ❌ BAD: CPU work on IO dispatcher
launch(Dispatchers.IO) {
val result = heavyComputation() // Wastes IO thread
}
// ✅ GOOD: Use Default for CPU work
launch(Dispatchers.Default) {
val result = heavyComputation()
}
Mistake 4: Not handling exceptions
// ❌ BAD: Exception cancels the parent scope and its siblings
launch {
throw Exception("Error") // propagates to the parent, then to the uncaught-exception handler
}
// ✅ GOOD: Handle exceptions
launch {
try {
riskyOperation()
} catch (e: Exception) {
handleError(e)
}
}
// Or use CoroutineExceptionHandler
val handler = CoroutineExceptionHandler { _, exception ->
println("Caught: $exception")
}
launch(handler) {
throw Exception("Error")
}
An uncaught exception in launch is not silent. It cancels the parent job, which cancels every sibling, and if it reaches the root it goes to the thread’s uncaught-exception handler, which on Android crashes the app. That is structured concurrency working as designed: one failed child fails the whole unit of work. If siblings should survive each other’s failures, use supervisorScope or a SupervisorJob (which is what viewModelScope uses).
CoroutineExceptionHandler is frequently misplaced. It is only consulted for root coroutines (launched directly in a scope, or children of a supervisor). Passing it to a child launch inside another coroutine does nothing, because the child hands the exception to its parent instead, and the handler never runs. It also does not apply to async, whose exception is stored and rethrown by await(); an async that nobody awaits still fails its parent scope. When in doubt, try/catch inside the coroutine is the explicit and predictable option.
Practical guide
When to use coroutines
✅ Use coroutines for:
- Network / DB (
Dispatchers.IO) - Parallel
async/await - Composable concurrency with clear scopes
- Android UI operations
- Sequential async operations
- Thousands of concurrent tasks
When to use threads
⚠️ Threads may remain for:
- Legacy Java pools
- Blocking code you cannot wrap yet
- Interop constraints
- Libraries that require ExecutorService
- Very simple one-off tasks
Decision flowchart
Need concurrency?
├─ Yes
│ ├─ Kotlin project?
│ │ ├─ Yes → Use coroutines
│ │ └─ No → Use threads/executors
│ └─ Legacy Java?
│ └─ Use threads/executors
└─ No → Sequential code
Side-by-side code
Fetching multiple URLs
// Threads
fun fetchUrlsThreads(urls: List<String>): List<String> {
val executor = Executors.newFixedThreadPool(10)
val futures = urls.map { url ->
executor.submit<String> {
httpClient.get(url)
}
}
val results = futures.map { it.get() }
executor.shutdown()
return results
}
// Coroutines
suspend fun fetchUrlsCoroutines(urls: List<String>): List<String> {
return coroutineScope {
urls.map { url ->
async(Dispatchers.IO) {
httpClient.get(url)
}
}.awaitAll()
}
}
Retry logic
// Threads (complex)
fun retryThread(maxAttempts: Int, block: () -> String): String {
repeat(maxAttempts) { attempt ->
try {
return block()
} catch (e: Exception) {
if (attempt == maxAttempts - 1) throw e
Thread.sleep(1000 * (attempt + 1))
}
}
throw IllegalStateException()
}
// Coroutines (simple)
suspend fun retryCoroutine(maxAttempts: Int, block: suspend () -> String): String {
repeat(maxAttempts) { attempt ->
try {
return block()
} catch (e: Exception) {
if (attempt == maxAttempts - 1) throw e
delay(1000L * (attempt + 1))
}
}
throw IllegalStateException()
}
The coroutine version is shorter, but as written it has the cancellation bug described in section 5: if the caller is cancelled while block() is running, catch (e: Exception) catches the CancellationException, and the function waits and retries instead of stopping. Add catch (e: CancellationException) { throw e } above the general handler, or retry only on specific exception types such as IOException. In the thread version, the equivalent problem is InterruptedException: catching it as a generic Exception loses the interrupt, which is why thread code that catches it should restore the flag with Thread.currentThread().interrupt().
Best practices
Use structured concurrency
// ✅ GOOD: Scoped
suspend fun loadUserData() = coroutineScope {
val profile = async { fetchProfile() }
val posts = async { fetchPosts() }
UserData(profile.await(), posts.await())
}
Choose correct dispatcher
// CPU-bound
launch(Dispatchers.Default) {
val result = complexCalculation()
}
// I/O-bound
launch(Dispatchers.IO) {
val data = database.query()
}
// UI updates
launch(Dispatchers.Main) {
updateUI(data)
}
Handle cancellation
suspend fun longRunningTask() {
repeat(1000) { i ->
ensureActive() // Check cancellation
processItem(i)
}
}
Use withContext for switching
suspend fun loadData(): Data {
val data = withContext(Dispatchers.IO) {
database.query()
}
// Back to original dispatcher
return processData(data)
}
Avoid GlobalScope
// ❌ BAD
GlobalScope.launch { }
// ✅ GOOD
class MyViewModel : ViewModel() {
fun loadData() {
viewModelScope.launch {
// Cancelled when ViewModel is cleared
}
}
}
Frequently Asked Questions (FAQ)
Q. Why does calling Thread.sleep() inside a coroutine stall other coroutines?
A. Because Thread.sleep() blocks the underlying thread, and coroutines on Dispatchers.Default or Dispatchers.Main share a small pool of threads (or a single main thread). While that thread is blocked, no other coroutine scheduled on it can run. Use delay(), which suspends the coroutine and frees the thread; if you must call a blocking API you cannot change, wrap it in withContext(Dispatchers.IO) so it runs on the pool meant for blocking work.
Related Articles
- Rust Concurrency: Threads, Channels, Arc and Mutex
- Python asyncio in Practice: Event Loops and Concurrency Limits
- JavaScript Async Programming: Promises, async/await, and the Event Loop