Rust Concurrency: Threads, Channels, Arc and Mutex, and the Errors You Will Hit
Key takeaways
Rust prevents data races at compile time through the Send and Sync traits. This article walks through threads, channels, Arc<Mutex<T>>, atomics and scoped threads, the real compiler errors each one produces, and the bugs the compiler cannot catch: deadlocks, hung channel loops and poisoned locks.
Introduction
Most languages let you share memory between threads and then hope you remembered every lock. Rust takes a different approach: the type system encodes which values may cross a thread boundary (Send) and which may be shared by reference between threads (Sync). If your code compiles in safe Rust, it has no data races. That is a real guarantee, but a narrow one. It does not stop deadlocks, channels that never close, or holding a lock far longer than you meant to.
This article covers the standard library tools (std::thread, std::sync::mpsc, Arc, Mutex, atomics, and thread::scope) with the actual compiler errors each one tends to produce and the runtime bugs the compiler cannot see. All code below compiles on stable Rust (tested with rustc 1.89).
Spawning threads and why 'static shows up
use std::thread;
use std::time::Duration;
fn main() {
let handle = thread::spawn(|| {
for i in 1..10 {
println!("thread: {i}");
thread::sleep(Duration::from_millis(100));
}
});
for i in 1..5 {
println!("main: {i}");
thread::sleep(Duration::from_millis(100));
}
handle.join().unwrap();
}
thread::spawn creates an OS thread and returns a JoinHandle<T>. join() blocks until the thread finishes and returns Result<T, Box<dyn Any + Send>>; the Err case means the thread panicked. If you never call join() and main returns, the process exits and the spawned thread is killed mid-loop, which is why the output above is cut short without the final join.
The signature is the important part:
pub fn spawn<F, T>(f: F) -> JoinHandle<T>
where
F: FnOnce() -> T + Send + 'static,
T: Send + 'static;
'static means the closure may not hold references to anything that could be freed before the thread ends. The compiler cannot know when an unjoined thread will finish, so it assumes “possibly never”. That is exactly what this error is telling you:
let v = vec![1, 2, 3];
let handle = thread::spawn(|| {
println!("{:?}", v);
});
error[E0373]: closure may outlive the current function, but it borrows `v`, which is owned by the current function
--> src/main.rs:4:32
|
4 | let handle = thread::spawn(|| {
| ^^ may outlive borrowed value `v`
5 | println!("{:?}", v);
| - `v` is borrowed here
|
note: function requires argument type to outlive `'static`
help: to force the closure to take ownership of `v` (and any other referenced variables), use the `move` keyword
There are three honest fixes, and choosing between them is the actual design decision:
move: the thread ownsv. Afterwardsvis gone frommain, and using it givesE0382: borrow of moved value.Arc: several threads need the same data. Clone theArc(cheap, it only bumps a counter) and move each clone in.thread::scope(stable since Rust 1.63): the threads are guaranteed to be joined before the scope returns, so borrowing is allowed. This is usually the cleanest option when the work has a clear start and end.
use std::thread;
fn main() {
let v = vec![1, 2, 3];
thread::scope(|s| {
s.spawn(|| println!("borrowed: {:?}", v));
s.spawn(|| println!("also borrowed: {}", v.len()));
}); // both threads are joined here
println!("still mine: {:?}", v);
}
Send and Sync: what the compiler is actually checking
Send: ownership of a value can be moved to another thread.Sync: a shared reference&Tcan be used from several threads at once. Formally,T: Syncexactly when&T: Send.
Both are auto traits: a struct is Send/Sync if all its fields are. A few types are deliberately excluded, and the reasons tell you a lot about the design:
| Type | Send | Sync | Why |
|---|---|---|---|
Rc<T> | no | no | non-atomic reference count |
Arc<T> | yes (if T: Send + Sync) | yes (if T: Send + Sync) | atomic reference count |
RefCell<T> | yes (if T: Send) | no | borrow flag is not thread-safe |
Mutex<T> | yes (if T: Send) | yes (if T: Send) | the lock provides the synchronization |
MutexGuard<'_, T> | no | yes (if T: Sync) | some platforms require unlocking on the locking thread |
| raw pointers | no | no | the compiler cannot reason about them |
The classic way to meet this table is to write single-threaded code with Rc and then try to parallelize it:
let counter = Rc::new(Mutex::new(0));
let c = Rc::clone(&counter);
thread::spawn(move || {
*c.lock().unwrap() += 1;
});
error[E0277]: `Rc<Mutex<i32>>` cannot be sent between threads safely
= help: within `{closure@src/main.rs:7:27: 7:34}`, the trait `Send` is not implemented for `Rc<Mutex<i32>>`
note: required by a bound in `spawn`
The Mutex inside does not help, because the problem is not the integer but the reference count of the Rc itself. Replace Rc with Arc and it compiles. Note that Arc<T> alone only gives shared read access; for mutation you need something with interior mutability that is Sync, such as Mutex, RwLock or an atomic.
The same trait shows up in async code: holding a std::sync::MutexGuard across an .await inside tokio::spawn fails with “future cannot be sent between threads safely”, because the guard is not Send and the future may resume on a different worker thread. The fix is to drop the guard before the .await, or use tokio::sync::Mutex if you genuinely need to hold it across the await point.
Channels: message passing with mpsc
mpsc stands for multiple producer, single consumer. The Sender can be cloned; the Receiver cannot.
use std::sync::mpsc;
use std::thread;
fn main() {
let (tx, rx) = mpsc::channel();
for id in 0..3 {
let tx = tx.clone();
thread::spawn(move || {
tx.send(format!("worker {id} done")).unwrap();
});
}
drop(tx); // without this line, the loop below never ends
for msg in rx {
println!("{msg}");
}
println!("all senders gone, loop ended");
}
send moves the value into the channel, so after tx.send(val) the sending thread can no longer touch val. That is ownership doing the synchronization for you: there is never a moment where both threads can reach the same String.
The receiving side has three modes, and choosing the wrong one is a common source of bugs:
recv()blocks until a message arrives, and returnsErr(RecvError)once all senders are dropped and the buffer is empty.try_recv()never blocks; it returnsErr(TryRecvError::Empty)orErr(TryRecvError::Disconnected).recv_timeout(d)blocks for at mostd.
Iterating for msg in rx is just repeated recv(), which means the loop ends only when every Sender is gone.
The version of this I have hit most is exactly the commented line above. You clone tx for each worker, move the clones into the threads, and the workers finish, but main still owns the original tx. The channel is technically still open, so the for loop sits there forever. There is no error and no warning; the program just looks hung. Whenever a receive loop “never finishes”, the first thing I check now is which senders are still alive.
Bounded channels and backpressure
mpsc::channel() is unbounded: a fast producer and a slow consumer will grow the buffer until memory runs out. mpsc::sync_channel(n) caps the buffer at n messages and makes send block when it is full:
use std::sync::mpsc;
use std::thread;
fn main() {
let (tx, rx) = mpsc::sync_channel::<u32>(2);
let producer = thread::spawn(move || {
for i in 0..5 {
tx.send(i).unwrap(); // blocks while 2 items are waiting
println!("sent {i}");
}
});
for v in rx {
println!("got {v}");
}
producer.join().unwrap();
}
sync_channel(0) is a rendezvous channel: every send waits until a receiver takes the value. Since Rust 1.67 the standard library’s mpsc is implemented on top of the crossbeam-channel design, so performance is fine for most uses. If you need multiple consumers (a work queue), select! over several channels, or a cloneable receiver, reach for the crossbeam-channel crate or wrap the receiver in Arc<Mutex<Receiver<T>>> for simple cases.
Shared state with Arc<Mutex<T>>
use std::sync::{Arc, Mutex};
use std::thread;
fn main() {
let counter = Arc::new(Mutex::new(0));
let mut handles = vec![];
for _ in 0..10 {
let counter = Arc::clone(&counter);
handles.push(thread::spawn(move || {
let mut num = counter.lock().unwrap();
*num += 1;
})); // guard dropped here, lock released
}
for handle in handles {
handle.join().unwrap();
}
println!("result: {}", *counter.lock().unwrap()); // 10
}
lock() returns a MutexGuard, which derefs to the inner value and unlocks when it is dropped. That is the whole point of the design: you cannot reach the data without holding the lock, and you cannot forget to unlock. The data lives inside the mutex rather than next to it, which is the key difference from C++ std::mutex, where nothing links the mutex to the variables it is supposed to protect.
Where the guard actually gets dropped
Because unlocking is tied to drop, the question “how long is the lock held” becomes “how long does the guard live”, and temporaries make that less obvious than it looks. In both match and if let, a temporary guard created in the scrutinee lives for the entire body:
let queue = Mutex::new(vec![1, 2, 3]);
match queue.lock().unwrap().pop() {
Some(job) => {
// the guard from the line above is STILL alive here
queue.lock().unwrap().push(job * 10); // deadlock
}
None => {}
}
std::sync::Mutex is not reentrant, so locking it again on the same thread deadlocks (or panics, depending on the platform; the docs leave it unspecified). The Rust 2024 edition shortened the temporary’s lifetime for the else branch of if let, but inside the matched arm or the if let body the guard is still held in every edition. The fix is to end the statement first:
let next = queue.lock().unwrap().pop(); // guard dropped at the semicolon
if let Some(job) = next {
queue.lock().unwrap().push(job * 10); // fine
}
This bug is sneaky because the code reads correctly: each line only locks once. When a Rust service “freezes” under load with no panic, a re-lock through a long-lived temporary guard, or a guard held while calling a function that locks the same mutex, is where I look first, ahead of anything more exotic.
Poisoning
If a thread panics while holding a MutexGuard, the mutex is marked poisoned and every later lock() returns Err(PoisonError). That is why every example calls .unwrap(): it turns “another thread crashed mid-update” into a panic here too. Poisoning exists because the data may have been left half-modified. If you know the data is still valid, you can recover it:
match data.lock() {
Ok(v) => println!("ok: {:?}", *v),
Err(poisoned) => {
let v = poisoned.into_inner(); // take the guard anyway
println!("recovered: {:?}", *v);
}
}
Blindly unwrapping means one panicking worker can take down every other thread that touches the same lock. Whether that cascade is what you want depends on whether a half-applied update is worse than a crash.
Other tools for shared state
RwLock<T>: many readers or one writer. It helps only when reads are frequent, long and rarely overlap with writes; for short critical sections aMutexis often just as fast.- Atomics (
AtomicUsize,AtomicBool, …): for a single counter or flag, skip the lock entirely.
use std::sync::atomic::{AtomicUsize, Ordering};
use std::sync::Arc;
use std::thread;
let hits = Arc::new(AtomicUsize::new(0));
let handles: Vec<_> = (0..8)
.map(|_| {
let hits = Arc::clone(&hits);
thread::spawn(move || {
for _ in 0..1000 {
hits.fetch_add(1, Ordering::Relaxed);
}
})
})
.collect();
for h in handles {
h.join().unwrap();
}
assert_eq!(hits.load(Ordering::Relaxed), 8000);
Relaxed is correct for an independent counter. It is not correct when the atomic guards other data (for example “set ready = true after writing the buffer”); that needs Release on the store and Acquire on the load.
OnceLock<T>/LazyLock<T>: one-time initialization of global data, replacing most uses of thelazy_staticcrate.
A parallel sum, done three ways
A common first attempt splits the data, has each thread lock a shared Mutex<i32> and add its partial sum. It works, but the mutex is unnecessary: each thread can simply return its result through join(). With thread::scope, the threads can also borrow slices of the input instead of copying it:
use std::thread;
fn parallel_sum(nums: &[i64]) -> i64 {
let workers = thread::available_parallelism()
.map(|n| n.get())
.unwrap_or(4);
let chunk_size = nums.len().div_ceil(workers).max(1);
thread::scope(|s| {
let handles: Vec<_> = nums
.chunks(chunk_size)
.map(|chunk| s.spawn(move || chunk.iter().sum::<i64>()))
.collect();
handles.into_iter().map(|h| h.join().unwrap()).sum()
})
}
fn main() {
let nums: Vec<i64> = (1..=1_000_000).collect();
println!("sum = {}", parallel_sum(&nums)); // 500000500000
println!("nums still usable: {}", nums.len());
}
Details worth noticing:
chunks()handles the remainder for you, so there is no special case for the last thread..max(1)avoids a panic, becausechunks(0)panics on an empty input.- The element type is
i64. Withi32, the sum of 1..=1,000,000 overflows; in debug builds that panics inside the worker thread, which then shows up as anErrfromjoin(). - For a sum this small, spawning threads can easily cost more than the addition itself. Measure before assuming parallel is faster.
In real code, the third way is usually the right one. The rayon crate does the chunking and work-stealing for you:
use rayon::prelude::*;
fn parallel_sum(nums: &[i64]) -> i64 {
nums.par_iter().sum()
}
What the compiler will not catch
“Fearless concurrency” refers to data races. These are still your problem:
- Deadlocks from lock ordering. Thread A locks
accountsthenlog; thread B lockslogthenaccounts. Pick one global order and document it, or restructure so no code path needs two locks at once. - Long critical sections. Holding a guard while doing file or network I/O serializes every thread behind that I/O. Copy what you need out of the lock, drop the guard, then do the slow work.
- Logical races.
if map.lock().unwrap().contains_key(k) { map.lock().unwrap().insert(...) }takes the lock twice; another thread can insert between the check and the insert. Do the check and the update under a single guard (or use theentryAPI). - Unbounded queues. An unbounded channel with a slow consumer is a memory leak with extra steps.
- Detached threads at exit. Threads that were never joined are killed when
mainreturns, so buffered writes or cleanup in those threads may never happen.
Choosing between threads, rayon and async
| Situation | Reasonable default |
|---|---|
| CPU-bound work over a collection | rayon (par_iter) |
| A few long-lived workers, no extra dependencies | std::thread + channels |
| Borrowing local data for a bounded parallel step | std::thread::scope |
| Thousands of concurrent network connections | Tokio or another async runtime |
| Crash isolation between components | separate processes |
Async and threads are not rivals: async runtimes such as Tokio run tasks on a pool of OS threads, and blocking or CPU-heavy work inside an async task should go through tokio::task::spawn_blocking or a rayon pool so it does not stall the runtime’s worker threads.
Further reading
The Rust Book, chapter 16: Fearless Concurrency, the std::sync documentation, and the rayon documentation.