C++ std::async와 launch 정책: future, deferred 실행, 흔한 함정

이 글의 핵심

std::async의 반환값을 받지 않으면 임시 future의 소멸자가 작업 완료를 기다려서 비동기 코드가 사실상 순차 실행됩니다. 정책을 생략했을 때의 동작, get을 두 번 호출할 때의 오류, 데드락이 생기는 패턴을 짚고, std::thread와 비교해 언제 async를 쓰는 편이 나은지 판단할 수 있게 합니다.

들어가며

std::async는 함수를 비동기로 실행하고 std::future로 결과를 받는 C++11 API입니다. std::thread보다 간결하며, launch 정책으로 실행 시점을 제어할 수 있습니다.

std::thread와 가장 크게 다른 점은 결과와 예외를 돌려받는 통로가 기본으로 있다는 것입니다. std::thread로 실행한 함수의 반환값은 버려지므로 공유 변수와 동기화를 직접 만들어야 하고, 함수가 예외를 던지면 std::terminate()로 프로그램 전체가 종료됩니다. std::async는 반환값과 예외를 모두 std::future에 담아 두었다가 get()에서 돌려주므로, “작업 하나를 백그라운드에 맡기고 나중에 결과를 받는” 형태라면 가장 짧고 안전하게 쓸 수 있습니다. 대신 스레드 수를 제어하거나 작업을 취소하는 기능은 없어서, 그런 요구가 생기면 스레드 풀이나 다른 도구로 넘어가야 합니다.


기본 개념

std::async란?

std::async는 함수를 비동기로 실행하고 std::future로 결과를 받습니다.

#include <future>
#include <iostream>
int compute(int x) {
    return x * x;
}
int main() {
    auto future = std::async(compute, 10);
    
    int result = future.get();  // 결과 대기
    std::cout << "결과: " << result << std::endl;  // 100
    
    return 0;
}

launch 정책

정책실행 시점스레드 생성사용 시나리오
launch::async즉시✅ 새 스레드CPU 집약적 작업
launch::deferredget() 호출 시❌ 현재 스레드조건부 실행
async | deferred구현 의존자동 선택일반적 사용 (기본)

정책을 생략한 기본값(async | deferred)은 “구현이 알아서 고른다”는 뜻이라 몇 가지 보장이 사라집니다. 작업이 다른 스레드에서 돈다는 보장이 없으므로 thread_local 변수가 호출자의 것일 수도 있고, deferred로 선택되면 get()이나 wait()를 부르지 않는 한 영원히 실행되지 않습니다. 실제로 GCC(libstdc++)와 Clang(libc++)은 기본 정책에서 대개 새 스레드를 만들지만 이것은 구현의 선택일 뿐입니다. Scott Meyers가 Effective Modern C++ 항목 36에서 “비동기성이 필수라면 std::launch::async를 명시하라”고 권하는 이유가 이것이며, 이 글의 예제도 대부분 정책을 명시합니다.


실전 구현

기본 사용

#include <future>
#include <iostream>
#include <thread>
#include <chrono>
int compute(int x) {
    std::this_thread::sleep_for(std::chrono::seconds(1));
    return x * x;
}
int main() {
    auto future = std::async(compute, 10);
    
    std::cout << "계산 중..." << std::endl;
    
    int result = future.get();  // 대기
    std::cout << "결과: " << result << std::endl;  // 100
    
    return 0;
}

launch::async - 즉시 실행

#include <future>
#include <iostream>
#include <thread>
int main() {
    auto future = std::async(std::launch::async, []() {
        std::cout << "비동기 스레드 ID: " 
                  << std::this_thread::get_id() << std::endl;
        return 42;
    });
    
    std::cout << "메인 스레드 ID: " 
              << std::this_thread::get_id() << std::endl;
    
    int result = future.get();
    std::cout << "결과: " << result << std::endl;
    
    return 0;
}

출력:

메인 스레드 ID: 140735268771840
비동기 스레드 ID: 123145307557888
결과: 42

launch::deferred - 지연 실행

#include <future>
#include <iostream>
#include <thread>
int main() {
    auto future = std::async(std::launch::deferred, []() {
        std::cout << "지연 실행 스레드 ID: " 
                  << std::this_thread::get_id() << std::endl;
        return 42;
    });
    
    std::cout << "메인 스레드 ID: " 
              << std::this_thread::get_id() << std::endl;
    
    std::cout << "get 호출 전" << std::endl;
    int result = future.get();  // 이때 실행
    std::cout << "결과: " << result << std::endl;
    
    return 0;
}

출력:

메인 스레드 ID: 140735268771840
get 호출 전
지연 실행 스레드 ID: 140735268771840
결과: 42

주의: 같은 스레드에서 실행됨

deferred는 “나중에 다른 스레드에서”가 아니라 “나중에 결과를 요청한 스레드에서” 실행됩니다. 그래서 동시성이 전혀 없고, 결과가 필요 없어지면 비용도 들지 않는 지연 계산(lazy evaluation)에 가깝습니다. 계산 비용이 크지만 조건에 따라 결과를 쓰지 않을 수도 있는 값을 준비할 때 쓸 수 있습니다. 한 가지 함정은 wait_for입니다. deferred future에 wait_for를 호출하면 작업을 실행하지 않고 즉시 std::future_status::deferred를 돌려주므로, while (f.wait_for(10ms) != std::future_status::ready) 같은 폴링 루프는 무한 루프가 됩니다. 기본 정책으로 만든 future를 폴링한다면 deferred 상태도 반드시 처리해야 합니다.


여러 비동기 작업

#include <future>
#include <iostream>
#include <thread>
#include <chrono>
int compute1() {
    std::this_thread::sleep_for(std::chrono::seconds(1));
    return 10;
}
int compute2() {
    std::this_thread::sleep_for(std::chrono::seconds(1));
    return 20;
}
int compute3() {
    std::this_thread::sleep_for(std::chrono::seconds(1));
    return 30;
}
int main() {
    auto start = std::chrono::high_resolution_clock::now();
    
    auto f1 = std::async(std::launch::async, compute1);
    auto f2 = std::async(std::launch::async, compute2);
    auto f3 = std::async(std::launch::async, compute3);
    
    int total = f1.get() + f2.get() + f3.get();
    
    auto end = std::chrono::high_resolution_clock::now();
    auto duration = std::chrono::duration_cast<std::chrono::seconds>(end - start).count();
    
    std::cout << "총합: " << total << std::endl;      // 60
    std::cout << "시간: " << duration << "초" << std::endl;  // 1초
    
    return 0;
}

결과: 위 예제는 각 작업이 1초 동안 sleep만 하므로, 순차로 호출하면 약 3초, 세 작업을 동시에 띄우면 약 1초가 걸립니다. 대기 시간이 겹쳐지기 때문이며, 실제 CPU 계산이라면 코어 수와 스레드 생성 비용에 따라 개선 폭이 이보다 작아집니다.


예외 처리

#include <future>
#include <iostream>
#include <stdexcept>
int divide(int a, int b) {
    if (b == 0) {
        throw std::runtime_error("0으로 나눌 수 없음");
    }
    return a / b;
}
int main() {
    auto future = std::async(divide, 10, 0);
    
    try {
        int result = future.get();  // 예외 재던지기
    } catch (const std::exception& e) {
        std::cout << "예외: " << e.what() << std::endl;
    }
    
    return 0;
}

작업 스레드에서 던져진 예외는 std::exception_ptr로 공유 상태에 저장되었다가 get()에서 원래 타입 그대로 다시 던져집니다. 그래서 catch (const std::runtime_error&)로도 잡을 수 있고, 스택 트레이스는 작업 스레드가 아니라 get()을 호출한 위치를 가리킵니다. get()을 끝내 호출하지 않으면 예외는 조용히 버려지므로, 실패 여부가 중요한 작업이라면 모든 future에 대해 get()을 호출하는 경로를 보장해야 합니다.


future 상태 확인

#include <future>
#include <iostream>
#include <thread>
#include <chrono>
int longCompute() {
    std::this_thread::sleep_for(std::chrono::seconds(3));
    return 42;
}
int main() {
    auto future = std::async(std::launch::async, longCompute);
    
    // 타임아웃
    auto status = future.wait_for(std::chrono::seconds(1));
    
    if (status == std::future_status::ready) {
        std::cout << "완료" << std::endl;
    } else if (status == std::future_status::timeout) {
        std::cout << "타임아웃 (아직 실행 중)" << std::endl;
    }
    
    // 완료 대기
    future.wait();
    std::cout << "결과: " << future.get() << std::endl;
    
    return 0;
}

고급 활용

병렬 다운로드

#include <future>
#include <vector>
#include <string>
#include <iostream>
#include <thread>
#include <chrono>
std::string downloadFile(const std::string& url) {
    std::this_thread::sleep_for(std::chrono::seconds(1));
    return "Data from " + url;
}
int main() {
    std::vector<std::string> urls = {
        "http://example.com/file1",
        "http://example.com/file2",
        "http://example.com/file3"
    };
    
    std::vector<std::future<std::string>> futures;
    
    for (const auto& url : urls) {
        futures.push_back(std::async(std::launch::async, downloadFile, url));
    }
    
    for (auto& future : futures) {
        std::cout << future.get() << std::endl;
    }
    
    return 0;
}

URL이 3개일 때는 문제가 없지만, 이 패턴을 그대로 URL 1만 개에 적용하면 launch::async가 작업마다 스레드를 하나씩 만들어 스레드 1만 개가 동시에 생깁니다. 스레드마다 스택(Linux 기본 8MB 예약)이 잡히고, 스레드 수 한도에 걸리면 std::system_error: Resource temporarily unavailable 예외가 납니다. 저는 이 코드를 배치 도구에 옮겼다가 입력이 커진 날 정확히 이 에러로 실패하는 것을 본 적이 있습니다. 작업 수가 입력에 따라 늘어난다면 동시에 띄우는 개수를 제한하거나(예: 16개씩 묶어 기다리기), 고정 크기 스레드 풀을 쓰는 것이 맞습니다. 결과를 완료된 순서로 처리하는 API(when_any)도 표준에는 없어서, 위 코드는 첫 번째 다운로드가 가장 느리면 나머지가 끝나도 기다립니다.

병렬 맵 리듀스

#include <future>
#include <vector>
#include <numeric>
#include <iostream>
int sum_range(const std::vector<int>& data, size_t start, size_t end) {
    return std::accumulate(data.begin() + start, data.begin() + end, 0);
}
int parallel_sum(const std::vector<int>& data, size_t num_threads) {
    size_t chunk_size = data.size() / num_threads;
    std::vector<std::future<int>> futures;
    
    for (size_t i = 0; i < num_threads; ++i) {
        size_t start = i * chunk_size;
        size_t end = (i == num_threads - 1) ? data.size() : (i + 1) * chunk_size;
        
        futures.push_back(std::async(std::launch::async, sum_range, 
                                     std::ref(data), start, end));
    }
    
    int total = 0;
    for (auto& future : futures) {
        total += future.get();
    }
    
    return total;
}
int main() {
    std::vector<int> data(10000000, 1);
    
    int sum = parallel_sum(data, 4);
    std::cout << "합계: " << sum << std::endl;  // 10000000
    
    return 0;
}

여기서 std::ref(data)는 생략하면 안 되는 부분입니다. std::async는 std::thread와 마찬가지로 인자를 복사해서(decay-copy) 작업에 넘기므로, 함수 매개변수가 const std::vector<int>&여도 data를 그대로 넘기면 작업마다 1000만 개짜리 벡터가 통째로 복사됩니다. 결과는 맞게 나오니 버그로 드러나지 않고 느리기만 해서 놓치기 쉽습니다. std::ref로 참조를 넘기면 복사는 없어지지만, 대신 모든 작업이 끝날 때까지 data가 살아 있어야 한다는 책임이 호출자에게 생깁니다. 이 예제는 get()으로 모두 기다린 뒤 반환하므로 안전합니다. 참고로 이런 단순 병렬 합계는 C++17의 std::reduce(std::execution::par, ...)로도 표현할 수 있습니다.

타임아웃 패턴

#include <future>
#include <iostream>
#include <thread>
#include <chrono>
int slowCompute() {
    std::this_thread::sleep_for(std::chrono::seconds(5));
    return 42;
}
int main() {
    auto future = std::async(std::launch::async, slowCompute);
    
    auto status = future.wait_for(std::chrono::seconds(2));
    
    if (status == std::future_status::ready) {
        std::cout << "결과: " << future.get() << std::endl;
    } else {
        std::cout << "타임아웃: 기본값 사용" << std::endl;
        // future는 계속 실행 중 (취소 불가)
        // 기본값 반환 또는 다른 처리
    }
    
    return 0;  // ⚠️ 여기서 future 소멸자가 남은 3초를 기다림
}

이 “타임아웃”은 기다리기를 포기하는 것일 뿐 작업을 멈추지 않습니다. 더 큰 문제는 main이 끝날 때 future의 소멸자가 실행된다는 점입니다. std::async가 만든 future는 소멸될 때 작업이 끝날 때까지 블로킹하므로, 이 프로그램은 “타임아웃: 기본값 사용”을 출력한 뒤에도 남은 3초를 더 기다렸다가 종료합니다. 요청 처리 함수 안에서 이 패턴을 쓰면 응답은 2초 만에 보냈는데 함수 반환이 5초 뒤에 일어나는 식으로 드러납니다. 정말 작업을 중단해야 한다면 작업 함수가 주기적으로 확인하는 std::atomic<bool> 플래그나 C++20 std::stop_token을 넘기는 협력적 취소를 직접 구현해야 합니다.


성능 비교

async vs thread

테스트: 간단한 계산 작업

방식코드 복잡도결과 반환예외 처리오버헤드
std::async낮음future자동중간
std::thread높음수동 (공유 변수)수동낮음

병렬 실행 벤치마크

테스트: 위 예제의 3개 작업(각 1초 sleep)

실행 방식걸리는 시간
순차 실행약 3초 (1초 × 3)
병렬 실행 (async)약 1초 (가장 긴 작업 기준)

해석: sleep처럼 대기만 하는 작업은 동시에 기다리면 되므로 작업 수만큼 시간이 줄어듭니다. CPU를 계속 쓰는 계산은 코어 수 이상으로 빨라지지 않고, 작업이 짧으면 std::async가 스레드를 만드는 비용이 이득을 잠식하므로 직접 측정해 보고 적용하세요.

launch 정책 비교

#include <future>
#include <iostream>
#include <thread>
#include <chrono>
int compute() {
    std::this_thread::sleep_for(std::chrono::milliseconds(100));
    return 42;
}
int main() {
    // launch::async
    auto start1 = std::chrono::high_resolution_clock::now();
    auto f1 = std::async(std::launch::async, compute);
    auto end1 = std::chrono::high_resolution_clock::now();
    auto duration1 = std::chrono::duration_cast<std::chrono::microseconds>(end1 - start1).count();
    
    std::cout << "async 생성: " << duration1 << "us" << std::endl;
    // 스레드 생성 비용이 포함됨 (보통 수십 us 수준, OS·환경에 따라 다름)
    
    // launch::deferred
    auto start2 = std::chrono::high_resolution_clock::now();
    auto f2 = std::async(std::launch::deferred, compute);
    auto end2 = std::chrono::high_resolution_clock::now();
    auto duration2 = std::chrono::duration_cast<std::chrono::microseconds>(end2 - start2).count();
    
    std::cout << "deferred 생성: " << duration2 << "us" << std::endl;
    // 공유 상태 할당만 하므로 훨씬 짧음 (지연 실행)
    
    f1.get();
    f2.get();
    
    return 0;
}

실무 사례

사례 1: 웹 서버 - 병렬 요청 처리

#include <future>
#include <vector>
#include <string>
#include <iostream>
#include <thread>
#include <chrono>
struct Request {
    std::string url;
    std::string method;
};
std::string processRequest(const Request& req) {
    std::this_thread::sleep_for(std::chrono::milliseconds(100));
    return "Response from " + req.url;
}
int main() {
    std::vector<Request> requests = {
        {"http://api.example.com/users", "GET"},
        {"http://api.example.com/posts", "GET"},
        {"http://api.example.com/comments", "GET"}
    };
    
    std::vector<std::future<std::string>> futures;
    
    for (const auto& req : requests) {
        futures.push_back(std::async(std::launch::async, processRequest, req));
    }
    
    for (auto& future : futures) {
        std::cout << future.get() << std::endl;
    }
    
    return 0;
}

사례 2: 데이터 처리 - 병렬 파일 읽기

#include <future>
#include <vector>
#include <string>
#include <fstream>
#include <iostream>
std::string readFile(const std::string& filename) {
    std::ifstream file(filename);
    std::string content((std::istreambuf_iterator<char>(file)),
                        std::istreambuf_iterator<char>());
    return content;
}
int main() {
    std::vector<std::string> filenames = {
        "file1.txt",
        "file2.txt",
        "file3.txt"
    };
    
    std::vector<std::future<std::string>> futures;
    
    for (const auto& filename : filenames) {
        futures.push_back(std::async(std::launch::async, readFile, filename));
    }
    
    for (size_t i = 0; i < futures.size(); ++i) {
        try {
            std::string content = futures[i].get();
            std::cout << "파일 " << i << " 크기: " << content.size() << std::endl;
        } catch (const std::exception& e) {
            std::cout << "파일 " << i << " 읽기 실패: " << e.what() << std::endl;
        }
    }
    
    return 0;
}

사례 3: 게임 - 비동기 리소스 로딩

#include <future>
#include <vector>
#include <string>
#include <iostream>
#include <thread>
#include <chrono>
struct Texture {
    std::string name;
    int width, height;
};
Texture loadTexture(const std::string& filename) {
    std::this_thread::sleep_for(std::chrono::milliseconds(200));
    return {filename, 1024, 1024};
}
int main() {
    std::vector<std::string> textures = {
        "player.png",
        "enemy.png",
        "background.png"
    };
    
    std::vector<std::future<Texture>> futures;
    
    for (const auto& filename : textures) {
        futures.push_back(std::async(std::launch::async, loadTexture, filename));
    }
    
    std::cout << "로딩 중..." << std::endl;
    
    std::vector<Texture> loadedTextures;
    for (auto& future : futures) {
        loadedTextures.push_back(future.get());
    }
    
    std::cout << "로딩 완료: " << loadedTextures.size() << "개" << std::endl;
    
    return 0;
}

트러블슈팅

문제 1: future 소멸자 블로킹

증상: 프로그램이 예상치 않게 대기

// ❌ future 무시
std::async(std::launch::async, []() {
    std::this_thread::sleep_for(std::chrono::seconds(5));
});  // 소멸자에서 대기 (블로킹)
std::cout << "다음 작업" << std::endl;  // 5초 후 출력
// ✅ future 저장
auto future = std::async(std::launch::async, []() {
    std::this_thread::sleep_for(std::chrono::seconds(5));
});
std::cout << "다음 작업" << std::endl;  // 즉시 출력
future.get();  // 명시적 대기

이 동작은 C++11 표준화 당시 “비동기 작업이 참조하는 지역 변수보다 오래 살아 댕글링 참조를 만들지 않도록” 정한 규칙입니다. 반환값을 무시하면 문장 끝에서 임시 future가 소멸하며 작업 완료를 기다리므로, for 루프에서 std::async(...)를 반환값 없이 호출하면 모든 작업이 하나씩 순서대로 실행됩니다. C++20부터 std::async에 [[nodiscard]]가 붙어 GCC·Clang·MSVC가 경고를 내 주므로 경고를 켜고 빌드하면 쉽게 찾을 수 있습니다. 주의할 점은 이 블로킹 소멸자가 std::async가 만든 future에만 해당한다는 것입니다. std::promise나 std::packaged_task에서 얻은 future는 소멸할 때 기다리지 않습니다.

문제 2: get 여러 번 호출

증상: std::future_error 예외

auto future = std::async([]() { return 42; });
int r1 = future.get();  // OK
// int r2 = future.get();  // 에러: future_error
// ✅ get은 한 번만
// 여러 스레드에서 공유하려면 shared_future 사용

문제 3: 데드락

증상: 프로그램이 멈춤

#include <future>
#include <mutex>
std::mutex mtx;
// ❌ 데드락 가능
void bad_pattern() {
    std::lock_guard<std::mutex> lock(mtx);
    
    auto future = std::async(std::launch::async, []() {
        std::lock_guard<std::mutex> lock(mtx);  // 데드락
        return 42;
    });
    
    future.get();
}
// ✅ 락 범위 최소화
void good_pattern() {
    auto future = std::async(std::launch::async, []() {
        std::lock_guard<std::mutex> lock(mtx);
        return 42;
    });
    
    future.get();
}

문제 4: 작은 작업의 오버헤드

증상: 병렬 실행이 오히려 느림

// ❌ 작은 작업 (스레드 생성 비용 > 작업 시간)
auto future = std::async(std::launch::async, []() {
    return 1 + 1;  // 너무 간단
});
// ✅ 충분히 큰 작업만 병렬화
auto future = std::async(std::launch::async, []() {
    // 복잡한 계산 (수백 ms 이상)
    return compute_heavy();
});

기준: 작업 시간이 스레드 생성·종료 비용(보통 수십 µs 수준, 환경마다 다름)보다 충분히 커야 합니다. 경계가 애매하면 순차 버전과 병렬 버전을 같은 입력으로 직접 재 보는 것이 가장 확실합니다. 참고로 MSVC의 std::async(std::launch::async, ...)는 새 스레드 대신 내부 스레드 풀(Windows 스레드 풀)을 쓰기 때문에 생성 비용이 작은 대신, 작업 사이에 thread_local 값이 남아 있을 수 있다는 차이가 있습니다.


마무리

std::async는 비동기 작업을 간결하게 표현하고 std::future로 결과를 받을 수 있게 합니다.

핵심 요약

  1. 기본 사용
    • std::async(func, args...)로 비동기 실행
    • future.get()으로 결과 대기
  2. launch 정책
    • launch::async: 즉시 실행 (새 스레드)
    • launch::deferred: 지연 실행 (get 호출 시)
    • 기본: async | deferred (구현 의존)
  3. 예외 처리
    • 예외는 get() 호출 시 재던지기
    • try-catch로 처리
  4. 주의사항
    • future 소멸자는 블로킹
    • get()은 한 번만 호출 가능
    • 작은 작업은 오버헤드 주의

선택 가이드

상황방법
간단한 비동기 작업std::async
결과 반환 필요std::async + future.get()
여러 스레드에서 결과 공유shared_future
세밀한 스레드 제어std::thread

코드 예제 치트시트

// 기본 사용
auto future = std::async(func, args...);
int result = future.get();
// launch::async (즉시 실행)
auto f1 = std::async(std::launch::async, func);
// launch::deferred (지연 실행)
auto f2 = std::async(std::launch::deferred, func);
// 타임아웃
auto status = future.wait_for(std::chrono::seconds(1));
if (status == std::future_status::ready) {
    // 완료
}
// 예외 처리
try {
    future.get();
} catch (const std::exception& e) {
    // 처리
}

다음 단계

참고 자료

  • “C++ Concurrency in Action” - Anthony Williams
  • “Effective Modern C++” - Scott Meyers
  • cppreference: https://en.cppreference.com/w/cpp/thread/async 한 줄 정리: std::async는 비동기 작업을 간결하게 표현하며, launch 정책으로 실행 시점을 제어하고 future로 결과를 안전하게 받습니다.

자주 묻는 질문 (FAQ)

Q. std::async의 반환값을 받지 않았더니 코드가 비동기로 실행되지 않고 멈추는 이유는 무엇인가요?

A. std::async가 반환한 std::future를 변수에 담지 않으면 그 문장이 끝날 때 임시 future가 바로 소멸하고, async로 만든 future의 소멸자는 작업이 끝날 때까지 기다립니다. 그래서 겉으로는 비동기 호출처럼 보여도 실제로는 순차 실행과 같아집니다. future를 변수에 저장해 두고 결과가 필요한 시점에 get()이나 wait()를 호출해야 의도한 병렬 실행이 됩니다.


같이 보면 좋은 글