C++로 10GB 파일 올리기: S3 멀티파트 업로드, MinIO, 재시도·진행률, CDN과 서명 URL

들어가며: “10GB 동영상 업로드 시 메모리 부족 에러가 발생한다”

대용량 파일 업로드의 문제점

가장 단순한 파일 업로드 코드는 전체 파일을 메모리에 로드한 후 전송합니다. 이 방식으로 10GB 동영상을 업로드하면 다음과 같은 일이 생깁니다.

// ❌ 잘못된 방법: 전체 파일을 메모리에 로드
std::ifstream file("large_video.mp4", std::ios::binary);
std::vector<char> buffer(std::istreambuf_iterator<char>(file), {});
// 💥 10GB 메모리 사용! 서버 다운!
upload_to_s3(buffer);

문제점:

  • 메모리 부족 (OOM)

  • 네트워크 끊김 시 처음부터 재시도

  • 업로드 진행률 표시 불가

  • 동시 업로드 시 서버 과부하 해결책: 멀티파트 업로드

  • 파일을 5MB 청크로 분할

  • 각 청크를 독립적으로 업로드

  • 실패한 청크만 재시도

  • 병렬 업로드로 속도 향상 목표:

  • AWS S3 멀티파트 업로드 구현

  • MinIO 로컬 스토리지 통합

  • 재시도 로직 및 에러 처리

  • 진행률 표시 및 취소 기능

  • CDN 연동 및 서명된 URL 생성 요구 환경: C++17 이상, AWS SDK for C++, libcurl


실무에서 겪는 문제 시나리오

시나리오 1: 95% 업로드 후 네트워크 끊김

상황: 10GB 파일을 업로드하다 9.5GB 지점에서 연결이 끊겼습니다. 단일 업로드: 처음부터 다시 10GB 전송 → 이미 보낸 9.5GB가 전부 낭비 멀티파트: 이미 완료된 파트는 그대로 두고, 끊긴 파트와 남은 약 500MB만 이어서 전송

시나리오 2: 동시 100명 업로드 시 서버 다운

상황: 사용자 100명이 동시에 500MB 파일을 업로드합니다. 단일 업로드: 100 × 500MB = 50GB 메모리 → OOM 크래시 멀티파트: 100 × 5MB 청크 = 500MB 메모리 → 정상 동작

시나리오 3: 업로드 중 사용자 취소

상황: 사용자가 5GB 업로드 중 “취소” 버튼을 눌렀습니다. 단일 업로드: 이미 전송된 데이터는 버려지며, 서버 리소스 낭비 멀티파트: AbortMultipartUpload로 즉시 정리, 부분 업로드된 데이터 삭제

시나리오 4: CDN 캐시 갱신 필요

상황: S3에 새 파일을 올렸는데 CDN이 이전 버전을 서빙합니다. 해결: CreateInvalidation으로 해당 경로 캐시 무효화

시나리오 5: 민감 파일 다운로드 URL 유출

상황: 다운로드 링크가 SNS에 공유되어 무단 접근 발생 해결: 서명된 URL(Presigned URL)로 1시간 등 짧은 유효기간 설정

시나리오 2의 계산은 “청크 하나씩만 메모리에 둔다”는 전제에서 성립합니다. 아래 예제처럼 파트를 병렬로 올리면 메모리 사용량은 동시에 진행 중인 파트 수 × 청크 크기가 되므로, 병렬 수를 제한하지 않으면 멀티파트를 써도 메모리가 다시 커집니다. 이 점은 구현 절에서 다시 짚습니다.


시스템 아키텍처

전체 구조

flowchart TB
    Client[클라이언트]
    Server[C++ 서버]
    S3[AWS S3]
    MinIO[MinIO]
    CDN[CloudFront CDN]
    
    Client -->|1. 파일 업로드 요청| Server
    Server -->|2. 청크 분할| Server
    Server -->|3a. 멀티파트 업로드| S3
    Server -->|3b. 로컬 저장| MinIO
    S3 -->|4. CDN 배포| CDN
    CDN -->|5. 빠른 다운로드| Client
    
    style S3 fill:#ff9900
    style MinIO fill:#00bcd4
    style CDN fill:#4caf50

핵심 컴포넌트

// 진행률 콜백 (인터페이스보다 먼저 선언해야 함)
using ProgressCallback = std::function<void(
    size_t uploaded_bytes,
    size_t total_bytes,
    double percentage
)>;
// 파일 스토리지 인터페이스
class IFileStorage {
public:
    virtual ~IFileStorage() = default;
    
    // 파일 업로드 (멀티파트 자동 처리)
    virtual std::string upload(
        const std::string& file_path,
        const std::string& bucket,
        const std::string& key,
        ProgressCallback progress_cb = nullptr
    ) = 0;
    
    // 파일 다운로드
    virtual void download(
        const std::string& bucket,
        const std::string& key,
        const std::string& output_path
    ) = 0;
    
    // 서명된 URL 생성 (임시 다운로드 링크)
    virtual std::string generate_presigned_url(
        const std::string& bucket,
        const std::string& key,
        std::chrono::seconds expiration
    ) = 0;
};

AWS S3 멀티파트 업로드

멀티파트 업로드 흐름

sequenceDiagram
    participant Client
    participant Server
    participant S3
    
    Client->>Server: 파일 업로드 요청
    Server->>S3: 1. CreateMultipartUpload
    S3-->>Server: upload_id
    
    loop 각 청크 (5MB)
        Server->>S3: 2. UploadPart (part_number, data)
        S3-->>Server: ETag
    end
    
    Server->>S3: 3. CompleteMultipartUpload (upload_id, ETags)
    S3-->>Server: 완료
    Server-->>Client: 업로드 성공

S3 클라이언트 구현

#include <aws/core/Aws.h>
#include <aws/s3/S3Client.h>
#include <aws/s3/model/CreateMultipartUploadRequest.h>
#include <aws/s3/model/UploadPartRequest.h>
#include <aws/s3/model/CompleteMultipartUploadRequest.h>
#include <aws/s3/model/AbortMultipartUploadRequest.h>
#include <fstream>
#include <future>
class S3FileStorage : public IFileStorage {
protected:
    Aws::S3::S3Client client_;  // MinIOStorage가 교체할 수 있도록 protected
    static constexpr size_t CHUNK_SIZE = 5 * 1024 * 1024;  // 5MB
    static constexpr size_t MAX_RETRIES = 3;
    
public:
    S3FileStorage(const std::string& region) {
        Aws::Client::ClientConfiguration config;
        config.region = region;
        config.connectTimeoutMs = 30000;
        config.requestTimeoutMs = 60000;
        
        client_ = Aws::S3::S3Client(config);
    }
    
    std::string upload(
        const std::string& file_path,
        const std::string& bucket,
        const std::string& key,
        ProgressCallback progress_cb = nullptr
    ) override {
        // 1. 파일 크기 확인
        auto file_size = std::filesystem::file_size(file_path);
        
        // 2. 단일 업로드 vs 멀티파트 업로드 결정
        if (file_size < CHUNK_SIZE) {
            return simple_upload(file_path, bucket, key);
        }
        
        return multipart_upload(file_path, bucket, key, file_size, progress_cb);
    }
    
private:
    // 단일 업로드 (5MB 미만)
    std::string simple_upload(
        const std::string& file_path,
        const std::string& bucket,
        const std::string& key
    ) {
        Aws::S3::Model::PutObjectRequest request;
        request.SetBucket(bucket);
        request.SetKey(key);
        
        auto input_data = Aws::MakeShared<Aws::FStream>(
            "PutObjectInputStream",
            file_path.c_str(),
            std::ios_base::in | std::ios_base::binary
        );
        
        request.SetBody(input_data);
        
        auto outcome = client_.PutObject(request);
        if (!outcome.IsSuccess()) {
            throw std::runtime_error(
                "Upload failed: " + outcome.GetError().GetMessage()
            );
        }
        
        return "s3://" + bucket + "/" + key;
    }
    
    // 멀티파트 업로드 (5MB 이상)
    std::string multipart_upload(
        const std::string& file_path,
        const std::string& bucket,
        const std::string& key,
        size_t file_size,
        ProgressCallback progress_cb
    ) {
        // 1. 멀티파트 업로드 시작
        auto upload_id = initiate_multipart_upload(bucket, key);
        
        try {
            // 2. 청크 수 계산
            size_t num_chunks = (file_size + CHUNK_SIZE - 1) / CHUNK_SIZE;
            
            // 3. 각 청크 업로드 (병렬 처리)
            std::vector<std::future<Aws::S3::Model::CompletedPart>> futures;
            std::atomic<size_t> uploaded_bytes{0};
            
            for (size_t i = 0; i < num_chunks; ++i) {
                futures.push_back(std::async(
                    std::launch::async,
                    [&, i, upload_id]() {
                        return upload_part_with_retry(
                            file_path, bucket, key, upload_id,
                            i + 1,  // part_number는 1부터 시작
                            i * CHUNK_SIZE,
                            std::min(CHUNK_SIZE, file_size - i * CHUNK_SIZE),
                            uploaded_bytes,
                            file_size,
                            progress_cb
                        );
                    }
                ));
            }
            
            // 4. 모든 청크 완료 대기
            std::vector<Aws::S3::Model::CompletedPart> completed_parts;
            for (auto& future : futures) {
                completed_parts.push_back(future.get());
            }
            
            // 5. part_number 순서로 정렬 (중요!)
            std::sort(completed_parts.begin(), completed_parts.end(),
                 [](const Aws::S3::Model::CompletedPart& a, const Aws::S3::Model::CompletedPart& b) {
                    return a.GetPartNumber() < b.GetPartNumber();
                }
            );
            
            // 6. 멀티파트 업로드 완료
            complete_multipart_upload(bucket, key, upload_id, completed_parts);
            
            return "s3://" + bucket + "/" + key;
            
        } catch (...) {
            // 에러 발생 시 멀티파트 업로드 취소
            abort_multipart_upload(bucket, key, upload_id);
            throw;
        }
    }
    
    // 멀티파트 업로드 시작
    std::string initiate_multipart_upload(
        const std::string& bucket,
        const std::string& key
    ) {
        Aws::S3::Model::CreateMultipartUploadRequest request;
        request.SetBucket(bucket);
        request.SetKey(key);
        
        auto outcome = client_.CreateMultipartUpload(request);
        if (!outcome.IsSuccess()) {
            throw std::runtime_error(
                "Failed to initiate multipart upload: " +
                outcome.GetError().GetMessage()
            );
        }
        
        return outcome.GetResult().GetUploadId();
    }
    
    // 청크 업로드 (재시도 포함)
    Aws::S3::Model::CompletedPart upload_part_with_retry(
        const std::string& file_path,
        const std::string& bucket,
        const std::string& key,
        const std::string& upload_id,
        int part_number,
        size_t offset,
        size_t size,
        std::atomic<size_t>& uploaded_bytes,
        size_t total_size,
        ProgressCallback progress_cb
    ) {
        for (size_t retry = 0; retry < MAX_RETRIES; ++retry) {
            try {
                // 파일에서 청크 읽기
                std::ifstream file(file_path, std::ios::binary);
                file.seekg(offset);
                
                auto buffer = std::make_shared<std::vector<char>>(size);
                file.read(buffer->data(), size);
                
                // UploadPart 요청
                Aws::S3::Model::UploadPartRequest request;
                request.SetBucket(bucket);
                request.SetKey(key);
                request.SetUploadId(upload_id);
                request.SetPartNumber(part_number);
                request.SetContentLength(size);
                
                auto stream = Aws::MakeShared<Aws::StringStream>("UploadPartStream");
                stream->write(buffer->data(), size);
                request.SetBody(stream);
                
                auto outcome = client_.UploadPart(request);
                if (!outcome.IsSuccess()) {
                    throw std::runtime_error(outcome.GetError().GetMessage());
                }
                
                // 진행률 업데이트
                uploaded_bytes += size;
                if (progress_cb) {
                    progress_cb(
                        uploaded_bytes.load(),
                        total_size,
                        100.0 * uploaded_bytes.load() / total_size
                    );
                }
                
                // CompletedPart 반환
                Aws::S3::Model::CompletedPart part;
                part.SetPartNumber(part_number);
                part.SetETag(outcome.GetResult().GetETag());
                return part;
                
            } catch (const std::exception& e) {
                if (retry == MAX_RETRIES - 1) {
                    throw;
                }
                // 지수 백오프 (1초, 2초, 4초)
                std::this_thread::sleep_for(
                    std::chrono::seconds(1 << retry)
                );
            }
        }
        
        throw std::runtime_error("Upload part failed after max retries");
    }
    
    // 멀티파트 업로드 완료
    void complete_multipart_upload(
        const std::string& bucket,
        const std::string& key,
        const std::string& upload_id,
        const std::vector<Aws::S3::Model::CompletedPart>& parts
    ) {
        Aws::S3::Model::CompletedMultipartUpload completed_upload;
        for (const auto& part : parts) {
            completed_upload.AddParts(part);
        }
        
        Aws::S3::Model::CompleteMultipartUploadRequest request;
        request.SetBucket(bucket);
        request.SetKey(key);
        request.SetUploadId(upload_id);
        request.SetMultipartUpload(completed_upload);
        
        auto outcome = client_.CompleteMultipartUpload(request);
        if (!outcome.IsSuccess()) {
            throw std::runtime_error(
                "Failed to complete multipart upload: " +
                outcome.GetError().GetMessage()
            );
        }
    }
    
    // 멀티파트 업로드 취소
    void abort_multipart_upload(
        const std::string& bucket,
        const std::string& key,
        const std::string& upload_id
    ) {
        Aws::S3::Model::AbortMultipartUploadRequest request;
        request.SetBucket(bucket);
        request.SetKey(key);
        request.SetUploadId(upload_id);
        
        client_.AbortMultipartUpload(request);
    }
    
public:
    // 서명된 URL 생성 (1시간 유효)
    std::string generate_presigned_url(
        const std::string& bucket,
        const std::string& key,
        std::chrono::seconds expiration = std::chrono::hours(1)
    ) override {
        return client_.GeneratePresignedUrl(
            bucket, key,
            Aws::Http::HttpMethod::HTTP_GET,
            expiration.count()
        );
    }
};

이 구현에서 가장 먼저 손봐야 할 부분은 병렬도 제한이 없다는 점입니다. for 루프가 청크마다 std::async(std::launch::async, ...)를 호출하므로, 10GB 파일이면 약 2,000개의 스레드가 한꺼번에 만들어지고 각 스레드가 5MB 버퍼를 잡습니다. 게다가 buffer를 StringStream에 다시 써 넣으면서 같은 데이터가 두 벌 존재하므로, 최악의 경우 파일 크기의 두 배 가까운 메모리를 쓰게 되어 멀티파트를 쓰는 이유가 사라집니다. 실제 코드에서는 세마포어(C++20 std::counting_semaphore)나 고정 크기 스레드 풀로 동시 파트 수를 4~16개 정도로 묶고, 버퍼도 한 번만 만들도록 Aws::Utils::Stream::PreallocatedStreamBuf 같은 스트림 래퍼를 쓰는 편이 좋습니다. 처음 이 코드를 운영에 올렸을 때 흔히 겪는 증상이 “작은 파일은 잘 되는데 수 GB 파일에서만 프로세스가 OOM으로 죽는다”는 것인데, 대부분 이 병렬도 문제입니다.

또 진행률 콜백은 여러 업로드 스레드에서 동시에 호출됩니다. uploaded_bytes는 원자적으로 늘지만, 콜백 안에서 std::cout에 쓰면 출력이 뒤섞이고, UI 스레드의 위젯을 직접 건드리면 경쟁 조건이 생깁니다. 콜백에서는 값만 원자 변수나 큐에 넘기고, 화면 갱신은 한 스레드에서 주기적으로 하는 구조가 안전합니다. 재시도 로직도 실패 원인을 구분하지 않고 모든 예외를 재시도하는데, AccessDenied나 NoSuchUpload처럼 다시 해도 결과가 같은 에러까지 백오프하며 기다리는 것은 낭비이므로 SDK 에러의 ShouldRetry() 여부로 거르는 것이 좋습니다.


MinIO 로컬 스토리지

MinIO 설정

# Docker로 MinIO 실행
docker run -p 9000:9000 -p 9001:9001 \
  -e "MINIO_ROOT_USER=minioadmin" \
  -e "MINIO_ROOT_PASSWORD=minioadmin" \
  minio/minio server /data --console-address ":9001"
# 버킷 생성
mc alias set myminio http://localhost:9000 minioadmin minioadmin
mc mb myminio/uploads

MinIO 클라이언트 (S3 호환)

class MinIOStorage : public S3FileStorage {
public:
    MinIOStorage(const std::string& endpoint, int port = 9000)
        : S3FileStorage("us-east-1") {  // MinIO는 리전 무시
        
        Aws::Client::ClientConfiguration config;
        config.endpointOverride = endpoint + ":" + std::to_string(port);
        config.scheme = Aws::Http::Scheme::HTTP;  // HTTPS 사용 시 HTTPS로 변경
        config.verifySSL = false;
        
        // MinIO는 S3 호환 API 사용
        client_ = Aws::S3::S3Client(
            Aws::Auth::AWSCredentials("minioadmin", "minioadmin"),
            config,
            Aws::Client::AWSAuthV4Signer::PayloadSigningPolicy::Never,
            false  // virtual hosting 비활성화
        );
    }
};

MinIO에서 가장 흔한 함정은 주소 방식입니다. AWS SDK는 기본적으로 bucket.host 형태의 가상 호스트 방식 주소를 만드는데, 로컬 MinIO에는 uploads.localhost 같은 DNS 이름이 없어서 연결 자체가 실패하거나 NoSuchBucket이 납니다. 위 코드의 마지막 인자 false가 경로 방식(host/bucket/key)으로 바꾸는 설정이며, 이 인자를 빠뜨리는 실수가 가장 많습니다. 또 minioadmin 기본 계정과 verifySSL = false는 로컬 개발용일 뿐이므로, 공유 환경에 띄우는 MinIO에는 별도 계정과 TLS를 설정해야 합니다.


재시도 로직 및 에러 처리

일반적인 에러와 해결법

에러 1: “Connection timeout”

// 원인: 네트워크 불안정 또는 큰 파일
// 해결: 타임아웃 증가 + 재시도
config.connectTimeoutMs = 60000;  // 60초
config.requestTimeoutMs = 300000;  // 5분

에러 2: “Access Denied”

원인: IAM 권한 부족 해결: 필요한 권한을 IAM 정책에 추가

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "s3:PutObject",
        "s3:GetObject",
        "s3:DeleteObject",
        "s3:AbortMultipartUpload",
        "s3:ListMultipartUploadParts"
      ],
      "Resource": "arn:aws:s3:::my-bucket/*"
    }
  ]
}

에러 3: “EntityTooLarge”

원인: 파트 하나가 5GB를 넘거나, 단일 PutObject로 5GB를 넘는 파일을 올린 경우 (S3 제한) 해결: 청크 크기 조정 (5GB 이상 파일은 반드시 멀티파트)

// 5 * 1024 * 1024 * 1024는 int 연산이라 오버플로가 나므로 ULL 접미사 필요
static constexpr size_t MAX_CHUNK_SIZE = 5ULL * 1024 * 1024 * 1024;  // 5GB
if (chunk_size > MAX_CHUNK_SIZE) {
    chunk_size = MAX_CHUNK_SIZE;
}

에러 4: “InvalidPartOrder”

원인: CompleteMultipartUpload 호출 시 part_number 순서 오류 해결: CompletedPart를 part_number 기준으로 정렬 후 전달

std::sort(completed_parts.begin(), completed_parts.end(),
     [](const Aws::S3::Model::CompletedPart& a, const Aws::S3::Model::CompletedPart& b) {
        return a.GetPartNumber() < b.GetPartNumber();
    }
);

에러 5: “NoSuchUpload”

원인: upload_id가 이미 완료·취소되었거나, 수명 주기 규칙(AbortIncompleteMultipartUpload)으로 정리된 경우 해결: 새 upload_id로 처음부터 재시작

S3는 완료되지 않은 멀티파트 업로드를 스스로 지우지 않습니다. 프로세스가 크래시해 abort_multipart_upload가 호출되지 못하면, 올라간 파트는 버킷 목록(ListObjects)에는 보이지 않으면서 저장 요금은 계속 청구됩니다. 청구서에 설명되지 않는 저장 용량이 쌓여 있다면 ListMultipartUploads로 확인해 볼 만하며, 버킷에 “시작 후 N일이 지난 미완료 멀티파트 업로드 중단” 수명 주기 규칙을 걸어 두는 것이 사실상 필수입니다.


진행률 표시 및 취소

진행률 콜백 구현

#include <iostream>
#include <iomanip>
void print_progress(size_t uploaded, size_t total, double percentage) {
    const int bar_width = 50;
    int pos = static_cast<int>(bar_width * percentage / 100.0);
    
    std::cout << "\r[";
    for (int i = 0; i < bar_width; ++i) {
        if (i < pos) std::cout << "=";
        else if (i == pos) std::cout << ">";
        else std::cout << " ";
    }
    std::cout << "] " << std::fixed << std::setprecision(1) << percentage << "% "
              << "(" << uploaded / 1024 / 1024 << " / " 
              << total / 1024 / 1024 << " MB)" << std::flush;
}
// 사용 예시
storage.upload(
    "large_video.mp4",
    "my-bucket",
    "videos/large_video.mp4",
    print_progress
);

업로드 취소 기능

class CancellableUpload {
    std::atomic<bool> cancelled_{false};
    std::string upload_id_;
    
public:
    void cancel() {
        cancelled_ = true;
    }
    
    bool is_cancelled() const {
        return cancelled_;
    }
    
    // 업로드 중 취소 확인
    void upload_with_cancellation() {
        for (size_t i = 0; i < num_chunks; ++i) {
            if (is_cancelled()) {
                abort_multipart_upload(bucket, key, upload_id_);
                throw std::runtime_error("Upload cancelled by user");
            }
            
            upload_part(i);
        }
    }
};

CDN 연동 및 서명된 URL

CloudFront 배포 설정

#include <aws/cloudfront/CloudFrontClient.h>
#include <aws/cloudfront/model/CreateInvalidationRequest.h>
class CDNManager {
    Aws::CloudFront::CloudFrontClient client_;
    std::string distribution_id_;
    
public:
    CDNManager(const std::string& distribution_id)
        : distribution_id_(distribution_id) {}
    
    // CDN 캐시 무효화
    void invalidate_cache(const std::vector<std::string>& paths) {
        Aws::CloudFront::Model::InvalidationBatch batch;
        
        Aws::CloudFront::Model::Paths invalidation_paths;
        invalidation_paths.SetQuantity(paths.size());
        for (const auto& path : paths) {
            invalidation_paths.AddItems(path);
        }
        
        batch.SetPaths(invalidation_paths);
        batch.SetCallerReference(
            std::to_string(std::time(nullptr))
        );
        
        Aws::CloudFront::Model::CreateInvalidationRequest request;
        request.SetDistributionId(distribution_id_);
        request.SetInvalidationBatch(batch);
        
        auto outcome = client_.CreateInvalidation(request);
        if (!outcome.IsSuccess()) {
            throw std::runtime_error(
                "Cache invalidation failed: " +
                outcome.GetError().GetMessage()
            );
        }
    }
    
    // CDN URL 생성
    std::string get_cdn_url(const std::string& key) {
        return "https://d111111abcdef8.cloudfront.net/" + key;
    }
};

프로덕션 배포

환경 변수 설정

# .env 파일
AWS_ACCESS_KEY_ID=AKIAIOSFODNN7EXAMPLE
AWS_SECRET_ACCESS_KEY=wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY
AWS_REGION=ap-northeast-2
S3_BUCKET=my-production-bucket
CLOUDFRONT_DISTRIBUTION_ID=E1234ABCDEFGH
MINIO_ENDPOINT=http://localhost:9000

예시의 .env처럼 액세스 키를 파일에 두는 방식은 로컬 개발에서만 쓰는 편이 좋습니다. EC2·ECS·EKS에서 돌린다면 IAM 역할(인스턴스 프로파일, IRSA)을 붙이면 SDK의 기본 자격 증명 체인이 임시 자격 증명을 알아서 가져오므로 키를 이미지나 저장소에 둘 필요가 없습니다. 단, 임시 자격 증명으로 만든 서명된 URL은 요청한 유효기간과 상관없이 세션 토큰이 만료되는 시점에 함께 무효가 됩니다. “7일짜리 링크를 만들었는데 몇 시간 뒤 ExpiredToken이 난다”는 문의의 원인이 대부분 이것입니다. SigV4 서명 URL의 최대 유효기간은 7일입니다.

Docker 배포

FROM ubuntu:22.04
RUN apt-get update && apt-get install -y \
    build-essential \
    cmake \
    libcurl4-openssl-dev \
    libssl-dev \
    && rm -rf /var/lib/apt/lists/*
# AWS SDK 설치
RUN git clone --recurse-submodules https://github.com/aws/aws-sdk-cpp && \
    cd aws-sdk-cpp && \
    mkdir build && cd build && \
    cmake .. -DCMAKE_BUILD_TYPE=Release \
             -DBUILD_ONLY="s3;cloudfront" && \
    make -j$(nproc) && \
    make install
WORKDIR /app
COPY . .
RUN cmake -B build -DCMAKE_BUILD_TYPE=Release && \
    cmake --build build --parallel
CMD ["./build/file_storage_server"]

모니터링

#include <prometheus/counter.h>
#include <prometheus/histogram.h>
class StorageMetrics {
    prometheus::Counter& upload_total_;
    prometheus::Counter& upload_failed_;
    prometheus::Histogram& upload_duration_;
    
public:
    void record_upload_success(double duration_seconds) {
        upload_total_.Increment();
        upload_duration_.Observe(duration_seconds);
    }
    
    void record_upload_failure() {
        upload_failed_.Increment();
    }
};

실전 예시

예시 1: 비디오 업로드 서비스

int main() {
    // AWS SDK 초기화
    Aws::SDKOptions options;
    Aws::InitAPI(options);
    
    {
        S3FileStorage storage("ap-northeast-2");
        
        // 10GB 비디오 업로드
        auto start = std::chrono::steady_clock::now();
        
        auto url = storage.upload(
            "/path/to/large_video.mp4",
            "my-videos",
            "uploads/2026/03/large_video.mp4",
             [](size_t uploaded, size_t total, double pct) {
                std::cout << "Progress: " << pct << "%\n";
            }
        );
        
        auto end = std::chrono::steady_clock::now();
        auto duration = std::chrono::duration_cast<std::chrono::seconds>(
            end - start
        ).count();
        
        std::cout << "Upload completed in " << duration << " seconds\n";
        std::cout << "URL: " << url << "\n";
        
        // 서명된 URL 생성 (1시간 유효)
        auto presigned_url = storage.generate_presigned_url(
            "my-videos",
            "uploads/2026/03/large_video.mp4",
            std::chrono::hours(1)
        );
        
        std::cout << "Download URL: " << presigned_url << "\n";
    }
    
    Aws::ShutdownAPI(options);
    return 0;
}

예시 2: 로컬 파일 스토리지 (MinIO 없이 테스트용)

// 로컬 디스크 기반 스토리지 - MinIO/S3 없이 개발 시 사용
class LocalFileStorage : public IFileStorage {
    std::filesystem::path base_path_;
    
public:
    LocalFileStorage(const std::string& base_path) 
        : base_path_(base_path) {
        std::filesystem::create_directories(base_path_);
    }
    
    std::string upload(
        const std::string& file_path,
        const std::string& bucket,
        const std::string& key,
        ProgressCallback progress_cb = nullptr
    ) override {
        auto dest_dir = base_path_ / bucket / std::filesystem::path(key).parent_path();
        std::filesystem::create_directories(dest_dir);
        
        auto dest_path = base_path_ / bucket / key;
        auto file_size = std::filesystem::file_size(file_path);
        
        std::ifstream src(file_path, std::ios::binary);
        std::ofstream dst(dest_path, std::ios::binary);
        
        const size_t buf_size = 64 * 1024;  // 64KB
        std::vector<char> buffer(buf_size);
        size_t copied = 0;
        
        while (src.read(buffer.data(), buf_size) || src.gcount() > 0) {
            auto count = src.gcount();
            dst.write(buffer.data(), count);
            copied += count;
            if (progress_cb) {
                progress_cb(copied, file_size, 100.0 * copied / file_size);
            }
        }
        
        return "file://" + dest_path.string();
    }
    
    void download(const std::string& bucket, const std::string& key,
                  const std::string& output_path) override {
        auto src = base_path_ / bucket / key;
        std::filesystem::copy(src, output_path, 
            std::filesystem::copy_options::overwrite_existing);
    }
    
    std::string generate_presigned_url(const std::string& bucket,
        const std::string& key, std::chrono::seconds) override {
        return "file://" + (base_path_ / bucket / key).string();
    }
};

예시 3: 백업 시스템

class BackupManager {
    S3FileStorage storage_;
    
public:
    void backup_directory(const std::filesystem::path& dir) {
        for (const auto& entry : std::filesystem::recursive_directory_iterator(dir)) {
            if (entry.is_regular_file()) {
                auto relative_path = std::filesystem::relative(entry.path(), dir);
                
                storage_.upload(
                    entry.path().string(),
                    "backups",
                    "daily/" + relative_path.string()
                );
                
                std::cout << "Backed up: " << relative_path << "\n";
            }
        }
    }
};

예시 4: 완전한 파일 업로드 서버 (main 함수)

// file_upload_server.cpp - 빌드 후 실행 가능한 완전한 예제
#include <iostream>
#include <iomanip>
#include <string>
#include <chrono>
int main(int argc, char* argv[]) {
    if (argc < 4) {
        std::cerr << "Usage: " << argv[0] 
                  << " <file_path> <bucket> <key>\n";
        std::cerr << "Example: " << argv[0] 
                  << " video.mp4 my-bucket uploads/video.mp4\n";
        return 1;
    }
    
    std::string file_path = argv[1];
    std::string bucket = argv[2];
    std::string key = argv[3];
    
    Aws::SDKOptions options;
    Aws::InitAPI(options);
    
    try {
        // MinIO 로컬 테스트: MinIOStorage storage("localhost", 9000);
        S3FileStorage storage("ap-northeast-2");
        
        std::cout << "Uploading " << file_path << " to s3://" 
                  << bucket << "/" << key << "\n";
        
        auto start = std::chrono::steady_clock::now();
        
        auto result = storage.upload(file_path, bucket, key,
             [](size_t uploaded, size_t total, double pct) {
                std::cout << "\rProgress: " << std::fixed 
                          << std::setprecision(1) << pct << "% ("
                          << (uploaded / 1024 / 1024) << " / "
                          << (total / 1024 / 1024) << " MB)" << std::flush;
            }
        );
        
        auto end = std::chrono::steady_clock::now();
        auto secs = std::chrono::duration_cast<std::chrono::seconds>(end - start).count();
        
        std::cout << "\nUpload completed in " << secs << " seconds\n";
        std::cout << "Result: " << result << "\n";
        
        auto presigned = storage.generate_presigned_url(
            bucket, key, std::chrono::hours(1));
        std::cout << "Download URL (1h): " << presigned << "\n";
        
    } catch (const std::exception& e) {
        std::cerr << "Error: " << e.what() << "\n";
        Aws::ShutdownAPI(options);
        return 1;
    }
    
    Aws::ShutdownAPI(options);
    return 0;
}

MinIO·타임아웃·스로틀링 에러

에러 6: “SSL certificate problem”

원인: MinIO 로컬 사용 시 HTTPS 검증 실패 해결:

config.verifySSL = false;  // 로컬 MinIO만! 프로덕션에서는 true 유지
config.scheme = Aws::Http::Scheme::HTTP;

에러 7: “RequestTimeout” (대용량 청크)

원인: 5MB 청크 업로드가 60초 내 완료되지 않음 (느린 네트워크) 해결:

config.requestTimeoutMs = 300000;  // 5분
// 또는 청크 크기 축소
static constexpr size_t CHUNK_SIZE = 2 * 1024 * 1024;  // 2MB

에러 8: “SlowDown” (503)

원인: S3 요청 제한 초과 (접두사(prefix)당 초당 약 3,500건의 PUT/POST/DELETE) 해결: 병렬도 조절 + 지수 백오프

// 병렬 업로드 수 제한 (10개 → 5개)
const size_t MAX_CONCURRENT_PARTS = 5;
// 세마포어로 동시 업로드 수 제한

에러 9: “File not found” (업로드 시)

원인: 파일 경로 오류 또는 업로드 중 파일 삭제 해결:

if (!std::filesystem::exists(file_path)) {
    throw std::runtime_error("File not found: " + file_path);
}
auto file_size = std::filesystem::file_size(file_path);
if (file_size == 0) {
    throw std::runtime_error("Empty file: " + file_path);
}

에러 10: “InvalidPart” (멀티파트 완료 시)

원인: CompleteMultipartUpload에 넘긴 ETag가 실제 업로드된 파트와 다름 (따옴표를 직접 떼거나 붙여 가공한 경우, 다른 upload_id의 파트를 섞은 경우 등) 해결: UploadPart 응답의 ETag를 가공하지 않고 그대로 전달


성능 벤치마크

먼저 계산해 볼 한계

업로드 시간의 하한은 파일 크기 ÷ 실제 업로드 대역폭입니다. 예를 들어 100Mbps(약 12MB/s) 회선이라면 10GB 파일은 이론상으로도 약 14분이 걸리고, 청크 크기나 병렬 수를 어떻게 바꿔도 이보다 빨라질 수는 없습니다. 튜닝으로 줄일 수 있는 것은 이 하한과 실제 시간 사이의 낭비입니다.

청크 크기의 트레이드오프

청크 크기장점단점
작게 (5MB, S3 최소값)실패 시 재전송량이 작음, 메모리 사용이 작음파트 수가 많아져 요청당 오버헤드(HTTP 왕복, 서명 계산) 증가. S3는 파트 최대 10,000개라 큰 파일은 청크를 키워야 함
크게 (50MB 이상)요청 수가 줄어 오버헤드 감소실패 시 재전송량이 크고, 병렬 수 × 청크 크기만큼 메모리 필요

권장: 5MB ~ 10MB에서 시작하고, 파일 크기 ÷ 10,000이 이보다 크면 그 값 이상으로 청크를 키웁니다.

병렬도의 트레이드오프

순차 업로드는 파트 하나를 보내고 응답을 기다리는 동안 회선이 놀기 때문에, 특히 RTT가 큰 원격 리전에서 대역폭을 다 쓰지 못합니다. 병렬 업로드는 여러 파트를 동시에 보내 이 대기 시간을 겹쳐서 회선을 채웁니다. 회선이 가득 차면 병렬 수를 더 늘려도 빨라지지 않고, 메모리 사용(병렬 수 × 청크 크기)과 스로틀링(S3의 503 SlowDown) 위험만 커집니다. 적정값은 대역폭과 RTT에 따라 다르므로 4, 8, 16처럼 단계적으로 늘려 가며 처리량이 더 오르지 않는 지점을 찾으세요.

S3 vs MinIO 로컬

로컬 MinIO는 네트워크 왕복이 거의 없고 디스크 속도가 병목이 되므로, 원격 S3보다 훨씬 빠르게 나옵니다. 개발·테스트에서 MinIO로 잰 속도를 운영 S3 성능으로 기대하면 안 되고, 운영 환경의 리전·회선에서 따로 측정해야 합니다.


성능 비교

방식업로드 시간메모리 사용량재시도 가능
단일 업로드회선 한계에 가깝지만 실패하면 처음부터 다시구현에 따라 파일 전체를 버퍼링할 수 있음❌
멀티파트 (순차)파트마다 응답 대기가 끼어 회선을 다 못 씀청크 1개✅
멀티파트 (병렬 N개)대기 시간이 겹쳐져 회선 한계에 가까워짐청크 N개✅

결론: 병렬 멀티파트 업로드는 회선을 더 잘 채우고, 실패한 파트만 다시 보내며, 메모리는 청크 크기 × 병렬 수로 제한됩니다. 대용량 파일에서는 속도보다 재시도 가능성과 메모리 상한이 더 중요한 이유입니다.


프로덕션 패턴

패턴 1: 스토리지 추상화 (환경별 전환)

std::unique_ptr<IFileStorage> create_storage() {
    auto env = std::getenv("STORAGE_TYPE");
    if (env && std::string(env) == "minio") {
        return std::make_unique<MinIOStorage>("minio.local", 9000);
    }
    if (env && std::string(env) == "local") {
        return std::make_unique<LocalFileStorage>("/tmp/uploads");
    }
    return std::make_unique<S3FileStorage>("ap-northeast-2");
}

패턴 2: 업로드 큐 (동시 업로드 제한)

class UploadQueue {
    std::queue<std::function<void()>> queue_;
    std::mutex mutex_;
    std::condition_variable cv_;
    std::atomic<int> active_{0};
    int max_concurrent_{5};
    
public:
    void enqueue(std::function<void()> task) {
        std::unique_lock lock(mutex_);
        queue_.push([this, task]() {
            task();
            active_--;
            cv_.notify_one();
        });
        cv_.notify_one();
    }
    
    void run() {
        while (true) {
            std::unique_lock lock(mutex_);
            cv_.wait(lock, [this] {
                return active_ < max_concurrent_ && !queue_.empty();
            });
            if (queue_.empty()) break;
            auto task = std::move(queue_.front());
            queue_.pop();
            active_++;
            lock.unlock();
            std::thread(task).detach();
        }
    }
};

패턴 3: 업로드 재개 (Resumable Upload)

// 중단된 업로드의 upload_id와 완료된 part 목록을 DB/파일에 저장
struct UploadState {
    std::string upload_id;
    std::vector<std::pair<int, std::string>> completed_parts;
};
// 재시작 시 저장된 state 로드 후 완료되지 않은 part만 업로드

저장해 둔 파트 목록을 그대로 믿기보다는, 재시작할 때 ListParts로 S3에 실제로 올라간 파트와 ETag를 다시 조회하는 편이 정확합니다. 로컬 상태 파일을 쓰기 직전에 크래시하면 S3에는 있는데 로컬 기록에는 없는 파트가 생길 수 있기 때문입니다. 또 재개하기 전에 원본 파일이 바뀌지 않았는지(크기·수정 시각·해시)를 확인해야, 앞부분은 옛 파일이고 뒷부분은 새 파일인 객체가 만들어지는 사고를 막을 수 있습니다.

위 UploadQueue 패턴도 그대로 쓰기에는 손볼 곳이 있습니다. run()의 대기 조건이 “큐가 비어 있지 않을 때”만 깨어나므로 if (queue_.empty()) break;에는 도달하지 못해 종료 방법이 없고, detach()한 스레드는 프로세스 종료 시점에 정리할 방법이 없습니다. 실제로는 종료 플래그를 대기 조건에 넣고, 워커 스레드를 미리 만들어 두는 스레드 풀 형태로 바꾸는 편이 안전합니다.

패턴 4: 파일 검증 (업로드 후)

// MD5/SHA256 해시로 업로드 무결성 검증
#include <openssl/md5.h>
std::string compute_file_hash(const std::string& path) {
    std::ifstream file(path, std::ios::binary);
    unsigned char hash[MD5_DIGEST_LENGTH];
    MD5_CTX ctx;
    MD5_Init(&ctx);
    char buf[8192];
    while (file.read(buf, sizeof(buf)) || file.gcount()) {
        MD5_Update(&ctx, buf, file.gcount());
    }
    MD5_Final(hash, &ctx);
    return bytes_to_hex(hash, MD5_DIGEST_LENGTH);
}
// S3 GetObject의 ETag와 비교 (단일 PutObject + SSE-S3일 때만 ETag == MD5)

MD5·ETag 비교는 조건이 까다롭습니다. 멀티파트로 올린 객체의 ETag는 파일 전체의 MD5가 아니라 각 파트 MD5를 이어 붙인 값의 MD5-파트수 형태라서, 위 함수의 결과와 절대 일치하지 않습니다. SSE-KMS로 암호화된 객체도 ETag가 MD5가 아닙니다. 무결성 검증이 목적이라면 업로드 요청에 ChecksumAlgorithm(SHA256, CRC32 등)을 지정해 S3가 서버 측에서 검증하게 하는 편이 간단합니다. 참고로 MD5_Init 계열 함수는 OpenSSL 3.0에서 deprecated되어 경고가 나므로, 새 코드에서는 EVP_Digest API를 쓰는 것이 좋습니다.

패턴 5: 비용 최적화 (스토리지 클래스)

// 자주 접근하지 않는 파일은 Glacier로 이관
// S3 Lifecycle 규칙으로 30일 후 STANDARD_IA, 90일 후 GLACIER
// 또는 업로드 시 스토리지 클래스 지정
request.SetStorageClass(Aws::S3::Model::StorageClass::STANDARD_IA);

체크리스트

구현 체크리스트

  • AWS SDK 설치 및 설정
  • IAM 권한 설정 (s3:PutObject, s3:GetObject 등)
  • 버킷 생성 및 CORS 설정
  • 멀티파트 업로드 임계값 설정 (5MB)
  • 재시도 정책 설정 (최대 3회, 지수 백오프)
  • 진행률 콜백 구현
  • 에러 로깅 설정
  • 모니터링 설정 (CloudWatch)

프로덕션 체크리스트

  • HTTPS 사용
  • 서명된 URL로 다운로드 보안
  • CDN 연동
  • 캐시 무효화 전략
  • 비용 모니터링 (S3 요금)
  • 백업 정책
  • 로그 보관 정책

파트 크기와 중단된 업로드에서 놓치기 쉬운 것

S3 멀티파트 업로드는 마지막 파트를 제외한 각 파트가 최소 5MiB여야 하고, 업로드 하나당 파트는 최대 10,000개입니다. 그래서 파트 크기를 5MiB로 고정하면 약 48GiB를 넘는 파일은 올릴 수 없습니다. 파트 크기는 상수로 두기보다 파일 크기를 10,000으로 나눈 값과 최솟값 중 큰 쪽을 쓰도록 계산하는 편이 안전합니다.

운영에서 가장 늦게 발견되는 문제는 비용입니다. 클라이언트가 중간에 죽거나 취소해 CompleteMultipartUpload도 AbortMultipartUpload도 호출되지 않으면, 이미 올라간 파트는 객체 목록에 보이지 않은 채 저장 용량으로 계속 과금됩니다. 취소 경로에서 Abort를 호출하는 것과 별개로, 버킷 수명 주기 규칙에 AbortIncompleteMultipartUpload를 설정해 두면 이런 잔여 파트를 자동으로 정리할 수 있습니다. MinIO도 같은 S3 API를 따르므로 로컬 테스트 단계에서 이 경로를 함께 확인해 두세요.


자주 묻는 질문 (FAQ)

Q. 멀티파트 업로드 중 RequestTimeout이 반복되면 어떻게 해야 하나요?

A. 느린 네트워크에서 청크 하나가 제한 시간 안에 다 올라가지 못해 생기는 경우가 많습니다. 청크 크기를 줄이거나 타임아웃을 늘리고, 실패한 파트만 지수 백오프로 재시도하도록 만들면 전체 업로드를 처음부터 다시 하지 않아도 됩니다. 병렬 업로드 수가 너무 많아 대역폭을 나눠 쓰는 경우도 있으므로 동시 파트 수도 함께 조절하세요.

Q. 비용은 얼마나 드나요?

A. AWS S3 요금 (서울 리전 기준):

  • 스토리지: $0.025/GB/월
  • PUT 요청: $0.005/1,000건
  • GET 요청: $0.0004/1,000건
  • 데이터 전송 (out): $0.126/GB 요금은 리전과 시점에 따라 바뀌므로 AWS 요금 페이지에서 최신 값을 확인해야 합니다. 계산해 보면 업로드 쪽은 저렴합니다. 10GB를 5MB 파트로 올리면 PUT 요청이 약 2,000건이라 요청 비용은 1센트 수준이고, 저장 비용도 월 몇십 센트입니다. 비용을 좌우하는 것은 다운로드 전송량입니다. 같은 10GB 파일을 1,000번 내려받으면 전송량이 약 10TB가 되어 인터넷 전송 요금만 1,000달러 단위가 됩니다. 대용량 파일을 많이 배포한다면 CloudFront를 앞에 두어 전송 단가를 낮추고, 서명된 URL로 무단 재배포를 막는 것이 비용 측면에서도 중요합니다.

Q. MinIO vs S3, 어떤 걸 써야 하나요?

A.


같이 보면 좋은 글