Python File Automation | Organize, Rename, and Back Up Files

Key takeaways

Automate file workflows in Python: find and rename files, organize by extension, backups with shutil, duplicate detection, and log cleanup—patterns and code you can reuse.

Introduction

Automating file operations in Python saves a lot of time in real workflows: sorting downloads, rotating logs, pruning old backups, and catching duplicate files are all tasks that are simple in principle but tedious and error-prone when done by hand at scale. A human doing this work eventually skips a step, renames the wrong file, or forgets to check free disk space before a big copy — a script does the same check every single time.

This post assumes you already know the basics of reading and writing files in Python (open, context managers, CSV/JSON serialization). If you need that foundation first, see Python File Handling | Read, Write, CSV, JSON in this series. Here, the focus is different: treating the filesystem itself as the thing being manipulated — finding files by pattern, renaming them safely, moving them between directories, archiving them, and deciding what to keep or delete. That shift in focus is also why the code below leans so heavily on pathlib and shutil rather than raw open() calls.


Choosing your toolkit: os vs. shutil vs. pathlib

Before writing any automation script, it helps to know which of Python’s three main filesystem modules actually does what, because they overlap in confusing ways:

  • os / os.path is the oldest layer, largely a thin wrapper around POSIX system calls. Paths are plain strings, so operations like joining directories require os.path.join() instead of natural operator overloading. It is still the right tool when you need low-level control — for example, os.walk() lets you mutate the dirs list in place to prune subdirectories from a recursive walk before it descends into them, something pathlib.rglob() cannot do.
  • shutil sits on top of os and adds the bulk operations that neither os nor pathlib implement natively: recursive directory copy (copytree), recursive delete (rmtree), archiving (make_archive), and moving across filesystem boundaries (move, which silently falls back to copy+delete when a plain rename would fail). There is still no Path.copytree() in the standard library as of Python 3.12, so any script that needs to duplicate an entire directory tree ends up importing shutil regardless of how “pathlib-first” the rest of the code is.
  • pathlib is the modern, object-oriented interface introduced in Python 3.4. A Path object behaves like a string when you need one (it implements __fspath__), but paths compose with the / operator, glob patterns are a method call away (Path.glob, Path.rglob), and it abstracts away \\ vs / so the same code runs on Windows and Linux without manual string surgery.

In practice, idiomatic modern code uses pathlib for representing and navigating paths, and calls into shutil for the handful of bulk operations it doesn’t reimplement. That is the pattern used throughout this article: Path objects flow through the code, but shutil.copytree, shutil.move, and shutil.make_archive still get called directly when the job calls for them.


Finding files

Files with a given extension

from pathlib import Path
def find_files(directory, extension):
    """Find files with a specific extension."""
    path = Path(directory)
    return list(path.glob(f'**/*.{extension}'))
# Usage
pdf_files = find_files('.', 'pdf')
for file in pdf_files:
    print(file)

Path.glob('**/*.ext') walks the directory tree recursively and matches the pattern lazily — the generator underneath only touches the filesystem as you iterate, so wrapping it in list() here is a deliberate choice to get an eager, reusable collection rather than a one-shot generator. One gotcha worth knowing: on case-insensitive filesystems (NTFS on Windows, APFS on macOS by default), *.PDF and *.pdf will both match; on case-sensitive filesystems (most Linux ext4 setups) they will not. If a script needs to behave identically across platforms, normalize the extension with .lower() before comparing, or match with a case-insensitive regex instead of relying on glob alone.

import os
from datetime import datetime, timedelta
def find_old_files(directory, days=30):
    """Find files older than N days."""
    cutoff = datetime.now() - timedelta(days=days)
    old_files = []
    
    for root, dirs, files in os.walk(directory):
        for file in files:
            filepath = Path(root) / file
            mtime = datetime.fromtimestamp(filepath.stat().st_mtime)
            
            if mtime < cutoff:
                old_files.append(filepath)
    
    return old_files
# Usage
old_files = find_old_files('.', days=90)
print(f"{len(old_files)} old file(s)")

This function uses os.walk() instead of Path.rglob() on purpose: os.walk() yields (root, dirs, files) tuples, and mutating dirs in place (e.g. dirs[:] = [d for d in dirs if d not in ('.git', 'node_modules')]) is the standard way to skip entire subtrees during a large recursive scan — something you cannot do mid-iteration with rglob(). Two pitfalls to keep in mind here: st_mtime reflects the content modification time, not creation time (Linux has no reliable creation timestamp in the common case, though st_birthtime exists on some filesystems); and mtimes can be misleading after a git clone or tar extraction, since many tools reset mtimes to the extraction time rather than preserving the original timestamp. If “old” needs to mean “old on the server that created it” rather than “old on this filesystem,” mtime alone is not a reliable signal.


Renaming files

Batch rename

from pathlib import Path
def rename_files(directory, old_pattern, new_pattern):
    """Batch rename files in a directory."""
    path = Path(directory)
    
    for file in path.glob('*'):
        if old_pattern in file.name:
            new_name = file.name.replace(old_pattern, new_pattern)
            file.rename(file.parent / new_name)
            print(f"{file.name} → {new_name}")
# Usage
rename_files('.', 'old_', 'new_')

Path.rename() maps directly onto the operating system’s rename() system call, which is important: on POSIX systems, a rename within the same filesystem is atomic. Other processes reading the directory will see either the old name or the new name, never a half-renamed state, because the kernel updates a single directory entry rather than copying bytes. That is the property that makes rename-based patterns so valuable for automation — a crash mid-script cannot leave a partially renamed file behind.

There are two gotchas this simple version glosses over, though. First, name collisions are silent on POSIX: if new_name already exists, Path.rename() overwrites it without warning, permanently destroying whatever was there. A safer version checks first:

target = file.parent / new_name
if target.exists():
    raise FileExistsError(f"{target} already exists, refusing to overwrite")
file.rename(target)

Second, Windows behaves differently: os.rename() (and therefore Path.rename()) raises FileExistsError if the destination exists, rather than silently replacing it. If a script needs consistent “replace if exists, atomic either way” behavior across platforms, use os.replace() (or Path.replace()) instead of rename() — it is documented to succeed even when the target exists, on both POSIX and Windows.

Adding sequence numbers

def add_numbers(directory, extension):
    """Prefix files with a zero-padded sequence number."""
    path = Path(directory)
    files = sorted(path.glob(f'*.{extension}'))
    
    for i, file in enumerate(files, 1):
        new_name = f"{i:03d}_{file.name}"
        file.rename(file.parent / new_name)
        print(f"{file.name} → {new_name}")
# Usage
add_numbers('./images', 'jpg')
# photo.jpg → 001_photo.jpg

Sorting before renaming matters more than it looks: Path.glob() does not guarantee a stable order across platforms (it generally reflects filesystem directory-entry order, which on some filesystems is closer to insertion order than alphabetical order), so skipping sorted() can produce a different numbering sequence each run. This function is also a good place to flag a subtler bug pattern: because it renames files one at a time inside the same directory it is scanning, if extension matching were loose enough that a newly-created 001_photo.jpg could itself match the glob pattern on a later run, a script re-run mid-way could double-prefix files. Here the glob is evaluated once up front into a list before the loop starts, which avoids that trap — but it is a common mistake when someone “simplifies” the code to iterate path.glob(...) directly instead of a pre-materialized list.


Organizing files

Sort into folders by extension

import shutil
from pathlib import Path
def organize_files(directory):
    """Move files into subfolders named by extension."""
    path = Path(directory)
    
    for file in path.iterdir():
        if file.is_file():
            # Extension without dot
            ext = file.suffix[1:]  # .jpg → jpg
            
            if ext:
                # Create folder
                target_dir = path / ext
                target_dir.mkdir(exist_ok=True)
                
                # Move file
                shutil.move(str(file), str(target_dir / file.name))
                print(f"{file.name} → {ext}/")
# Usage
organize_files('./downloads')

Notice this uses shutil.move() rather than Path.rename(). The distinction matters: shutil.move() first attempts a plain rename, and only if that fails with OSError: Invalid cross-device link (moving across a filesystem boundary — a different drive letter on Windows, a different mounted volume, a Docker bind mount, etc.) does it transparently fall back to copy-then-delete. Path.rename() has no such fallback and simply raises the OSError. Since a “downloads → sorted subfolder” move usually stays on the same filesystem, either call would normally work here, but using shutil.move() makes the function resilient if it is ever pointed at a directory structure that spans volumes — at the cost of losing atomicity in that fallback case, since a crash mid-copy can leave a partial file at the destination while the original still exists at the source. path.iterdir() is intentionally non-recursive (unlike rglob), which is correct here — organizing already-sorted subfolders recursively would immediately try to re-sort files it just moved.


Automated backups

A one-off shutil.copytree() is easy to write, but production backup jobs almost always need three things a naive script does not give you for free: a decision about what actually needs copying (to avoid re-copying unchanged gigabytes every run), a way to detect that a copy failed partway through, and a retention policy so backups don’t quietly fill the disk. The diagram below sketches the decision flow a more robust, incremental version of the backup script would follow — the code snippets after it show the simpler full-copy approach that is often good enough for a nightly job over a modest-sized directory.

flowchart TD
    A["Walk source tree"] --> B{"mtime + size unchanged<br/>since last backup?"}
    B -- "Yes, unchanged" --> C["Skip file"]
    B -- "No, or unsure" --> D["Compute checksum (sha256)"]
    D --> E{"Checksum matches<br/>last recorded value?"}
    E -- "Match" --> C
    E -- "Differs" --> F["Copy to temp file<br/>on destination volume"]
    F --> G["Verify temp file checksum"]
    G --> H["os.replace() to final name<br/>(atomic on same filesystem)"]
    H --> I["Record new checksum + mtime<br/>in manifest"]

Backup script

import shutil
from pathlib import Path
from datetime import datetime
def backup_directory(source, backup_root):
    """Back up a directory tree."""
    timestamp = datetime.now().strftime('%Y%m%d_%H%M%S')
    backup_name = f"backup_{timestamp}"
    backup_path = Path(backup_root) / backup_name
    
    # Copy tree
    shutil.copytree(source, backup_path)
    print(f"Backup done: {backup_path}")
    
    # Zip
    shutil.make_archive(str(backup_path), 'zip', backup_path)
    shutil.rmtree(backup_path)  # remove unzipped folder
    print(f"Archive created: {backup_path}.zip")
# Usage
backup_directory('./project', './backups')

This is a full backup: every run copies the entire source tree, regardless of what changed since the last run. That is simple and easy to reason about, but it does not scale well — copying and re-zipping 50 GB of mostly-unchanged files every night wastes disk I/O, CPU on compression, and time. The alternative, incremental backup, only copies files that changed since the last snapshot. The naive way to detect “changed” is comparing st_mtime, but mtime alone is not trustworthy as the sole signal: a touch without content changes updates mtime with no real change, while some tools (archive extraction, certain sync utilities) can reset mtimes without the content actually changing, and clock skew between machines can make timestamps lie in either direction. That is why the diagram above computes a checksum (MD5 or SHA-256) as the source of truth whenever mtime looks like it changed — mtime is used as a cheap first filter to avoid hashing every file on every run, and the hash is the final word on whether the content actually differs.

It is also worth being explicit that shutil.copytree() and make_archive() are not atomic with respect to the rest of the filesystem: if the process is killed midway through copytree, backup_path is left as a partially-copied directory, and the subsequent make_archive + rmtree sequence never runs, so a stale partial folder can accumulate under backup_root. A production version typically writes to a *.tmp name and renames (or moves) it to the final name only after the copy and verification succeed — the same atomic-rename-as-a-commit pattern used for the renaming functions earlier in this article.

For anything beyond “back up a single project directory nightly,” reach for a purpose-built tool rather than hand-rolling deduplication and incremental logic: restic, Borg, and rclone already implement content-addressed deduplication, encryption, and incremental snapshotting correctly and efficiently. The scripts in this section are meant for small, self-contained automation tasks — not as a replacement for a real backup system protecting anything you can’t afford to lose.

Pruning old backups

def cleanup_old_backups(backup_dir, keep_count=5):
    """Keep only the N most recent backups."""
    path = Path(backup_dir)
    backups = sorted(path.glob('backup_*.zip'), key=lambda x: x.stat().st_mtime)
    
    for backup in backups[:-keep_count]:
        backup.unlink()
        print(f"Deleted: {backup.name}")
# Usage
cleanup_old_backups('./backups', keep_count=5)

Sorting by st_mtime rather than parsing the timestamp out of the filename is a small but deliberate robustness choice — it keeps working even if the naming scheme changes later, as long as file modification order still reflects creation order. One edge case to handle defensively: if keep_count is larger than the number of existing backups, backups[:-keep_count] evaluates to an empty slice (not an error), so the function safely does nothing rather than deleting everything — but if keep_count is ever 0, backups[:-0] is backups[:0], which is also empty, silently keeping all backups instead of deleting all of them as a caller might expect. That kind of off-by-slice bug is worth a unit test if this function is depended on in a real retention policy.


Finding duplicates

Hash-based duplicate detection

import hashlib
from collections import defaultdict
def find_duplicates(directory):
    """Find duplicate files using MD5 hashes."""
    hashes = defaultdict(list)
    
    for file in Path(directory).rglob('*'):
        if file.is_file():
            with open(file, 'rb') as f:
                file_hash = hashlib.md5(f.read()).hexdigest()
            hashes[file_hash].append(file)
    
    duplicates = {h: files for h, files in hashes.items() if len(files) > 1}
    
    for hash_val, files in duplicates.items():
        print(f"\nDuplicate group ({hash_val[:8]}...):")
        for file in files:
            print(f"  - {file}")
    
    return duplicates
# Usage
duplicates = find_duplicates('./documents')

MD5 being cryptographically broken (collisions can be engineered deliberately) is irrelevant here, because this is detecting accidental duplicates among files you already trust, not defending against an adversary crafting a malicious collision — MD5’s speed is the relevant property, not its security. What does matter in production is I/O cost: f.read() loads the entire file into memory before hashing, which is fine for documents but will blow up memory usage (or at least be needlessly slow) on a directory containing large video or archive files. Two standard optimizations fix that:

  1. Pre-filter by file size before hashing anything — two files of different sizes obviously cannot be identical, so grouping by file.stat().st_size first and only hashing within groups that share a size avoids reading most files at all.
  2. Hash in chunks (hashlib.md5() updated incrementally via .update() in a loop reading fixed-size blocks) instead of f.read(), so memory use stays constant regardless of file size.

Path.rglob('*') here also means every file and directory is visited (the file.is_file() check filters out directories), and it will silently descend into symlinked subdirectories inconsistently depending on platform and Python version — worth keeping in mind if a documents folder might contain symlinks back into itself, which can turn a duplicate scan into an effectively infinite walk.


Real-world example

Log cleanup script

from pathlib import Path
import gzip
from datetime import datetime, timedelta
def cleanup_logs(log_dir, archive_days=7, delete_days=30):
    """
    Log maintenance:
    - Older than archive_days: gzip
    - Older than delete_days: delete
    """
    path = Path(log_dir)
    now = datetime.now()
    
    for log_file in path.glob('*.log'):
        mtime = datetime.fromtimestamp(log_file.stat().st_mtime)
        age = (now - mtime).days
        
        if age >= delete_days:
            log_file.unlink()
            print(f"Deleted: {log_file.name} ({age} days)")
        
        elif age >= archive_days:
            gz_path = log_file.with_suffix('.log.gz')
            
            with open(log_file, 'rb') as f_in:
                with gzip.open(gz_path, 'wb') as f_out:
                    f_out.writelines(f_in)
            
            log_file.unlink()
            print(f"Compressed: {log_file.name} → {gz_path.name}")
# Usage
cleanup_logs('./logs', archive_days=7, delete_days=30)

This script combines several of the ideas above into one maintenance job, and it is also a good case study in crash-safety ordering. Look closely at the compress branch: the gzip write happens before log_file.unlink(), which is the correct order — if the process is killed while gzip.open(...).writelines(...) is still running, the original .log file is untouched and the partially-written .gz file is simply overwritten (since gzip.open(..., 'wb') truncates) the next time the job runs. Reversing that order — deleting the source first, then compressing — would risk permanent data loss if the process died in between. The general pattern is: never delete or move the only copy of data until the new copy is confirmed written, and ideally verified (e.g. by attempting to read back the last few bytes, or checking the gzip trailer’s CRC).

Two production pitfalls apply directly to a script like this running unattended on a schedule. First, permission errors: a log file currently held open for writing by the application that produced it can be safe to read on Linux (the read succeeds even while another process writes to the same inode) but may raise PermissionError on Windows, where an exclusively-locked file blocks even read access from another process. A scheduled job should catch PermissionError per-file and log-and-continue rather than letting one locked file abort the entire batch. Second, partial-write visibility: if any external process (a log shipper, a monitoring agent) watches the log directory for new .gz files, it could pick up the gzip file mid-write, before writelines finishes flushing — the safe fix is writing to a temporary name (log_file.name + '.gz.tmp') and calling Path.replace() to the final .log.gz name only after the write completes, which is once again the same atomic-rename-as-commit idea from the renaming and backup sections above.


Making file automation safe

File automation checklist

# ✅ Safer file operations
# 1. Back up first
# 2. Dry-run mode (preview before destructive steps)
# 3. Logging
# ✅ Error handling
try:
    shutil.move(src, dst)
except PermissionError:
    print("Permission denied")
except FileNotFoundError:
    print("File not found")
# ✅ Progress feedback
from tqdm import tqdm
for file in tqdm(files, desc="Processing"):
    process(file)

A few of these points deserve elaboration beyond the checklist comments. Dry-run mode is worth implementing as a first-class parameter (dry_run: bool = False) rather than a comment reminder — have every destructive call (unlink, rename, rmtree, move) gated behind if not dry_run:, with a print/log statement outside the gate so the exact same code path reports what it would do. This turns every script in this article into something safe to run against production directories for a first pass. Logging should go to a file with timestamps (Python’s logging module, not bare print), because automation scripts are usually run unattended via cron or Task Scheduler — by the time someone notices a problem, the terminal output that would have explained it is long gone. And on the exception handling: catching PermissionError and FileNotFoundError separately (rather than a bare except Exception) matters because it lets you decide different responses — a missing file might mean “someone already cleaned it up, safe to skip,” while a permission error might mean “this needs investigation,” and collapsing both into one generic handler hides that distinction from whoever reads the logs later.