Python File Handling | Read, Write, CSV, JSON, and pathlib

Key takeaways

Learn Python file I/O: text files, CSV with csv module, JSON with json, pathlib paths, encoding, and the with statement—practical patterns included.

Introduction

Almost every non-trivial Python program eventually has to talk to the filesystem — reading a configuration file at startup, writing structured logs, exporting a report as CSV, or persisting application state as JSON between runs. Unlike, say, a database connection, file I/O looks deceptively simple: open(), read or write, close(). In practice it is one of the more error-prone corners of a codebase, because the failure modes are quiet. A file left open after an exception doesn’t crash immediately — it leaks a file descriptor that eventually exhausts a limit under load. A file opened without an explicit encoding works fine on your Linux laptop and produces garbled text the moment a teammate runs the same script on Windows. A CSV writer without newline='' silently inserts blank rows on Windows but not on macOS.

This post walks through the four building blocks you’ll use constantly: plain text files, CSV (tabular data), JSON (nested data), and pathlib (path manipulation). For each one, the goal isn’t just to show the API — it’s to explain why the API is shaped the way it is, and where the common mistakes come from, so the patterns actually stick.


Reading and writing text files

Python’s open() function is the entry point for all file I/O, text and binary alike. The two arguments that matter most for text files are the mode ('r', 'w', 'a', plus optional '+', 'b') and the encoding. Getting either wrong is the single most common source of file-handling bugs in real projects.

Reading

# Read entire file
with open('data.txt', 'r', encoding='utf-8') as f:
    content = f.read()
    print(content)
# Read all lines into a list
with open('data.txt', 'r', encoding='utf-8') as f:
    lines = f.readlines()
    for line in lines:
        print(line.strip())
# Iterate line by line (memory-friendly)
with open('data.txt', 'r', encoding='utf-8') as f:
    for line in f:
        print(line.strip())

These three patterns look interchangeable but behave very differently under load. f.read() loads the entire file into a single string in memory — fine for a 2 KB config file, dangerous for a 4 GB log file, where it can spike memory usage and stall the process. f.readlines() is just as memory-hungry, because it still materializes every line as a list before you touch a single element. The third form — iterating over f directly — is the one you should reach for by default when the file’s size isn’t bounded, because the file object is itself an iterator: it pulls one line at a time from an internal buffer and never holds more than a chunk of the file in memory at once. The trade-off is that you lose random access; you can’t jump back to line 3 after you’ve already consumed it without reopening the file or calling f.seek(0).

Notice the explicit encoding='utf-8' on every call. If you omit it, Python falls back to locale.getpreferredencoding(), which is utf-8 on most Linux and macOS setups but is frequently cp1252 or cp949 on Windows depending on the system locale. That means the exact same script, run unmodified on two different machines, can raise a UnicodeDecodeError on one and run silently on the other — or worse, run silently on both but corrupt non-ASCII characters (accented letters, Korean text, emoji) without ever raising an error. Always pass encoding='utf-8' explicitly unless you have a specific reason to target a legacy encoding; treat the encoding argument as non-optional in production code, not as a stylistic nicety.

The with statement here is doing more than tidiness — it’s a correctness guarantee. open() returns a context manager whose __exit__ method calls f.close() unconditionally, including when an exception propagates out of the block. Without it, an exception raised between open() and a manual f.close() skips the close entirely, and the underlying OS file descriptor stays allocated until the garbage collector eventually reclaims the object — which, in CPython, usually happens quickly via reference counting, but is not something you should rely on, especially in long-running processes or under PyPy, where garbage collection timing is different. The sequence looks like this:

sequenceDiagram
    participant Code as Your code
    participant CM as with block
    participant OS as OS file handle

    Code->>CM: open("data.txt")
    CM->>OS: acquire file descriptor
    Code->>CM: read / process lines
    alt exception raised mid-loop
        CM->>OS: __exit__ still runs\ncloses descriptor
        CM-->>Code: exception re-raised after cleanup
    else normal completion
        CM->>OS: __exit__ runs\ncloses descriptor
    end

The key point the diagram makes is that __exit__ runs on both paths — success and failure — before the exception (if any) continues propagating up the call stack. That’s the entire reason with exists instead of a bare open()/close() pair.

Writing

# Overwrite (w)
with open('output.txt', 'w', encoding='utf-8') as f:
    f.write("First line\n")
    f.write("Second line\n")
# Append (a)
with open('output.txt', 'a', encoding='utf-8') as f:
    f.write("Third line\n")
# Write multiple lines
lines = ["line 1\n", "line 2\n", "line 3\n"]
with open('output.txt', 'w', encoding='utf-8') as f:
    f.writelines(lines)

The distinction between 'w' and 'a' is a common source of production incidents that have nothing to do with syntax and everything to do with intent. 'w' truncates the file to zero length the instant it’s opened — even if you never call .write() afterward. That means opening a file with 'w' inside a retry loop, or inside a function that’s occasionally called with the wrong path, can silently destroy existing data before your code has a chance to fail gracefully. 'a' never truncates; it always appends at the end, which is what you want for logs, audit trails, or anything you’re accumulating over time rather than replacing.

A subtlety worth internalizing: f.write() does not add a newline for you, unlike print(). If you forget the \n, consecutive write() calls run together on the same line. f.writelines() has the same behavior — it does not insert separators between the strings you pass it, so the list itself must already contain the line terminators, as shown above. This is a frequent source of confusion for people coming from languages where a “writeln”-style function exists by default.


CSV with the csv module

CSV looks like the simplest possible format — commas and newlines — which is exactly why hand-rolling a parser with str.split(',') is a trap. Real-world CSV has quoted fields containing commas, embedded newlines inside quoted fields, escaped quote characters, and platform-specific line endings. The csv module in the standard library handles all of that correctly according to RFC 4180-style conventions, which is why you should use it even for “trivial” CSV work.

Reading CSV

import csv
# As rows (lists)
with open('data.csv', 'r', encoding='utf-8') as f:
    reader = csv.reader(f)
    header = next(reader)  # first row (header)

    for row in reader:
        print(row)
# As dict rows (recommended)
with open('data.csv', 'r', encoding='utf-8') as f:
    reader = csv.DictReader(f)

    for row in reader:
        print(row['name'], row['age'])

csv.reader gives you each row as a plain list of strings, positional and terse — you access row[0], row[1], and so on. That’s fine for a one-off script, but it couples your code to column order: if someone inserts a new column in the source file, every downstream index shifts silently and you get wrong values instead of an error. csv.DictReader reads the first row as a header automatically and hands you each subsequent row as an ordered dict keyed by column name, so row['name'] keeps working even if the column order in the file changes. The cost is a small amount of overhead per row (building a dict instead of a list), which is irrelevant for anything short of very high-throughput parsing. Default to DictReader unless you have a specific performance reason not to — it’s more defensive and more readable.

Note also that every value coming out of csv.reader or csv.DictReader is a string, including things that look like numbers. row['age'] is '25', not 25. You need to convert explicitly (int(row['age'])) and handle the ValueError that comes from malformed or missing numeric fields — CSV has no type system, so nothing enforces that a column that’s usually numeric doesn’t occasionally contain an empty string or a stray label.

Writing CSV

import csv
# Write rows as lists
data = [
    ['Name', 'Age', 'City'],
    ['Alice', 25, 'Seoul'],
    ['Bob', 30, 'Busan']
]
with open('output.csv', 'w', newline='', encoding='utf-8') as f:
    writer = csv.writer(f)
    writer.writerows(data)
# Write dict rows
data = [
    {'name': 'Alice', 'age': 25, 'city': 'Seoul'},
    {'name': 'Bob', 'age': 30, 'city': 'Busan'}
]
with open('output.csv', 'w', newline='', encoding='utf-8') as f:
    fieldnames = ['name', 'age', 'city']
    writer = csv.DictWriter(f, fieldnames=fieldnames)

    writer.writeheader()
    writer.writerows(data)

The single most important detail in this whole section is newline='' on the open() call — easy to miss, and the Python documentation is explicit that it’s required whenever you’re writing CSV with the csv module. Here’s why it matters: the csv writer manages its own line-ending logic internally (\r\n per the CSV convention), but Python’s text-mode file object also does newline translation by default, converting every \n it sees into the platform’s native line ending on write. On Windows, that means \n becomes \r\n — and if the csv writer had already written \r\n, the file object’s translation layer doubles it into \r\r\n, which many CSV parsers (including Excel) interpret as an extra blank row between every record. Passing newline='' disables the file object’s own newline translation and lets the csv module own line endings exclusively, which is the only combination that produces correct output on every platform.

DictWriter mirrors DictReader: you pass fieldnames explicitly (there’s no header to infer it from when writing), call writeheader() once, then writerows() with a list of dicts whose keys match fieldnames. If a dict is missing a key that’s in fieldnames, DictWriter raises a ValueError by default rather than silently leaving the cell blank — which is usually what you want, since a missing key is more often a bug than a legitimate absence.


JSON with the json module

JSON maps naturally onto Python’s dict, list, str, int, float, bool, and None, which is why json.load/json.dump feel almost frictionless compared to CSV. The friction shows up at the edges — types JSON doesn’t have a native representation for.

Reading JSON

import json
# From file
with open('data.json', 'r', encoding='utf-8') as f:
    data = json.load(f)
    print(data)
# From string
json_str = '{"name": "Alice", "age": 25}'
data = json.loads(json_str)
print(data['name'])  # Alice

json.load() takes a file object and reads directly from it; json.loads() (note the trailing s, for “string”) parses a string you already have in memory — for example, a JSON payload you received over HTTP rather than from disk. Mixing these up is a very common TypeError for beginners: passing a string to json.load() fails because it expects something with a .read() method, and passing a file object to json.loads() fails for the mirror-image reason. If the input isn’t valid JSON — a trailing comma, single quotes instead of double, an unescaped control character — both functions raise json.JSONDecodeError, which carries a .pos, .lineno, and .colno attribute pointing at exactly where parsing failed. Catching and logging those attributes is far more useful for debugging a malformed config file than catching a bare Exception.

Writing JSON

import json
data = {
    "name": "Alice",
    "age": 25,
    "hobbies": ["reading", "movies", "sports"],
    "address": {
        "city": "Seoul",
        "district": "Gangnam"
    }
}
# Write to file
with open('output.json', 'w', encoding='utf-8') as f:
    json.dump(data, f, ensure_ascii=False, indent=2)
# Serialize to string
json_str = json.dumps(data, ensure_ascii=False, indent=2)
print(json_str)

Two keyword arguments here deserve real explanation rather than being treated as boilerplate. ensure_ascii defaults to True in the standard library, which means by default json.dump/json.dumps escape every non-ASCII character into a \uXXXX sequence — so "서울" becomes "서울" in the output file. That’s technically valid JSON and any compliant parser reads it back correctly, but it makes the file unreadable to a human and roughly doubles the byte size of any non-Latin text. Passing ensure_ascii=False writes the actual UTF-8 characters instead, which is almost always what you want for human-inspectable output or when interoperating with systems that expect readable Unicode. indent=2 pretty-prints the output with two-space indentation instead of the default single-line compact form — useful for config files and debugging, but worth skipping (leave indent=None) for machine-to-machine payloads where the extra whitespace is pure overhead.

The type mismatch that trips people up most often is datetime: Python’s datetime.datetime objects are not JSON-serializable out of the box, and json.dump raises TypeError: Object of type datetime is not JSON serializable the moment it encounters one. The two common fixes are converting to an ISO string before serializing (dt.isoformat()) or supplying a custom default= function to json.dump that knows how to convert your non-standard types. The same applies to set (JSON has no set type — convert to list) and Decimal (convert to float or str depending on whether you can tolerate precision loss).


Paths with pathlib

pathlib.Path is the modern, object-oriented replacement for the older os.path string-manipulation functions (os.path.join, os.path.exists, os.path.splitext, and so on). The core ergonomic win is that path segments compose with the / operator instead of nested function calls, and the result is a first-class object with methods, not just a string you have to remember to pass back into os.path.* functions every time.

Using pathlib

from pathlib import Path
# Current directory
current = Path.cwd()
print(current)
# Join paths
data_dir = Path('data')
file_path = data_dir / 'users.json'
print(file_path)  # data/users.json
# Exists?
if file_path.exists():
    print("file exists")
# Create directory
data_dir.mkdir(exist_ok=True)
# Read/write text
file_path.write_text("Hello, World!", encoding='utf-8')
content = file_path.read_text(encoding='utf-8')
print(content)
# Path parts
print(file_path.name)    # users.json
print(file_path.stem)    # users
print(file_path.suffix)  # .json
print(file_path.parent)  # data

data_dir / 'users.json' works because Path overloads __truediv__, which is a deliberate, readable use of operator overloading rather than a gimmick — it mirrors how a path actually looks on disk and, critically, produces the correct separator for the current platform automatically (/ on Linux/macOS, \ on Windows) without you writing any conditional logic. That cross-platform correctness is the main practical reason to prefer pathlib over building paths with string concatenation or os.path.join in new code: a hardcoded 'data' + '/' + 'users.json' will quietly work in dev on Linux and then break on a Windows teammate’s machine or in certain Windows-hosted CI runners.

mkdir(exist_ok=True) is worth calling out specifically: without exist_ok=True, calling .mkdir() on a directory that already exists raises FileExistsError, which is often not what you want in idempotent setup code that might run more than once against the same environment. exist_ok=True makes the call a no-op if the directory is already there instead of treating it as an error condition.

One gotcha worth knowing: file_path.exists() followed later by file_path.write_text(...) is not atomic — between the check and the write, another process (or another part of your own program, in a threaded or async context) could create, delete, or modify that file. This “time-of-check to time-of-use” (TOCTOU) gap is rarely a problem in a single-threaded script but matters a lot in concurrent code or anything touching shared/network filesystems; in those cases, prefer catching the specific exception (FileNotFoundError, FileExistsError) over pre-checking with .exists().

write_text()/read_text() are convenience wrappers that open the file, read or write the whole content, and close it in one call — ideal for small files like config or a single JSON blob, but you lose the line-by-line iteration and streaming behavior that open() gives you, so they’re the wrong tool for large files.


A log analyzer and a config helper

The two examples below aren’t toy demonstrations — they’re patterns you’ll reuse directly in real projects, so it’s worth understanding the design decisions behind each one, not just the syntax.

Log analysis

from collections import Counter
from pathlib import Path
def analyze_log(log_file):
    """Count error types from a log file."""
    error_counts = Counter()

    with open(log_file, 'r', encoding='utf-8') as f:
        for line in f:
            if 'ERROR' in line:
                error_type = line.split(':')[1].strip()
                error_counts[error_type] += 1

    return error_counts
# Usage
errors = analyze_log('app.log')
for error, count in errors.most_common(5):
    print(f"{error}: {count} occurrences")

This function deliberately iterates the file with for line in f rather than f.readlines(), for the memory reason discussed in the text-files section — log files are exactly the case where size is unbounded and often large, so streaming line-by-line keeps memory usage flat regardless of whether the log is 10 KB or 10 GB. Counter (from the collections module) is the right data structure here instead of a plain dict with manual if key not in counts checks, because Counter[key] += 1 on a missing key just initializes it to 0 first — no KeyError, no boilerplate. most_common(5) then does the sorting-and-slicing work for you in one call.

The fragile part of this function, worth flagging explicitly rather than hiding, is line.split(':')[1].strip() — it assumes every ERROR line has a colon-delimited format with the error type as the second field. That’s a reasonable assumption for a specific log format you control, but it will raise IndexError the moment a line matches 'ERROR' but doesn’t have that exact shape (for example, a stack trace continuation line that happens to contain the word “ERROR”). In production code you’d want to either use a proper log-parsing library, a regular expression with named groups, or at minimum wrap the split in a try/except IndexError so one malformed line doesn’t crash the whole analysis run. That’s exactly the kind of gap the next post in this series — exception handling — is designed to close.

Simple config helper

import json
from pathlib import Path
class Config:
    def __init__(self, config_file='config.json'):
        self.config_file = Path(config_file)
        self.data = self.load()

    def load(self):
        if self.config_file.exists():
            with open(self.config_file, 'r', encoding='utf-8') as f:
                return json.load(f)
        return {}

    def save(self):
        with open(self.config_file, 'w', encoding='utf-8') as f:
            json.dump(self.data, f, ensure_ascii=False, indent=2)

    def get(self, key, default=None):
        return self.data.get(key, default)

    def set(self, key, value):
        self.data[key] = value
        self.save()
# Usage
config = Config()
config.set('database_url', 'localhost:5432')
print(config.get('database_url'))

This small class ties together everything covered above: pathlib.Path for the file location, json.load/json.dump for serialization, and with for safe I/O. The design choice worth noting is load() returning {} instead of raising when the config file doesn’t exist yet — that makes Config() safe to instantiate on a fresh install where no config has been written yet, rather than forcing every caller to handle a FileNotFoundError on first run. The trade-off is that set() calls save() on every single call, which means every config.set(...) does a full file rewrite. That’s perfectly fine for a handful of settings changed occasionally, but it would be a poor choice for a hot path that updates config dozens of times per second — at that point you’d want to batch writes or debounce the save. Sizing the pattern to the actual write frequency, rather than reaching for this exact shape everywhere, is the real skill here.


Encoding mismatches and other file-handling failures

A few failure modes come up repeatedly enough in real Python codebases that they’re worth calling out directly rather than leaving them to be discovered the hard way:

  • Encoding mismatches across platforms. As covered above, omitting encoding='utf-8' lets Python fall back to the OS locale encoding. The bug this causes is particularly nasty because it often doesn’t raise an exception at all — it silently mojibake-corrupts non-ASCII text, and the corruption isn’t noticed until someone reads the output file. Always pin the encoding explicitly.
  • 'w' mode truncating data unexpectedly. Because 'w' truncates on open regardless of whether you ever call .write(), a bug in argument handling that routes execution into a 'w'-mode open() call on the wrong path can destroy a file before any exception has a chance to stop the program. If a function’s job is “append if the file exists,” reach for 'a', not a conditional 'w'.
  • Missing newline='' when writing CSV. Covered in detail above — this is the classic cause of mysterious blank rows appearing in CSV output that only shows up on Windows, making it painful to reproduce if your CI runs on Linux.
  • json.dump failing on datetime, set, or Decimal. These are common in real application data models but have no native JSON representation. Decide on a conversion strategy (ISO strings for dates, lists for sets) before you hit the TypeError in production.
  • Checking Path.exists() and then acting on it separately. As noted in the pathlib section, this has a TOCTOU race in concurrent contexts. Prefer catching the specific exception the operation itself raises.
  • Reading an entire large file with .read() or .readlines(). Fine for small files, but it’s easy to write code against a small test file and only discover the memory problem once it runs against a multi-gigabyte file in production. Default to line-by-line iteration unless you specifically need random access or the whole content in memory at once.

None of these are exotic — they’re the ordinary, boring bugs that show up in code review or, worse, in production logs. Knowing the mechanism behind each one (why 'w' truncates immediately, why newline translation doubles line endings, why encoding fallback is locale-dependent) makes them easy to spot in a diff instead of something you debug after the fact.


Next in the series

Exception handling is the natural next step, since several of the failures above (malformed CSV rows, JSONDecodeError, missing files) are really exception-handling problems in disguise. Later, file automation puts pathlib and these file APIs to work on whole directories.