Python Functions: Parameters, Defaults, Closures, and the Traps Behind Them

Key takeaways

Python functions are objects created when `def` runs. That one fact explains the mutable default trap, late-binding closures, and why decorators need functools.wraps.

Why functions deserve more than a syntax overview

Most Python tutorials treat def as a way to name a block of code. That is true, but the detail that matters is this: def is an executable statement that creates a function object at runtime. The function object carries its code, its default values, and references to the variables it closes over. Almost every “why does Python do that?” moment with functions — shared default lists, loops of lambdas that all return the same number, decorators that break help() — follows directly from that one fact.

This article walks through the parameter system, return values, closures, and a first look at decorators, with the traps explained rather than just listed. Every output shown was produced by running the snippet on CPython 3.11.


Defining, calling, and returning

def greet(name: str) -> str:
    """Return a greeting for the given name."""
    return f"Hello, {name}!"

print(greet("Bob"))  # Hello, Bob!

The docstring is not a comment: it is stored on greet.__doc__ and is what help(greet) and IDEs show. The type hints are also stored (greet.__annotations__) but not enforced at runtime — greet(42) will happily run. Hints pay off through tools like mypy or pyright, not through the interpreter.

Returning several values

def get_user_info():
    return "Bob", 25, "Seoul"      # this is one tuple

name, age, city = get_user_info()  # tuple unpacking
name, _, city = get_user_info()    # _ signals "ignored"

Python never returns “multiple values”; it returns one tuple and the caller unpacks it. Once a function returns more than three or four items, positional unpacking becomes fragile — adding a field silently shifts every caller. At that point a dataclass or NamedTuple is safer, because callers access fields by name.

Implicit None

A function with no return (or a bare return) returns None. This bites people with in-place methods:

nums = [3, 1, 2]
nums = nums.sort()   # sort() works in place and returns None
print(nums)          # None

The convention in the standard library is deliberate: methods that mutate return None so you cannot mistake them for methods that return a new object (sorted(nums) returns a new list).

When None is a legitimate “not found” result, compare with is None, not truthiness. if not user: also treats 0, "", and [] as missing, which is wrong for a user ID of 0.

def find_user(user_id):
    users = {1: "Bob", 2: "Alice"}
    return users.get(user_id)   # None when missing

if find_user(5) is None:
    print("Not found")

The parameter system

Positional and keyword arguments

def introduce(name, age, city):
    print(f"{name} ({age}) - {city}")

introduce("Bob", 25, "Seoul")
introduce(age=25, name="Bob", city="Seoul")
introduce("Bob", age=28, city="Daejeon")   # positional first, then keywords

Keyword arguments make call sites self-documenting, especially for booleans: send(msg, True, False) tells a reader nothing, send(msg, retry=True, log=False) does.

Positional-only (/) and keyword-only (*)

Since Python 3.8 you can declare exactly how each parameter may be passed:

def f(a, b, /, c, *, d):
    return a, b, c, d

f(1, 2, 3, d=4)       # OK
f(1, 2, c=3, d=4)     # OK
f(1, b=2, c=3, d=4)
# TypeError: f() got some positional-only arguments passed as keyword arguments: 'b'
f(1, 2, 3, 4)
# TypeError: f() takes 3 positional arguments but 4 were given

Why bother? Keyword-only parameters are ideal for options and flags: callers must name them, so you can later reorder or add options without breaking anyone. Positional-only parameters make the name private: you can rename a to something better later, and no caller can depend on the old name. Many built-ins already work this way (len(obj=...) is an error).

*args and **kwargs

def complex_func(a, b, *args, **kwargs):
    print(a, b, args, kwargs)

complex_func(1, 2, 3, 4, x=5, y=6)
# 1 2 (3, 4) {'x': 5, 'y': 6}

*args collects extra positional arguments into a tuple; **kwargs collects extra keyword arguments into a dict. The full order in a signature is:

positional-only, /, regular, *args (or bare *), keyword-only, **kwargs

The trade-off: **kwargs makes a function accept anything, including typos. connect(host="x", timout=5) does not fail — the misspelled option is silently ignored unless you check for unknown keys. Use **kwargs when you genuinely forward arguments (wrappers, decorators, subclass __init__), and prefer explicit keyword-only parameters when you know the options.


The mutable default argument trap

This is the most famous Python function bug, and it follows directly from “def runs once”:

def add_item_bad(item, items=[]):
    items.append(item)
    return items

print(add_item_bad("a"))           # ['a']
print(add_item_bad("b"))           # ['a', 'b']   <- not ['b']
print(add_item_bad.__defaults__)   # (['a', 'b'],)

The [] is evaluated a single time, when the def statement executes, and stored in __defaults__. Every call that omits items receives the same list object. The fix is the None sentinel:

def add_item_safe(item, items=None):
    if items is None:
        items = []
    items.append(item)
    return items

The same rule applies to anything evaluated in a default, not only mutable containers. def log(msg, ts=time.time()) records the time the module was imported, not the time of the call — I verified that two calls a few milliseconds apart return the identical timestamp.

I have been burned by this in a less obvious form: a helper that built HTTP headers used headers={} as a default and added an auth token to it. In a long-running worker process, the token from the first request stayed in the shared dict and was sent on later requests that were supposed to be anonymous. Nothing crashed, which is exactly why it took a while to notice. Linters catch this pattern (Ruff and flake8-bugbear flag it as B006), and I now treat that warning as an error rather than a style nit.


Lambdas

students = [("Bob", 85), ("Alice", 90), ("Carol", 80)]
print(sorted(students, key=lambda s: s[1], reverse=True))
# [('Alice', 90), ('Bob', 85), ('Carol', 80)]

A lambda is just an anonymous function limited to a single expression. It shines as a key= argument. Beyond that, its limits are real: it has no name in tracebacks (it shows up as <lambda>), cannot contain statements, and cannot hold a docstring. PEP 8 explicitly discourages square = lambda x: x ** 2; if you are naming it, write a def.

For transforming and filtering, a comprehension is usually clearer than map/filter with a lambda:

numbers = [1, 2, 3, 4, 5]
squares = [x ** 2 for x in numbers]
evens = [x for x in numbers if x % 2 == 0]

For key functions that just pick an item or attribute, operator.itemgetter(1) and operator.attrgetter("age") are slightly faster and read well.


Scope, closures, and nonlocal

Python resolves names with the LEGB rule: Local, Enclosing function, Global, Built-in. A crucial detail is that a variable is local if the function assigns to it anywhere in its body, decided at compile time:

x = 10
def show():
    print(x)   # looks like it should read the global...
    x = 5      # ...but this assignment makes x local for the whole function

show()
# UnboundLocalError: cannot access local variable 'x' where it is not associated with a value

Closures

An inner function that references variables from its enclosing function keeps them alive after the outer function returns:

def make_multiplier(n):
    def multiply(x):
        return x * n
    return multiply

times_3 = make_multiplier(3)
print(times_3(10))   # 30

To assign to an enclosing variable you need nonlocal; without it, the same compile-time rule as above kicks in:

def make_counter():
    count = 0
    def increment():
        nonlocal count   # remove this line -> UnboundLocalError on the first call
        count += 1
        return count
    return increment

counter = make_counter()
counter(); counter()
print(counter())   # 3

The late-binding closure trap

Closures capture variables, not values. The variable is looked up when the inner function is called:

callbacks = [lambda: i for i in range(3)]
print([f() for f in callbacks])   # [2, 2, 2]

By the time any lambda runs, the loop has finished and i is 2. The idiomatic fix is to bind the current value as a default argument, which — ironically, given the mutable default trap above — works precisely because defaults are evaluated at definition time:

callbacks = [lambda i=i: i for i in range(3)]
print([f() for f in callbacks])   # [0, 1, 2]

functools.partial(func, i) is the cleaner alternative when the callback is a real function. I usually hit this when registering handlers in a loop, such as one button callback or one scheduled job per item: everything works in a quick test with a single item, and with many items every handler quietly acts on the last one.


Decorators: a first look

A decorator is a function that takes a function and returns a replacement. @my_decorator above a def is shorthand for add = my_decorator(add).

import functools
import time

def measure_time(func):
    @functools.wraps(func)
    def wrapper(*args, **kwargs):
        start = time.perf_counter()
        try:
            return func(*args, **kwargs)
        finally:
            print(f"{func.__name__}: {time.perf_counter() - start:.4f}s")
    return wrapper

@measure_time
def add(a, b):
    """Add two numbers."""
    return a + b

Three details here are not optional polish:

  • functools.wraps copies __name__, __doc__, and other metadata onto the wrapper. Without it, add.__name__ is 'wrapper' and add.__doc__ is None — I checked — which breaks help(), logging that prints function names, and frameworks that register routes or tasks by name.
  • time.perf_counter() instead of time.time(): time.time() is wall-clock time and can jump when the system clock is adjusted; perf_counter is monotonic and high resolution, which is what you want for measuring durations.
  • try/finally so the timing still prints when the wrapped function raises.

The standard library ships a very useful decorator for pure functions:

from functools import lru_cache

@lru_cache(maxsize=None)
def fib(n):
    return n if n < 2 else fib(n - 1) + fib(n - 2)

print(fib(100))          # 354224848179261915075
print(fib.cache_info())  # CacheInfo(hits=98, misses=101, maxsize=None, currsize=101)

Caching turns exponential recursion into linear work, but it has conditions: all arguments must be hashable (fib([1, 2]) raises TypeError: unhashable type: 'list'), the function must be pure (same inputs, same output), and an unbounded cache holds every result for the life of the process. Decorators with arguments, class-based decorators, and stacking order are covered in Python Decorators.


Recursion and higher-order functions

def factorial(n):
    return 1 if n <= 1 else n * factorial(n - 1)

Recursion reads well for trees and nested data, but CPython has no tail-call optimization and a default recursion limit of 1000 (sys.getrecursionlimit()). factorial(5000) raises RecursionError: maximum recursion depth exceeded. Raising the limit with sys.setrecursionlimit only moves the problem and can crash the interpreter with a real C stack overflow; for deep or linear recursion, rewrite as a loop or use an explicit stack.

Because functions are objects, passing them around is ordinary:

def apply_operation(func, value):
    return func(value)

print(apply_operation(str.upper, "hi"))   # HI

This is the foundation of sorted(key=...), callbacks, and dependency injection in tests: pass a fake function instead of the real network call.