Python Comprehensions and Generator Expressions: Scope, Late Binding, Laziness and Memory
Key takeaways
A comprehension is a small function that Python builds and calls for you. That explains most of its behavior: the loop variable does not leak, class attributes are invisible inside it, lambdas created in it all see the last value, and a generator expression runs only when consumed, exactly once. This article covers list, dict and set comprehensions and generator expressions with those rules explained, plus walrus filtering, real memory measurements, and where a plain loop is the better choice.
Introduction
A comprehension builds a list, dict or set from an iterable in a single expression: [expr for x in iterable if cond]. The syntax is easy to learn from examples. What examples usually leave out is the rule that explains the surprising behavior: a comprehension runs in its own scope, like a small function that Python defines and calls immediately. A generator expression is the same thing, except the function is a generator that nobody has started yet.
With that rule in mind, the questions people actually get stuck on have straightforward answers: why the loop variable does not leak, why a class attribute is “not defined” inside a comprehension in the class body, why a list of lambdas all return the same value, and why a generator is empty the second time. This article covers the syntax first, then those rules, then memory and readability limits. All output shown was produced on CPython 3.11.
Syntax and Clause Order
squares = [i ** 2 for i in range(10)]
evens = [i for i in range(10) if i % 2 == 0]
lengths = {w: len(w) for w in ["apple", "fig"]} # dict
domains = {e.split("@")[1] for e in ["[email protected]", "[email protected]", "[email protected]"]} # set
total = sum(i ** 2 for i in range(10)) # generator expression
if at the end filters; if/else at the front transforms
These are two different constructs that happen to share keywords:
print([i for i in range(6) if i % 2 == 0]) # [0, 2, 4] filter
print(['even' if i % 2 == 0 else 'odd' for i in range(4)]) # ['even', 'odd', 'even', 'odd']
print([i ** 2 if i > 0 else 0 for i in range(-2, 3) if i != 0]) # [0, 0, 1, 4]
The trailing if decides whether an element appears at all, so it cannot have an else. The leading x if cond else y is an ordinary conditional expression that computes the element’s value, so it must have an else. Writing [x for x in data if x > 0 else 0] is a SyntaxError, and the fix depends on which you meant: drop negatives, or replace them with zero.
Multiple for clauses read like nested loops, left to right
print([(x, y) for x in range(2) for y in range(3)])
# [(0, 0), (0, 1), (0, 2), (1, 0), (1, 1), (1, 2)]
matrix = [[1, 2, 3], [4, 5, 6], [7, 8, 9]]
print([num for row in matrix for num in row]) # flatten: [1, 2, 3, 4, 5, 6, 7, 8, 9]
print([[row[i] for row in matrix] for i in range(3)]) # transpose: [[1, 4, 7], [2, 5, 8], [3, 6, 9]]
The flattening form trips people up because the element expression comes first, but the clauses are in the order you would write nested for statements: for row in matrix: then for num in row:. A comprehension inside another comprehension (the transpose) is different: the inner one builds a whole row per outer iteration, so the result stays nested.
Dict comprehensions: duplicate keys keep the last value
items = [('a', 1), ('b', 2), ('a', 3)]
print({k: v for k, v in items}) # {'a': 3, 'b': 2}
No error, no warning: the second 'a' silently overwrites the first. That is the same as assigning in a loop, and it is exactly what happens when you invert a mapping with non-unique values ({v: k for k, v in d.items()}). If duplicates are possible and you need all of them, grouping is a loop, not a comprehension:
from collections import defaultdict
groups = defaultdict(list)
for k, v in items:
groups[k].append(v)
print(dict(groups)) # {'a': [1, 3], 'b': [2]}
A related trap in string parsing: {p.split('=')[0]: p.split('=')[1] for p in s.split(',')} splits every pair twice and drops everything after a second =. For "URL=http://x/?a=b" it produces 'http://x/?a'. dict(p.split('=', 1) for p in s.split(',')) splits once, keeps the rest of the value intact, and is shorter.
Scope: Why a Comprehension Behaves Like a Function
The loop variable does not leak (Python 3)
x = 'outer'
squares = [x * x for x in range(3)]
print(x) # outer
for y in range(3):
pass
print(y) # 2
A for statement binds its variable in the surrounding scope, so y is still 2 after the loop. A comprehension’s x lives in the comprehension’s own scope and disappears when it finishes. In Python 2, list comprehensions leaked their variable, which occasionally overwrote a variable of the same name; Python 3 fixed that by giving them a scope. (CPython 3.12 inlines list, dict and set comprehensions for speed, per PEP 709, but deliberately keeps this visible behavior the same.)
The class-body trap
The same scoping rule has a less welcome consequence:
class Config:
factor = 10
scaled = [factor * n for n in range(3)]
# NameError: name 'factor' is not defined
Names defined in a class body are not visible to functions nested inside it; that is why methods have to write self.factor or Config.factor. The comprehension is such a nested scope. The exception is the first iterable, which Python evaluates in the enclosing scope before entering the comprehension, so this works:
class Config:
factor = 10
scaled = [f * n for f in [factor] for n in range(3)]
print(Config.scaled) # [0, 10, 20]
That trick is legal, but it is a workaround. If a class attribute needs computing, a small module-level helper function reads better.
I ran into the class-body version while moving a block of module-level constants into a class, and it looked like a Python bug until I remembered that comprehensions are functions in disguise. Since then, whenever a comprehension raises NameError for a name I can clearly see defined a few lines above, the first thing I check is whether I am inside a class body.
Late binding: lambdas created in a comprehension
funcs = [lambda: i for i in range(3)]
print([f() for f in funcs]) # [2, 2, 2]
Each lambda refers to the variable i in the comprehension’s scope, not to the value i had when the lambda was created. The comprehension finishes before any lambda is called, so they all look up i and find its final value, 2. This is the same late-binding rule that applies to closures in a for loop (see Python Functions); the comprehension just makes it easier to write by accident, because it reads like “one function per value”.
The fixes bind the current value at creation time:
funcs = [lambda i=i: i for i in range(3)] # default argument evaluated now
print([f() for f in funcs]) # [0, 1, 2]
from functools import partial
def power(base, exp):
return base ** exp
fs = [partial(power, exp=e) for e in range(3)]
print([f(2) for f in fs]) # [1, 2, 4]
The default-argument trick is common but has a cost: the function now has an optional parameter, and a caller who passes an argument silently overrides the captured value. partial makes the binding explicit and is harder to misuse. This bug is well known in GUI code that creates one callback per button in a comprehension, where every button ends up acting on the last item. When I see a list or dict of callbacks built this way, I check how the loop value is bound before reading anything else.
Generator Expressions: Lazy and One-Shot
Replacing the brackets with parentheses gives a generator expression. Nothing is computed when it is created; values are produced one at a time as something iterates over it.
total = sum(i ** 2 for i in range(1_000_000)) # no million-element list is built
When a generator expression is the only argument to a call, the extra parentheses can be dropped, as above.
They can be consumed only once
gen = (n * n for n in range(4))
print(list(gen)) # [0, 1, 4, 9]
print(list(gen)) # []
gen = (n for n in range(10))
print(5 in gen) # True
print(list(gen)) # [6, 7, 8, 9]
A generator is an iterator, and an exhausted iterator stays exhausted without raising an error. The second example is subtler: in consumes values up to and including the match, so the “rest” of the generator starts at 6. The bugs this causes look like missing data. A typical case is a function that computes len(list(rows)) for logging and then loops over rows again, and the second loop does nothing. If you need more than one pass, materialize a list once, or accept a callable that creates a fresh generator.
Only the first iterable is evaluated immediately
g = (n for n in undefined_name)
# NameError raised right here, at creation
g = (1 / n for n in [1, 0])
print("created fine")
list(g)
# ZeroDivisionError: division by zero, raised during iteration
The outermost iterable is evaluated when the generator expression is created; everything else (the element expression, the conditions, inner for clauses) runs lazily. That matters when the source changes between creation and consumption:
data = [1, 2, 3]
g = (n * 10 for n in data)
data.append(4)
print(list(g)) # [10, 20, 30, 40] same list object, mutated
data = [1, 2, 3]
g = (n * 10 for n in data)
data = [100]
print(list(g)) # [10, 20, 30] the generator kept the original object
The generator holds a reference to the original list object, so mutations are visible but rebinding the name is not. Laziness also moves exceptions: an error in the element expression surfaces wherever the generator is consumed, which may be far from the line that defined it. That makes tracebacks harder to read, and it is a reason to keep generator expressions close to where they are consumed.
Short-circuiting is where generators pay off
calls = 0
def big(n):
global calls
calls += 1
return n * n > 100
print(any([big(n) for n in range(1000)]), calls) # True 1000
calls = 0
print(any(big(n) for n in range(1000)), calls) # True 12
With a list, all 1000 calls happen before any sees the first element. With a generator, any stops at the first true value. all, next(gen, default) and itertools.islice benefit in the same way. For sum or max, which must see every element anyway, the gain is memory rather than work.
Memory: Measuring It Properly
import sys
lst = [i for i in range(100_000)]
gen = (i for i in range(100_000))
print(sys.getsizeof(lst), sys.getsizeof(gen)) # 800984 200
The generator object is small and constant in size, while the list grows with the input. But sys.getsizeof measures only the object you give it. For a list, that is the header plus an array of 8-byte pointers on a 64-bit build; the element objects are not included. To see the real allocation, use tracemalloc:
import tracemalloc
tracemalloc.start()
lst = [i * 1000 for i in range(100_000)]
current, peak = tracemalloc.get_traced_memory()
print(current) # 4000896: pointer array plus 100,000 int objects
tracemalloc.stop()
del lst
tracemalloc.start()
total = sum(i * 1000 for i in range(100_000))
current, peak = tracemalloc.get_traced_memory()
print(peak) # 528: only a few objects alive at any moment
tracemalloc.stop()
Here i * 1000 keeps the values out of CPython’s cache of small integers, which would otherwise make some ints look free. Exact numbers vary with the Python version and platform. The list costs roughly five times what getsizeof reports, because each int is a separate object, while the generator never holds more than a handful of values. This is why sum([...]) with brackets is worth fixing in code that processes large inputs, and harmless for ten elements.
On speed: a list comprehension is usually somewhat faster than the equivalent for loop with append, because it appends with a dedicated bytecode instead of looking up and calling the append method on every iteration. The difference is real but modest, varies between Python versions, and rarely matters compared to what the element expression itself does. Measure with timeit on your own data before rewriting a loop for speed. map with a lambda is typically no faster than a comprehension, since it still calls a Python function per element.
Walrus Inside Comprehensions
A common need is “compute something, keep it only if it passes a test”. Without assignment, you either compute it twice or nest a generator:
def parse(s):
try:
return int(s)
except ValueError:
return None
raw = ['3', 'x', '-4', '10']
print([v for s in raw if (v := parse(s)) is not None]) # [3, -4, 10]
The assignment expression := (Python 3.8+) computes parse(s) once, tests it, and reuses the value in the element expression. It also avoids a filter that is subtly wrong: [int(x) for x in data if x.isdigit()] rejects '-2', because - is not a digit, even though int accepts it. Deciding validity with a different rule from the one the conversion uses is a quiet source of dropped data.
Two scoping details are specific to walrus. Its target is bound in the enclosing scope, deliberately, so that you can inspect the last value afterwards:
print(v) # 10: the walrus target is visible after the comprehension
And it may not rebind the iteration variable: [i := 0 for i in range(3)] is a SyntaxError (“assignment expression cannot rebind comprehension iteration variable”). Before walrus, the idiom for a computed intermediate was a one-element inner loop, for parts in [line.split(',')]. It still works, and CPython 3.9+ optimizes it, but walrus says what it means.
Where a Plain Loop Is Better
Comprehensions are for building a collection from an iterable. They become a liability in a few recognizable cases:
- Side effects.
[results.append(x) for x in data]or[print(x) for x in data]builds a list ofNonevalues just to throw it away, and hides the side effect inside an expression that looks like data construction. Write aforloop. - More than two
forclauses, or conditions spread over several of them. The reader has to rebuild the nesting in their head. A triple loop filtered byif x < y < zisitertools.combinations(range(10), 3), which states the intent directly and generates only the valid tuples instead of filtering 1000 candidates. - Chained conditional expressions.
'A' if s >= 90 else 'B' if s >= 80 else 'C' if ...is correct but hard to review. A small function withifstatements, called from the comprehension, keeps the comprehension readable and the logic testable. - Exception handling. There is no
tryinside a comprehension. Move it into a helper, asparseabove does. - Debugging. You cannot put a breakpoint or a
printon one iteration of a comprehension. When one misbehaves, the fastest route is often to rewrite it as a loop temporarily.
One pitfall is not about comprehensions at all, but comprehensions are its standard fix:
m = [[0] * 3] * 3
m[0][0] = 1
print(m) # [[1, 0, 0], [1, 0, 0], [1, 0, 0]] three references to one row
m = [[0] * 3 for _ in range(3)]
m[0][0] = 1
print(m) # [[1, 0, 0], [0, 0, 0], [0, 0, 0]]
* repeats references, so all three rows are the same list. The comprehension evaluates [0] * 3 once per iteration, creating a new row each time. ([0] * 3 itself is fine, because ints are immutable.) The data types article covers aliasing in more depth.
Comprehension, generator or plain loop?
Use a list comprehension when you need the result more than once, or need len and indexing; a generator expression when it feeds a single consumer such as sum, any or a for loop; and a plain loop when the body has side effects, branches heavily, or needs try.
Exercises
- Squares of the multiples of 3 from 1 to 20. Expected:
[9, 36, 81, 144, 225, 324]. - Turn
[('Alice', 25), ('Bob', 30), ('Charlie', 35)]into{name: age}. - Flatten
[[1, 2, 3], [4, 5, 6], [7, 8, 9]], keeping even numbers only. Expected:[2, 4, 6, 8]. - Label
[5, 15, 30, 8, 22, 28]as"cold"(below 10),"mild"(10 to 25) or"hot"(above 25). - Explain why
handlers = {name: lambda: print(name) for name in ["a", "b"]}printsbfor both keys, and fix it.
Answers
print([x ** 2 for x in range(1, 21) if x % 3 == 0])
print({name: age for name, age in [('Alice', 25), ('Bob', 30), ('Charlie', 35)]})
matrix = [[1, 2, 3], [4, 5, 6], [7, 8, 9]]
print([num for row in matrix for num in row if num % 2 == 0])
temps = [5, 15, 30, 8, 22, 28]
print(['cold' if t < 10 else 'mild' if t <= 25 else 'hot' for t in temps])
# ['cold', 'mild', 'hot', 'cold', 'mild', 'hot']
# 5: each lambda looks up `name` when called, after the comprehension has ended
handlers = {name: (lambda name=name: print(name)) for name in ["a", "b"]}
handlers["a"]() # a
Related Articles
- Python Functions: Parameters, Defaults, Closures, and the Traps Behind Them
- Python Data Types Explained: Mutability, Aliasing, Float Rounding, is vs == and Truthiness
- Python Decorators: @decorator Syntax and functools.wraps
- Arrays vs Linked Lists for Coding Interviews: Access Costs, Two Pointers and Common Traps