Python Data Types Explained: Mutability, Aliasing, Float Rounding, is vs == and Truthiness

Key takeaways

Python variables are names bound to objects, and the object's type decides whether it can change in place. That one idea explains shared-list bugs, why strings need join, why tuples can be dict keys, and why is and == disagree. Numbers add their own traps: 0.1 + 0.2, banker's rounding and when to reach for Decimal.

Introduction

Most tutorials on Python data types list the types and their methods. That is useful reference material, but it does not explain the bugs beginners actually hit: a list that changes “by itself,” a default argument that remembers old values, a price that comes out as 2.67 instead of 2.68, or a check with is that works in the REPL and fails in production.

All of those come from a small number of ideas:

  1. A variable is a name bound to an object, not a box that holds a value.
  2. Some objects are mutable (list, dict, set) and some are immutable (int, float, str, tuple, frozenset).
  3. Floats are binary approximations; ints are exact and unbounded.

This article works through the built-in types with those ideas in mind. The container types (list, tuple, set) get a focused treatment here; for a side-by-side performance and memory comparison of the three, see list vs tuple vs set. All outputs below were produced with Python 3.11.


Names, objects and mutability

a = [1, 2, 3]
b = a
b.append(4)
print(a, a is b)
[1, 2, 3, 4] True

b = a did not copy the list. It created a second name for the same object. append changes that object in place, so the change is visible through both names. Compare with an immutable type:

x = 10
y = x
y += 1
print(x, y)   # 10 11

y += 1 cannot change the integer 10 (ints are immutable), so Python builds a new object 11 and rebinds y to it. x still points to 10. Same syntax, different outcome, and the only difference is mutability.

Copying: shallow vs deep

c = a.copy()      # also: list(a), a[:]
c.append(5)
print(a, c)
[1, 2, 3, 4] [1, 2, 3, 4, 5]

A shallow copy creates a new outer list but reuses the inner objects. With nested lists that matters:

import copy

m = [[1, 2], [3, 4]]
shallow = m.copy()
shallow[0][0] = 100
print(m)                  # [[100, 2], [3, 4]]  inner list was shared

deep = copy.deepcopy(m)
deep[0][0] = 999
print(m, deep)            # [[100, 2], [3, 4]] [[999, 2], [3, 4]]

The two aliasing bugs everyone hits once

Multiplying a nested list:

grid = [[0] * 3] * 3
grid[0][0] = 1
print(grid)
[[1, 0, 0], [1, 0, 0], [1, 0, 0]]

* 3 repeats the reference to one inner list three times. Build each row separately instead:

grid = [[0] * 3 for _ in range(3)]
grid[0][0] = 1
print(grid)   # [[1, 0, 0], [0, 0, 0], [0, 0, 0]]

Mutable default arguments:

def add_item(item, bucket=[]):
    bucket.append(item)
    return bucket

print(add_item("a"))   # ['a']
print(add_item("b"))   # ['a', 'b']   ← the same list again

The default [] is evaluated once, when the function is defined, and then reused on every call. The standard fix is a None sentinel:

def add_item(item, bucket=None):
    if bucket is None:
        bucket = []
    bucket.append(item)
    return bucket

I have seen this one survive code review more than once, because the function returns the right answer the first time it is called in a test. It only shows up when the same process calls it repeatedly, such as a long-running web worker where a “per-request” list slowly accumulates entries from earlier requests. Linters like Pylint (dangerous-default-value) and Ruff (B006) flag it, which is a good reason to turn them on.


Numbers: int, float and Decimal

int is exact and unbounded

print(2 ** 100)
1267650600228229401496703205376

Python integers grow as needed; there is no overflow like a 32- or 64-bit int in C or Java. The cost is that very large integers are slower than machine integers, but for everyday code you never have to think about overflow.

Division has two operators, and floor division rounds toward negative infinity, not toward zero:

print(7 / 2, 7 // 2, -7 // 2, -7 % 3)   # 3.5 3 -4 2
print(int(3.9), int(-3.9))              # 3 -3   int() truncates toward zero

float is a binary approximation

print(0.1 + 0.2)
print(0.1 + 0.2 == 0.3)
0.30000000000000004
False

This is not a Python bug; every language using IEEE 754 doubles behaves the same way. 0.1 cannot be represented exactly in binary, just as 1/3 cannot be written exactly in decimal. You can see the value that is really stored:

from decimal import Decimal
print(Decimal(0.1))
0.1000000000000000055511151231257827021181583404541015625

Practical consequences:

  • Never compare floats with ==. Use math.isclose(0.1 + 0.2, 0.3), which returns True.
  • Summing many floats accumulates error: sum([0.1] * 10) is 0.9999999999999999, while math.fsum([0.1] * 10) is 1.0.
  • Formatting hides the error but does not remove it: f"{0.1 + 0.2:.2f}" prints 0.30.

round() uses banker’s rounding

print(round(0.5), round(1.5), round(2.5), round(3.5))
print(round(2.675, 2))
0 2 2 4
2.67

Exact halves round to the nearest even number, which avoids a systematic upward bias when rounding large sets of values. The second result is a different issue: 2.675 is actually stored as 2.67499999999999982236431605997495353221893310546875, so it is below the halfway point before rounding even starts.

When the rules must match what humans expect, such as invoices, use Decimal built from strings:

from decimal import Decimal, ROUND_HALF_UP

print(Decimal("0.1") + Decimal("0.2"))                                   # 0.3
print(Decimal("2.675").quantize(Decimal("0.01"), rounding=ROUND_HALF_UP)) # 2.68

Decimal(0.1) (from a float) inherits the float’s error, which is why the string form matters. The trade-off is speed: Decimal arithmetic is noticeably slower than float, so use it where exactness is required, not for scientific computation. A common alternative for money is storing integer cents.

I once traced an off-by-one-cent report back to exactly this chain: prices parsed as floats, multiplied by quantities, rounded with round(x, 2). Every step looked reasonable, and totals were right most of the time. Switching the parsing to Decimal(str_value) fixed it; tweaking the rounding would not have.

bool is a subclass of int

print(isinstance(True, int), True + True)   # True 2
print({1: "int", 1.0: "float", True: "bool"})  # {1: 'bool'}

Because 1 == 1.0 == True and they hash the same, they are the same dictionary key. The first key object is kept and the last value wins. It is rare in practice, but it explains surprising results when mixing numeric and boolean keys.


Strings are immutable

s = "hello"
s[0] = "H"
TypeError: 'str' object does not support item assignment

Every string method returns a new string; the original never changes:

s = "Python"
print(s.upper(), s, s.replace("P", "J"))   # PYTHON Python Jython

A common beginner mistake is calling name.strip() on its own line and expecting name to change. You need name = name.strip().

Immutability has a performance side. Building a string with += in a loop creates new string objects repeatedly; for a large number of pieces, collect them in a list and join once:

parts = []
for i in range(3):
    parts.append(str(i))
print(",".join(parts))   # 0,1,2

(CPython has an optimization that sometimes makes += on strings fast, but it is an implementation detail and does not apply in all cases, so join is the reliable pattern.)

Python also refuses to mix strings and numbers implicitly:

"Age: " + 30
TypeError: can only concatenate str (not "int") to str

Use an f-string (f"Age: {30}") or convert explicitly with str(). The reverse conversion, int("42"), raises ValueError for input like "42.5" or "abc", so validate user input.


Lists

Lists are ordered, mutable sequences. The operations worth knowing well:

nums = [1, 2, 3, 4, 5]
print(nums[1:4], nums[::-1])   # [2, 3, 4] [5, 4, 3, 2, 1]

sub = nums[1:3]    # slicing creates a new list
sub[0] = 100
print(nums, sub)   # [1, 2, 3, 4, 5] [100, 3]
  • append(x) adds one element; extend(iterable) adds each element. Appending [3, 4] to [1, 2] gives [1, 2, [3, 4]], while extending gives [1, 2, 3, 4].
  • list.sort() sorts in place and returns None. Writing nums = nums.sort() leaves you with None. Use sorted(nums) when you want a new list.
  • x in list is a linear scan. For repeated membership checks on large collections, convert to a set once.
  • Removing from the front (pop(0)) shifts every remaining element. For queues use collections.deque.

Tuples and hashability

Tuples are ordered and immutable, which makes them hashable as long as their contents are hashable. That is why they can be dictionary keys and set members while lists cannot:

locations = {(10, 20): "A", (30, 40): "B"}
print(locations[(10, 20)])   # A

{[1, 2]: "x"}
TypeError: unhashable type: 'list'

(Newer Python versions word this error slightly differently, but it names the unhashable type either way.)

Immutability is shallow. A tuple that contains a list can still have that list modified, and then it is no longer hashable:

t = (1, 2, [3, 4])
t[2].append(5)
print(t)       # (1, 2, [3, 4, 5])
hash(t)        # TypeError: unhashable type: 'list'

A one-element tuple needs a trailing comma: (42,). (42) is just the integer 42 in parentheses. Unpacking is where tuples shine in everyday code:

first, *rest, last = (1, 2, 3, 4, 5)
print(first, rest, last)   # 1 [2, 3, 4] 5

a, b = 1, 2
a, b = b, a                # swap without a temp variable

Dictionaries

Dicts map hashable keys to values with average O(1) lookup, implemented as hash tables (see hash tables for how that works).

Missing keys

person = {"name": "Alice", "age": 25}
person["job"]
KeyError: 'job'

Choose deliberately between person["job"] (a missing key is a bug, fail loudly) and person.get("job", "N/A") (a missing key is expected). Using .get() everywhere hides typos in key names, which then surface much later as a mysterious None.

Insertion order is guaranteed

Since Python 3.7, dicts preserve insertion order as part of the language specification (CPython 3.6 already did this as an implementation detail):

d = {}
d["b"] = 1; d["a"] = 2; d["c"] = 3
print(list(d))   # ['b', 'a', 'c']

That makes dict.fromkeys a handy order-preserving deduplication: list(dict.fromkeys([3, 1, 3, 2, 1])) gives [3, 1, 2].

Counting and grouping

from collections import defaultdict, Counter

words = ["apple", "banana", "apple", "cherry", "banana", "apple"]
print(Counter(words).most_common(2))   # [('apple', 3), ('banana', 2)]

groups = defaultdict(list)
for name, grade in [("Alice", "A"), ("Bob", "B"), ("Carol", "A")]:
    groups[grade].append(name)
print(dict(groups))   # {'A': ['Alice', 'Carol'], 'B': ['Bob']}

Do not add or remove keys while iterating over a dict; Python raises RuntimeError: dictionary changed size during iteration. Iterate over list(d) if you need to delete as you go.


Sets

Sets hold unique, hashable elements with O(1) average membership tests, and they are unordered.

a = {1, 2, 3, 4}
b = {3, 4, 5, 6}
print(a | b, a & b, a - b, a ^ b)
{1, 2, 3, 4, 5, 6} {3, 4} {1, 2} {1, 2, 5, 6}

Two traps: {} creates an empty dict, not a set (type({}) is dict; use set()), and set(my_list) removes duplicates but loses the original order. Use the dict.fromkeys trick above when order matters.


is vs ==

== compares values. is compares identity: whether two names refer to the same object.

print([] == [], [] is [])   # True False

x = int("1000")
y = int("1000")
print(x == y, x is y)       # True False

Where it gets confusing is small integers. CPython pre-creates the integers from -5 to 256 and reuses them, so is happens to return True for those. Constants in the same compiled block can also be shared, which is why x = 1000; y = 1000; x is y may print True in a script. None of this is guaranteed by the language. Python 3.8 and later even emit a warning when you write x is 1000 (on 3.11: SyntaxWarning: "is" with a literal. Did you mean "=="?; the exact wording varies by version).

The rule is short: use is for None (and other singletons like sentinel objects), == for everything else.

This is a bug I have seen pass every local test: a check like if status_code is 200 works during development because 200 is in the small-int cache, and then an equivalent check against a value above 256 fails. Nothing about the code looks wrong until you know about the cache.


Truthiness

Every object has a truth value. These are falsy; almost everything else is truthy:

for v in [0, 0.0, "", [], {}, set(), None]:
    print(repr(v), bool(v))    # all False

for v in ["0", [0], " "]:
    print(repr(v), bool(v))    # all True

if items: is idiomatic for “non-empty.” The trap is using truthiness where 0 or "" are legitimate values:

def get_limit(limit=None):
    limit = limit or 10
    return limit

print(get_limit(0))   # 10, but the caller asked for 0

When None means “not provided,” test for it explicitly: if limit is None: limit = 10.


Choosing a type

TypeOrderedDuplicatesMutableHashableTypical use
listYesYesYesNoSequences you change
tupleYesYesNoIf contents areFixed records, dict keys
dictInsertion order (3.7+)Keys uniqueYesNoLookup by key
setNoNoYesNoUniqueness, membership, set math
strYesYesNoYesText
int / floatn/an/aNoYesExact integers / approximate reals
Decimaln/an/aNoYesMoney and exact decimal rounding