Python Classes: Class vs Instance Attributes, super() and the MRO, @property and Dataclasses

Key takeaways

Python classes are namespaces with attribute lookup rules, not rigid blueprints. Knowing how instance and class attributes resolve, how super() follows the MRO, and how __eq__ affects __hash__ prevents the bugs that bite most intermediate Python code.

Classes are namespaces with lookup rules

The usual explanation of classes, “a class is a blueprint, an instance is the thing built from it”, is fine for day one and then quietly misleads you. In Python, a class is an object that holds a dictionary of attributes, and an instance is another object with its own dictionary plus a pointer to its class. Almost every confusing behavior in this article comes from one rule:

When you read obj.name, Python checks the instance’s __dict__, then the class, then each base class in MRO order. When you assign obj.name = value, Python writes to the instance’s __dict__ (unless a descriptor such as @property intercepts it).

Reads search upward; writes land on the instance. Keep that asymmetry in mind and the rest follows.

class Person:
    species = "human"              # class attribute: one copy, on the class

    def __init__(self, name, age):
        self.name = name           # instance attributes: one copy per object
        self.age = age

    def greet(self):
        return f"Hi, I'm {self.name}."

p = Person("Alice", 25)
print(p.greet())          # Hi, I'm Alice.
print(p.__dict__)         # {'name': 'Alice', 'age': 25}
print(p.species)          # human  (found on the class, not the instance)

self is not a keyword. p.greet() is sugar for Person.greet(p): the method is a plain function stored on the class, and the attribute lookup binds the instance as the first argument. If you forget self in a method definition, you get TypeError: Person.greet() takes 0 positional arguments but 1 was given, which is Python telling you it passed the instance and the function had nowhere to put it.


The shared mutable class attribute bug

This is the class-related bug I see most often in code review, and it survives tests because a single instance behaves correctly.

class Team:
    members = []                    # looks like a default, is actually shared

    def __init__(self, name):
        self.name = name

    def add(self, person):
        self.members.append(person)

a = Team("a")
b = Team("b")
a.add("Alice")
print(b.members, a.members is b.members)
# ['Alice'] True

self.members.append is a read of members followed by a mutation. There is no instance attribute, so the lookup finds the class list and mutates it. Every team shares one list.

The same rule produces the opposite surprise with immutable values:

class Counter:
    total = 0
    def __init__(self):
        self.total += 1             # read from class, write to instance

c1, c2 = Counter(), Counter()
print(Counter.total, c1.total, c2.total, c1.__dict__)
# 0 1 1 {'total': 1}

self.total += 1 reads 0 from the class, then assigns 1 to a new instance attribute that shadows it. The class counter never moves. If you really want a per-class counter, write Counter.total += 1 (or type(self).total += 1, which also counts per subclass).

The fix for the first bug is to create mutable state in __init__:

class Team:
    def __init__(self, name):
        self.name = name
        self.members = []           # fresh list per instance

Class attributes are the right tool for constants and for values that are genuinely shared (a registry, a default that is immutable). It is the same trap as mutable default arguments covered in Python functions: a mutable object created once at definition time and reused forever.


__new__ vs __init__

Person("Alice", 25) does two steps. __new__ creates the object, then __init__ initializes it. For almost every class you only write __init__; __init__ must return None, so it can set attributes but cannot replace the object being built.

You need __new__ when the object’s value must be decided at creation time, which mainly means subclassing immutable types:

class Upper(str):
    def __new__(cls, value):
        return super().__new__(cls, value.upper())

    def __init__(self, value):
        print("init got", repr(value), "self is", repr(self))

u = Upper("hi")
# init got 'hi' self is 'HI'

By the time __init__ runs, the string is already 'HI' and cannot be changed, because str is immutable. Trying to do the uppercasing in __init__ is a common first attempt that silently does nothing useful. Other legitimate uses of __new__ (caching instances, returning a subclass) exist, but if you find yourself reaching for it in ordinary application code, a @classmethod factory is usually clearer.

__del__ also exists, but do not use it for cleanup: when it runs depends on reference counting and garbage collection, and on interpreter shutdown it may not run usefully at all. Resources such as files, sockets and locks belong in a context manager (with), as described in Python exception handling.


Inheritance, super() and the MRO

Single inheritance is straightforward: the subclass overrides what it needs and calls super() to reuse the parent’s behavior.

class Employee:
    def __init__(self, name, salary):
        self.name = name
        self.salary = salary

class Manager(Employee):
    def __init__(self, name, salary, team_size):
        super().__init__(name, salary)   # without this, self.name never exists
        self.team_size = team_size

Forgetting the super().__init__ call does not fail at construction. It fails later with AttributeError: 'Manager' object has no attribute 'name' in some unrelated method, which is why it is worth checking every overridden __init__ explicitly.

What super() actually calls

super() does not mean “my parent”. It means “the next class after me in the MRO of the object’s actual type”. With multiple inheritance, the difference matters:

class Base:
    def __init__(self):
        print("Base")

class A(Base):
    def __init__(self):
        print("A enter"); super().__init__(); print("A exit")

class B(Base):
    def __init__(self):
        print("B enter"); super().__init__(); print("B exit")

class C(A, B):
    def __init__(self):
        print("C enter"); super().__init__(); print("C exit")

C()
print([k.__name__ for k in C.__mro__])
C enter
A enter
B enter
Base
B exit
A exit
C exit
['C', 'A', 'B', 'Base', 'object']

Inside A.__init__, super() called B.__init__, a class A has never heard of. That is the C3 linearization at work: the MRO lists each class once, keeps every class before its bases, and preserves the left-to-right order of the class C(A, B) statement. It is also why Base ran exactly once instead of twice.

Two consequences:

  1. Cooperative inheritance only works if everyone calls super(). If A.__init__ called Base.__init__(self) directly, B.__init__ would be skipped entirely. Mixing explicit parent calls with super() is the classic source of “my mixin’s setup never ran”.
  2. Signatures have to be compatible along the chain, because you do not know which class your super() call reaches. The usual pattern is to accept **kwargs, consume your own arguments, and pass the rest along.

Python refuses hierarchies it cannot linearize:

class X: pass
class Y(X): pass
class Z(X, Y): pass
# TypeError: Cannot create a consistent method resolution order (MRO) for bases X, Y

X is listed before Y, but Y is a subclass of X and must come first. Swapping to class Z(Y, X) fixes it.

My honest take after untangling a few of these: multiple inheritance of stateful classes is rarely worth it. Small mixins that add methods and no __init__ (a JsonMixin with to_json, for example) are fine. When two parents both need constructor arguments, composition (holding an instance as an attribute) is usually easier to reason about than a correct cooperative super() chain.


Encapsulation: underscores, name mangling and @property

Python has no private keyword. There are two conventions:

  • _name: “internal, don’t touch from outside”. Purely a convention; nothing enforces it.
  • __name (two leading underscores, no trailing ones): name mangling. Inside the class body, self.__balance is rewritten to self._Account__balance.
class Account:
    def __init__(self):
        self.__balance = 0

acc = Account()
acc.__balance
# AttributeError: 'Account' object has no attribute '__balance'
print(acc._Account__balance)   # 0  (still reachable)

Name mangling exists to avoid accidental collisions with subclasses that use the same attribute name, not to provide security. In practice, a single underscore is the norm; double underscores mostly make debugging and subclassing more annoying.

@property instead of getters and setters

Java-style get_balance() / set_balance() methods are unnecessary in Python because you can start with a plain attribute and later turn it into a property without changing any caller:

class Celsius:
    def __init__(self, degrees):
        self.degrees = degrees        # goes through the setter below

    @property
    def degrees(self):
        return self._degrees

    @degrees.setter
    def degrees(self, value):
        if value < -273.15:
            raise ValueError(f"{value} is below absolute zero")
        self._degrees = value

    @property
    def fahrenheit(self):             # read-only computed attribute
        return self._degrees * 9 / 5 + 32

t = Celsius(100)
print(t.fahrenheit)                   # 212.0
Celsius(-300)
# ValueError: -300 is below absolute zero
t.fahrenheit = 1
# AttributeError: property 'fahrenheit' of 'Celsius' object has no setter

Note that __init__ assigns self.degrees, not self._degrees, so validation also applies at construction. Two pitfalls: storing the value under the same name as the property (self.degrees = value inside the setter) recurses until RecursionError, and a property that does expensive work (a database query) looks like cheap attribute access to callers, which is a readability trap. If it is slow, make it a method.


Dunder methods and the __eq__ / __hash__ interplay

Special (“dunder”) methods let your objects work with built-in syntax: __repr__ for debugging output, __len__ and __getitem__ for containers, __add__ for +, __eq__ for ==.

class Vector:
    def __init__(self, x, y):
        self.x, self.y = x, y

    def __repr__(self):
        return f"Vector({self.x!r}, {self.y!r})"

    def __add__(self, other):
        if not isinstance(other, Vector):
            return NotImplemented     # let Python try other.__radd__ or raise
        return Vector(self.x + other.x, self.y + other.y)

print(Vector(1, 2) + Vector(3, 4))    # Vector(4, 6)
Vector(1, 2) + 1
# TypeError: unsupported operand type(s) for +: 'Vector' and 'int'

Return NotImplemented (the constant, not the NotImplementedError exception) for types you don’t handle. Python then tries the reflected operation on the other operand and produces a proper TypeError if nothing works, instead of your method crashing with an AttributeError on other.x.

Always define __repr__. Without it, logs and debugger output show <__main__.Vector object at 0x0000017E56274EC0>, which tells you nothing. __str__ falls back to __repr__, so one good __repr__ covers both.

Defining __eq__ removes __hash__

This one surprises people who add equality to an existing class:

class Point:
    def __init__(self, x, y):
        self.x, self.y = x, y

    def __eq__(self, other):
        if not isinstance(other, Point):
            return NotImplemented
        return (self.x, self.y) == (other.x, other.y)

print(Point.__hash__)   # None
{Point(1, 2)}
# TypeError: unhashable type: 'Point'

Python sets __hash__ to None when a class defines __eq__ without __hash__, because the default identity hash would break the rule that equal objects must have equal hashes. The fix is to hash the same fields equality compares: def __hash__(self): return hash((self.x, self.y)).

The catch is mutability. If you mutate a hashed object after putting it in a dict, the dict can no longer find it:

k = Point(1, 2)          # with __hash__ defined as above
d = {k: "v"}
k.x = 99
print(Point(1, 2) in d, k in d)   # False False

The entry is still in the dict, stored under the old hash, and unreachable by either key. So: hash only objects whose equality fields never change, which in practice means making them immutable.


Dataclasses: less boilerplate, fewer of these bugs

For classes that are mostly fields, @dataclass generates __init__, __repr__ and __eq__ from type-annotated class attributes, and it actively guards against the mutable default bug:

from dataclasses import dataclass, field

@dataclass
class Cart:
    owner: str
    items: list = []
# ValueError: mutable default <class 'list'> for field items is not allowed: use default_factory
@dataclass
class Cart:
    owner: str
    items: list[str] = field(default_factory=list)   # new list per instance

x, y = Cart("x"), Cart("y")
x.items.append("apple")
print(x, y)
# Cart(owner='x', items=['apple']) Cart(owner='y', items=[])
print(Cart("a") == Cart("a"), Cart.__hash__)
# True None

Note Cart.__hash__ is None: a plain dataclass gets __eq__ and therefore loses hashing, exactly as described above. The options that change this:

  • frozen=True makes fields read-only (FrozenInstanceError: cannot assign to field 'amount') and, together with the default eq=True, generates a __hash__ from the fields. This is the safe way to get value objects usable as dict keys.
  • slots=True (Python 3.10+) generates __slots__, so instances have no __dict__. They use less memory and typos in attribute names fail loudly: AttributeError: 'P' object has no attribute 'y' and no __dict__ for setting new attributes.
@dataclass(frozen=True, slots=True)
class Money:
    amount: int
    currency: str = "USD"

m = Money(5)
print(hash(m) == hash(Money(5)), {m, Money(5)})
# True {Money(amount=5, currency='USD')}

One oddity I ran into while testing this on Python 3.13: with frozen=True, slots=True together, assigning an attribute that is not a declared field produces a confusing TypeError: super(type, obj): obj (instance of Money) is not an instance or subtype of type (Money). rather than a clean FrozenInstanceError. The object is still protected; the message is just misleading. If you see that error on a frozen, slotted dataclass, look for a typo in an attribute assignment.

Dataclasses are not always the answer. Generated __eq__ compares every field, which is wrong for entities identified by an ID (two User rows with the same ID and a different cached field are the same user). Use field(compare=False) for such fields, or write the class by hand.


@classmethod, @staticmethod and abstract base classes

@classmethod receives the class as cls, which makes it the standard way to write alternative constructors that also work for subclasses:

from datetime import date

class Person:
    def __init__(self, name, age):
        self.name = name
        self.age = age

    @classmethod
    def from_birth_year(cls, name, birth_year):
        return cls(name, date.today().year - birth_year)

cls(...) rather than Person(...) means Student.from_birth_year(...) returns a Student. @staticmethod receives neither the instance nor the class; it is a function namespaced under the class. If it does not logically belong to the class, a module-level function is usually simpler.

Abstract base classes enforce that subclasses implement required methods, and they fail at instantiation rather than at the first call:

from abc import ABC, abstractmethod

class Shape(ABC):
    @abstractmethod
    def area(self): ...

class Square(Shape):
    def __init__(self, side):
        self.side = side
    # forgot area()

Square(2)
# TypeError: Can't instantiate abstract class Square without an implementation for abstract method 'area'

(Older versions word this as Can't instantiate abstract class Square with abstract method area.) Much Python code relies on duck typing instead: any object with an area() method works. ABCs are worth it when you publish an interface for others to implement; for type checking without inheritance, typing.Protocol gives the same contract structurally.


A small example that uses the pieces together

from dataclasses import dataclass, field
from datetime import date, timedelta

@dataclass(frozen=True)
class Book:
    isbn: str
    title: str

@dataclass
class Member:
    member_id: str
    name: str
    borrowed: set[str] = field(default_factory=set)

class Library:
    LOAN_DAYS = 14                      # immutable class constant: fine to share

    def __init__(self):
        self.books: dict[str, Book] = {}
        self.on_loan: dict[str, str] = {}   # isbn -> member_id

    def add(self, book: Book) -> None:
        self.books[book.isbn] = book

    def borrow(self, member: Member, isbn: str) -> date:
        if isbn not in self.books:
            raise KeyError(f"unknown book {isbn}")
        if isbn in self.on_loan:
            raise ValueError(f"{self.books[isbn].title!r} is already on loan")
        self.on_loan[isbn] = member.member_id
        member.borrowed.add(isbn)
        return date.today() + timedelta(days=self.LOAN_DAYS)

lib = Library()
lib.add(Book("978-1", "Python Basics"))
dana = Member("M001", "Dana")
due = lib.borrow(dana, "978-1")
print(dana)   # Member(member_id='M001', name='Dana', borrowed={'978-1'})

Book is a frozen value object (hashable, safe to share), Member holds mutable state created per instance via default_factory, and Library keeps loan state in one place instead of an available flag duplicated on each book. Errors are raised as exceptions instead of returned as strings, so callers cannot silently ignore them.


Next in the series

Python modules and packages comes next, then decorators, which explains how @property and @classmethod are built.