Matplotlib Basics: Line, Bar, Histogram and Scatter Plots, Subplots and Report-Ready Styling
Key takeaways
Matplotlib tutorial: line plots, bar charts, histograms, scatter, subplots, styling, and report-ready figures. pyplot vs object-oriented API, fonts, and common pitfalls.
Introduction
Matplotlib is the de facto standard plotting library in Python. It is not part of the standard library, but pandas’ .plot(), Seaborn and many scientific packages draw through it, so knowing its model pays off even if you mostly use those higher-level tools.
That model has two layers. A Figure is the whole image (the window or the PNG file); it contains one or more Axes, and each Axes is one plot area with its own x and y axis, title, legend and data. Most confusion in Matplotlib comes from mixing up those two, for example calling plt.title() when you meant the title of one specific subplot. The plt.* functions (the pyplot interface) act on “the current figure” and “the current axes”, which is convenient for a quick single chart. The object-oriented interface (fig, ax = plt.subplots(), then ax.plot(...)) names the target explicitly and is what the later sections use for anything with more than one chart.
Matplotlib basics
Installation
pip install matplotlib
First plot
import matplotlib.pyplot as plt
# Data
x = [1, 2, 3, 4, 5]
y = [2, 4, 6, 8, 10]
# Plot
plt.plot(x, y)
plt.xlabel('X')
plt.ylabel('Y')
plt.title('Line plot')
plt.show()
plt.plot creates a figure and axes implicitly the first time it is called, draws the line, and plt.show() opens a window (or renders inline in Jupyter). In a plain script, show() blocks until you close the window. On a server or in a Docker container without a display, there is no window to open; Matplotlib then falls back to a non-interactive backend and warns that the figure cannot be shown, and the right approach is to call plt.savefig('plot.png') instead, optionally forcing the backend with matplotlib.use('Agg') before importing pyplot.
Line plots
Basic line plot
import matplotlib.pyplot as plt
import numpy as np
x = np.linspace(0, 10, 100)
y1 = np.sin(x)
y2 = np.cos(x)
plt.plot(x, y1, label='sin(x)', color='blue', linestyle='-')
plt.plot(x, y2, label='cos(x)', color='red', linestyle='--')
plt.xlabel('x')
plt.ylabel('y')
plt.title('Trigonometric functions')
plt.legend()
plt.grid(True)
plt.show()
np.linspace(0, 10, 100) produces 100 evenly spaced x values, and Matplotlib connects consecutive points with straight segments, so the number of points decides how smooth a curve looks: with 10 points, sin(x) would visibly turn into a jagged polyline. Calling plt.plot twice draws both lines on the same axes, and plt.legend() builds the legend from the label= arguments; a line without a label simply does not appear in it, and calling legend() when no line has a label produces an empty legend and a “No artists with labels found” warning. Style arguments can be combined into a format string ('r--' for a red dashed line), which is shorter but less readable than the keywords shown here.
Bar charts
Vertical bars
categories = ['A', 'B', 'C', 'D']
values = [25, 40, 30, 55]
plt.bar(categories, values, color='skyblue')
plt.xlabel('Category')
plt.ylabel('Value')
plt.title('Bar chart')
plt.show()
Horizontal bars
plt.barh(categories, values, color='lightgreen')
plt.xlabel('Value')
plt.ylabel('Category')
plt.title('Horizontal bar chart')
plt.show()
Passing strings as x values makes Matplotlib treat them as categories and place them in the order given, one unit apart. That order is the one thing worth controlling: bars sorted by value are much easier to compare than bars in arbitrary order, so sort the data first (for example with sorted(zip(values, categories)) or df.sort_values(...) in pandas). Horizontal bars are the better choice when category names are long, since the labels then run along the y axis instead of overlapping under each bar. Note that barh draws the first category at the bottom; plt.gca().invert_yaxis() puts it at the top, which matches reading order. Unlike line charts, bar charts should start at zero: a y axis that begins at 20 makes a bar of 25 look a fraction of one of 40, which misrepresents the data.
Histogram
A histogram answers a different question than the line and bar charts above: instead of plotting an exact value per category, bins=30 groups the 1,000 sampled values into 30 equal-width ranges and plots how many values fall into each — the shape that emerges is what tells you whether the underlying data is roughly normal, skewed, or has multiple peaks. bins is worth tuning deliberately rather than leaving at a default: too few bins hides real structure by lumping everything together, too many bins turns the histogram into noisy, spiky bars that overfit the specific sample rather than showing the underlying distribution.
# Normal-ish sample
data = np.random.randn(1000)
plt.hist(data, bins=30, color='purple', alpha=0.7, edgecolor='black')
plt.xlabel('Value')
plt.ylabel('Frequency')
plt.title('Histogram')
plt.show()
Scatter plot
A scatter plot is the right tool specifically when you want to see the relationship between two continuous variables rather than a single variable’s distribution — every point here is positioned by its x/y pair, with two more dimensions of information layered on through c (color) and s (size). This four-dimensional-in-one-chart trick is genuinely useful for exploratory analysis (spot a cluster where both color and size correlate with position), but it’s easy to overdo — past two or three encoded dimensions, a scatter plot usually becomes harder to read than several simpler ones.
x = np.random.rand(50)
y = np.random.rand(50)
colors = np.random.rand(50)
sizes = 1000 * np.random.rand(50)
plt.scatter(x, y, c=colors, s=sizes, alpha=0.5, cmap='viridis')
plt.colorbar()
plt.xlabel('X')
plt.ylabel('Y')
plt.title('Scatter plot')
plt.show()
Multiple subplots
plt.subplots(2, 2, ...) is the object-oriented API this guide’s introduction alluded to — instead of pyplot’s implicit “whatever was drawn most recently is the current axes” state machine, axes here is an explicit 2D array of Axes objects you index directly (axes[0, 0], axes[1, 1]), which is what makes it possible to build up several unrelated charts in one figure without pyplot’s implicit state ever becoming ambiguous about which chart a given .plot() call belongs to. This is exactly the pattern the pitfalls section below recommends once a script needs more than one or two charts.
fig, axes = plt.subplots(2, 2, figsize=(10, 8))
# Top-left
axes[0, 0].plot([1, 2, 3], [1, 4, 9])
axes[0, 0].set_title('Line')
# Top-right
axes[0, 1].bar(['A', 'B', 'C'], [3, 7, 5])
axes[0, 1].set_title('Bar')
# Bottom-left
axes[1, 0].hist(np.random.randn(100), bins=20)
axes[1, 0].set_title('Histogram')
# Bottom-right
axes[1, 1].scatter(np.random.rand(50), np.random.rand(50))
axes[1, 1].set_title('Scatter')
plt.tight_layout()
plt.show()
Practical example
Sales visualization
import matplotlib.pyplot as plt
import pandas as pd
# Data
sales_data = pd.DataFrame({
'month': ['Jan', 'Feb', 'Mar', 'Apr', 'May', 'Jun'],
'sales': [150, 180, 165, 220, 250, 240],
'profit': [30, 45, 35, 60, 75, 70]
})
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(12, 5))
# Revenue trend
ax1.plot(sales_data['month'], sales_data['sales'],
marker='o', linewidth=2, markersize=8)
ax1.set_title('Monthly Sales Trend', fontsize=14, fontweight='bold')
ax1.set_xlabel('Month')
ax1.set_ylabel('Sales ($1,000s)')
ax1.grid(True, alpha=0.3)
# Profit bars
ax2.bar(sales_data['month'], sales_data['profit'],
color='green', alpha=0.7)
ax2.set_title('Monthly Profit', fontsize=14, fontweight='bold')
ax2.set_xlabel('Month')
ax2.set_ylabel('Profit ($1,000s)')
plt.tight_layout()
plt.savefig('sales_report.png', dpi=300)
plt.show()
Combining a line chart and a bar chart side by side in one figure, as ax1/ax2 do here, is a deliberate choice tied to what each chart type communicates best: a line naturally reads as a trend over time (is revenue going up or down), while a bar naturally reads as a comparison between discrete categories (which month had the highest profit) — using the same chart type for both would work, but would blur that distinction for the reader.
Styling and CJK font setup
Styling
# Built-in style
plt.style.use('seaborn-v0_8')
# Font for CJK labels (OS-specific; example: a Korean font on Windows)
plt.rcParams['font.family'] = 'Malgun Gothic'
plt.rcParams['axes.unicode_minus'] = False
# Figure size
plt.figure(figsize=(10, 6))
# Color palette
colors = ['#FF6B6B', '#4ECDC4', '#45B7D1']
Style names changed in Matplotlib 3.6: the old 'seaborn' styles were renamed to 'seaborn-v0_8' (and variants such as 'seaborn-v0_8-whitegrid'), and in 3.8 the old names were removed, so tutorials that use plt.style.use('seaborn') now fail with an OSError saying it is not a valid style. plt.style.available lists what your installed version supports. Style and rcParams changes are global for the rest of the session; with plt.style.context('ggplot'): applies a style to just the figures created inside the block.
The font family must be installed on the machine that renders the figure. Setting 'Malgun Gothic' works on Windows but falls back, with a findfont: Font family 'Malgun Gothic' not found warning, on a Linux server, and the labels then render as empty boxes. For portable code, list several candidates (plt.rcParams['font.family'] = ['Malgun Gothic', 'AppleGothic', 'Noto Sans CJK KR']) or ship a font file and register it with matplotlib.font_manager.fontManager.addfont(path).
axes.unicode_minus = False next to the CJK font line isn’t a coincidence — it’s fixing a real, specific interaction: many CJK fonts (this one included) don’t include a glyph for Matplotlib’s default Unicode minus sign character, so negative axis labels render as a missing-character box (☐) instead of a minus sign once you switch to a CJK font. Setting this to False tells Matplotlib to fall back to the plain ASCII hyphen for negative numbers, which every font can render.
Going deeper
Scatter with regression line and residual histogram (runnable)
Synthetic data with numpy, linear fit via polyfit, and a residual histogram—includes savefig for reports.
import numpy as np
import matplotlib.pyplot as plt
rng = np.random.default_rng(42)
x = np.linspace(0, 10, 80)
y = 2.5 * x + 1.0 + rng.normal(0, 1.8, size=x.shape)
coef = np.polyfit(x, y, 1)
y_hat = np.poly1d(coef)(x)
residuals = y - y_hat
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(11, 4), constrained_layout=True)
ax1.scatter(x, y, alpha=0.7, label="Observed")
ax1.plot(x, y_hat, color="crimson", linewidth=2, label="Linear fit")
ax1.set_title("Scatter with regression line")
ax1.legend()
ax1.grid(True, alpha=0.3)
ax2.hist(residuals, bins=18, color="steelblue", edgecolor="black", alpha=0.85)
ax2.set_title("Residual Distribution")
ax2.grid(True, alpha=0.3)
fig.savefig("regression_residuals.png", dpi=200)
plt.show()
Plotting the residuals (y - y_hat, the gap between each observed point and the fitted line) alongside the fit itself is a genuine diagnostic step, not just a second chart for variety — a good linear fit should leave residuals that look roughly like a symmetric, bell-shaped cluster around zero with no visible pattern. If the residual histogram instead looks skewed, or a residual-vs-x scatter shows a curve, that’s a signal the underlying relationship isn’t actually linear and a straight-line fit is the wrong model, something the scatter plot with its fit line alone can be surprisingly easy to miss.
Common mistakes
- Calling
savefigaftershow()in a script often yields an empty image. Once the window is closed,show()has released the figure, so the followingplt.savefig()saves a new, blank current figure. Save first, then show, or callfig.savefig()on a figure object you kept. - Missing fonts for non-Latin labels—configure per OS.
- Mixing OO API and pyplot state so artists land on the wrong axes. After
fig, (ax1, ax2) = plt.subplots(1, 2), a bareplt.title('...')titles only the most recently used axes, which is rarely the one you meant; useax1.set_title(...). - Labels cut off in saved files: long tick labels or a y-label can fall outside the figure.
plt.tight_layout(),constrained_layout=True, orsavefig(..., bbox_inches='tight')fix it. - Memory growth in loops: every
plt.figure()orplt.subplots()stays alive until closed, and after 20 open figures Matplotlib warns that too many figures are open. Close each one withplt.close(fig)after saving.
The last one is the problem I would warn about first for anyone generating reports: a script that saves one chart per customer or per day works fine for ten files and then slowly eats memory for a thousand, because pyplot keeps a reference to every figure it created.
Caveats
- Journals often prefer vector formats (PDF/SVG); for raster, set dpi explicitly.
- Consider colorblind-friendly palettes (
cividis, etc.).
In production
- Share styles via matplotlibrc or
plt.style.context. - For batch reports, call plt.close(fig) to free memory.
Alternatives
| Tool | Best for |
|---|---|
| Matplotlib | Fine control, papers, non-interactive backends |
| Seaborn | Quick statistical plots |
| Plotly | Interactive web charts |
Further reading
Related Articles
- Pandas Basics
- NumPy Basics
- Python Data Preprocessing
- Hands-on Data Analysis with Python
- Python Environment Setup
- Python Comprehensions | List· Dict