Python Task Scheduling | Automate Jobs with schedule
Key takeaways
Schedule recurring Python jobs with the schedule library and APScheduler, understand cron syntax and DST pitfalls, and avoid the duplicate-execution trap that shows up once you scale past a single process.
Introduction
Every backend eventually needs something to happen on a timer: a nightly backup, an hourly price check, a report that goes out every Monday morning. The naive instinct is to write a while True: sleep(...) loop and move on. That works fine on a laptop script you run manually. It quietly breaks the moment your code runs as a long-lived service, gets restarted by a deploy, or gets scaled to more than one instance.
This article covers the two most common Python scheduling tools — the schedule library and APScheduler — along with the OS-level schedulers (cron, Windows Task Scheduler) that back them in production. More importantly, it covers why the simple approach fails once real infrastructure gets involved: duplicate execution across replicas, missed runs after downtime, and daylight-saving-time bugs that only show up twice a year. Those are the failures that actually page someone at 3 a.m., and understanding them up front will save you a rewrite later.
Task scheduling means running work at defined times or intervals without manual intervention. The interesting engineering problem is not “how do I call a function every hour” — that part is trivial. It’s “how do I guarantee it runs exactly once per interval, even when the process restarts, even when there are three copies of it running, and even when the clock jumps an hour in autumn.”
Why the scheduler you pick depends on your deployment shape
Before writing any code, it is worth being honest about how the service will actually run in production, because that answers which tool is appropriate.
- Single script, run manually or via a simple
systemdservice, one instance ever:scheduleis fine. It is a thin wrapper around checking “has enough time passed” in a loop, and that is exactly the problem you have. - Long-running service, but you might redeploy or crash-restart it: you need APScheduler with a persistent job store, not
schedule’s in-memory state, because in-memory schedules forget everything on restart. - Containerized service running with more than one replica (Kubernetes Deployment with
replicas: 2+, multiple Gunicorn/uWSGI workers, autoscaling): neitherschedulenor a bare in-process APScheduler is safe by default. Every replica runs its own copy of the scheduler, and every copy thinks it is the only one, so every copy fires the job. This is the single most common scheduling bug in production Python services, and it is invisible in local development because you only ever run one instance there. - Distributed task processing across many workers, or jobs that need retries, chaining, or a dashboard: Celery with
celery beat(or an external cron-like trigger hitting a queue) is the right tool. It separates “decide it’s time to run this job” from “actually run this job,” which is exactly the separation that fixes the duplicate-execution problem.
Keep that mapping in mind while reading the examples below — the code is nearly identical between the “fine” and “will duplicate-fire in production” cases, so the tool alone does not save you. The deployment shape does.
schedule library
Install
pip install schedule
Basics
import schedule
import time
def job():
print("Job running!")
# Every 10 seconds
schedule.every(10).seconds.do(job)
# Every minute
schedule.every(1).minutes.do(job)
# Daily at 9:00 AM
schedule.every().day.at("09:00").do(job)
# Every Monday at 10:00 AM
schedule.every().monday.at("10:00").do(job)
# Run loop
while True:
schedule.run_pending()
time.sleep(1)
schedule keeps all of its state — the registered jobs and their next run times — as plain Python objects living in the process’s memory. run_pending() just walks that list every time it is called and fires anything whose next-run time has passed. There is no persistence, no cross-process coordination, and no concept of “this job was supposed to run while I was down.” That simplicity is the whole appeal: no database, no broker, just a loop. It is also exactly what makes it unsuitable for anything beyond a single always-on process, because the moment that process restarts, every job’s history is gone and the schedule starts fresh from “now.”
One subtlety that trips people up: schedule.every().day.at("09:00") uses the local system time of wherever the process runs. If your container’s base image time zone is UTC and you deploy across regions, or if the host’s time zone setting differs from what you tested locally, “9:00” silently means something different in production than it did on your laptop. Always pin the container’s TZ environment variable explicitly rather than assuming it matches your dev machine.
Real-world example
Automated backup
import schedule
import time
from datetime import datetime
import shutil
from pathlib import Path
def backup_files():
"""Copy a data folder to backups."""
timestamp = datetime.now().strftime('%Y%m%d_%H%M%S')
backup_name = f"backup_{timestamp}"
source = Path('./data')
backup_dir = Path('./backups')
backup_dir.mkdir(exist_ok=True)
backup_path = backup_dir / backup_name
shutil.copytree(source, backup_path)
print(f"[{datetime.now()}] Backup complete: {backup_name}")
# Daily at midnight
schedule.every().day.at("00:00").do(backup_files)
# Every Sunday at 11:00 PM
schedule.every().sunday.at("23:00").do(backup_files)
print("Backup scheduler started...")
while True:
schedule.run_pending()
time.sleep(60)
This looks harmless, but it hides a real failure mode: what happens if the process is down at midnight? Say the host reboots for a kernel update at 23:55 and comes back up at 00:20. schedule has no idea a midnight backup was supposed to happen — it only checks “is it past the scheduled time and have I not already run today,” and depending on how you structure the check, you can either silently skip that day’s backup entirely, or (worse) fire it immediately on startup with no rate limiting if you built your own “catch-up” logic carelessly. Neither is good: silently skipping a backup means you find out you have a gap in your backup history exactly when you need a restore point that isn’t there.
If backups (or any job) genuinely cannot be allowed to silently disappear when the process happens to be down at the scheduled instant, schedule is the wrong tool. This is precisely the “misfire” problem that APScheduler has an explicit policy for, covered below.
APScheduler
Install
pip install apscheduler
Advanced scheduling
from apscheduler.schedulers.blocking import BlockingScheduler
from datetime import datetime
scheduler = BlockingScheduler()
def job1():
print(f"[{datetime.now()}] Job 1")
def job2():
print(f"[{datetime.now()}] Job 2")
# Cron-style
scheduler.add_job(job1, 'cron', hour=9, minute=0) # daily at 9:00
# Interval
scheduler.add_job(job2, 'interval', minutes=30) # every 30 minutes
# Weekdays at 9:00
scheduler.add_job(
job1,
'cron',
day_of_week='mon-fri',
hour=9,
minute=0
)
scheduler.start()
APScheduler’s real advantage over schedule is not the cron-style syntax — it is the job store abstraction. By default, BlockingScheduler() still keeps jobs in memory, so it has the exact same “forgets everything on restart” problem as schedule. The difference is that APScheduler lets you swap in a persistent store:
from apscheduler.schedulers.blocking import BlockingScheduler
from apscheduler.jobstores.sqlalchemy import SQLAlchemyJobStore
scheduler = BlockingScheduler(
jobstores={'default': SQLAlchemyJobStore(url='sqlite:///jobs.sqlite')},
job_defaults={
'coalesce': True, # collapse missed runs into a single run
'misfire_grace_time': 3600, # allow up to 1 hour of lateness before treating as a miss
},
)
coalesce and misfire_grace_time are the settings that actually matter for correctness. Without a grace time, if the process was down when a job was due, APScheduler treats that run as “missed” and simply drops it rather than firing late — which is the right default for something like “send an hourly status ping” but the wrong default for “send the daily digest email,” which should probably still go out even if it’s an hour late. coalesce=True matters when a job store has accumulated several missed runs (say, the process was down for six hours and the job was supposed to run every 15 minutes): instead of firing 24 backlogged runs back-to-back the instant the process comes back up, it collapses them into one. Get this wrong and you can accidentally hammer a downstream API with a burst of catch-up calls the moment a deploy finishes.
Timezone and DST pitfalls in cron expressions
Both cron-style APScheduler triggers and real crontab entries are specified in wall-clock time, and wall-clock time is not a monotonic, uniform thing — it jumps forward or backward twice a year in regions that observe daylight saving time. A job scheduled for hour=2, minute=30 in a US time zone that observes DST will, on the “spring forward” night, have a 2:30 a.m. that literally does not exist (clocks jump from 1:59:59 straight to 3:00:00), and on “fall back” night, 1:30 a.m. happens twice. Different schedulers handle this differently — some skip the nonexistent run, some fire it once anyway at the nearest valid time, some fire the ambiguous run twice. The only reliable fix is to either run your servers in UTC (recommended for anything backend, since UTC has no DST) or explicitly test your scheduler’s documented DST behavior rather than assuming. I have seen a “runs at 2 a.m. daily” job silently not run for exactly one day per year for over a year before anyone noticed, because nobody was specifically watching for a once-a-year gap.
flowchart TD
A["Job scheduled: 02:30 local time"] --> B{"DST transition night?"}
B -->|No| C["Runs normally"]
B -->|"Spring forward\n(clock skips 02:00-03:00)"| D["02:30 does not exist\nJob may be skipped or shifted"]
B -->|"Fall back\n(01:00-02:00 repeats)"| E["02:30 unambiguous,\nbut earlier hour runs twice"]
D --> F["Recommendation: schedule in UTC"]
E --> F
The multi-replica trap: why the same job runs twice
This is the failure mode worth spending the most time on, because it is the one that does not show up in development and does show up in production the moment you scale past one instance.
Picture a typical setup: a FastAPI or Django service packaged into a container image, deployed as a Kubernetes Deployment, and — because “the scheduler is part of the app” felt convenient at the time — the same container also starts an in-process APScheduler on boot to send a daily digest email. It works perfectly in staging, where the deployment runs a single replica. Then traffic grows, someone bumps replicas: 1 to replicas: 3 for availability, and now three independent processes each start their own copy of that scheduler on their own boot, each with no idea the other two exist. At 8:00 a.m., all three fire the “send daily digest” job. Users get three identical emails.
I have personally chased exactly this bug, and the debugging path is worth describing because it is not obvious from the symptom alone. The first report is usually “customer says they got the report twice” — not “getting it three times,” because email deduplication or a flaky network can hide the third copy, which makes you initially suspect a client-side rendering bug or a double-send inside the email library itself. You check the job code, see a single scheduler.add_job(...) call, and it looks completely correct in isolation — because it is. The bug isn’t in the job registration; it’s in the assumption that only one instance of that registration code would ever execute at once. The tell, once you know to look for it, is in the timestamps: if you log the pod/hostname alongside the job’s execution timestamp, you’ll see two or three different hostnames firing within the same second, not one process firing twice. That immediately rules out “double-registered job” and points straight at “multiple replicas, each running its own scheduler.”
The fixes, roughly in order of how much infrastructure they require:
- Run the scheduler in exactly one place. Split the scheduler out of the web-serving replicas into its own single-replica Deployment (or a Kubernetes
CronJobfor genuinely periodic work). This is the simplest fix and often the right one — the scheduler doesn’t need to scale with request traffic, so it shouldn’t live inside the pods that do. - Use a shared job store with locking. APScheduler’s
SQLAlchemyJobStoreorRedisJobStore, combined with only one scheduler instance actually calling.start()(gated behind a leader-election check, or an environment variable set only on one replica), lets multiple processes see the same schedule state without all of them executing it. - Move to
celery beatplus workers. Celery’s design already separates “the beat process decides a task is due” from “a worker executes it,” and you run exactly onebeatprocess regardless of how many workers you scale to. This is the standard answer once a project outgrows in-process scheduling, precisely because it eliminates this entire class of bug by construction rather than by convention. - Use the platform’s own scheduler. Kubernetes
CronJob, a cloud provider’s scheduled function trigger, or your CI system’s scheduled pipelines all guarantee single-invocation semantics without your application code needing to reason about it at all.
sequenceDiagram
participant P1 as Replica 1
participant P2 as Replica 2
participant P3 as Replica 3
participant DB as Job store
Note over P1,P3: In-process scheduler, no coordination
P1->>P1: 08:00 fires "send digest"
P2->>P2: 08:00 fires "send digest"
P3->>P3: 08:00 fires "send digest"
Note over P1,P3: Result: 3 identical emails sent
Note over P1,P3: Fixed with a leader lock
P1->>DB: acquire lock "digest-job"
DB-->>P1: lock granted
P2->>DB: acquire lock "digest-job"
DB-->>P2: lock denied
P3->>DB: acquire lock "digest-job"
DB-->>P3: lock denied
P1->>P1: fires "send digest" once
If you only take one thing from this section: treat “does this run correctly with N replicas” as a real question you answer before deploying a scheduler change, not something you discover from a support ticket. It costs nothing to ask in code review; it costs a support queue full of duplicate-email complaints to discover later.
Automating web scraping
Periodic data collection
import schedule
import requests
from bs4 import BeautifulSoup
import pandas as pd
from datetime import datetime
def scrape_prices():
"""Collect product prices."""
urls = [
'https://shop1.com/product/123',
'https://shop2.com/product/456'
]
prices = []
for url in urls:
try:
response = requests.get(url, timeout=10)
soup = BeautifulSoup(response.text, 'html.parser')
price_text = soup.select_one('.price').text
price = int(price_text.replace(',', '').replace('$', ''))
prices.append({
'url': url,
'price': price,
'timestamp': datetime.now()
})
except Exception as e:
print(f"Error: {url} - {e}")
df = pd.DataFrame(prices)
df.to_csv('price_history.csv', mode='a', header=False, index=False)
print(f"[{datetime.now()}] Price scrape finished")
# Every hour
schedule.every(1).hours.do(scrape_prices)
while True:
schedule.run_pending()
time.sleep(60)
Scraping jobs deserve extra caution around timeouts and error isolation, which the try/except inside the loop is already doing correctly: one failing URL should never take down the whole batch. What this snippet does not handle, and what real scraping schedules eventually need, is backoff on repeated failures. If shop2.com starts returning a CAPTCHA page or blocks your IP outright, an hourly schedule will hammer it once an hour forever, logging the same error each time, until a human notices. A more resilient version tracks consecutive failures per URL and skips (or slows down) a source after N failures in a row, rather than treating every run as independent. This also matters ethically and legally — a scraper that keeps retrying a source that has started rejecting it is indistinguishable from abusive traffic, and getting your source IP blocked entirely is a worse outcome than missing an hour of price data.
The .price CSS selector is also a single point of fragility: any redesign of the target site’s markup silently breaks the parse, and soup.select_one('.price') returns None, which then raises AttributeError on .text — caught by the broad except Exception, but logged only as a generic error string with no indication that the cause was “the selector no longer matches.” In a long-running scrape job, logging which selector failed (not just that something failed) is the difference between a five-minute fix and an afternoon of re-discovering the same root cause.
Email notifications
Daily report email
import smtplib
from email.mime.text import MIMEText
from email.mime.multipart import MIMEMultipart
from datetime import datetime
import schedule
def send_daily_report():
"""Send a daily HTML report by email."""
report = generate_report()
msg = MIMEMultipart()
msg['From'] = '[email protected]'
msg['To'] = '[email protected]'
msg['Subject'] = f"Daily report - {datetime.now().strftime('%Y-%m-%d')}"
msg.attach(MIMEText(report, 'html'))
try:
server = smtplib.SMTP('smtp.gmail.com', 587)
server.starttls()
server.login('[email protected]', 'password')
server.send_message(msg)
server.quit()
print("Report sent")
except Exception as e:
print(f"Send failed: {e}")
def generate_report():
"""Build simple HTML for the report."""
return """
<html>
<body>
<h1>Daily report</h1>
<p>Total revenue: 1,000,000 KRW</p>
<p>New customers: 50</p>
</body>
</html>
"""
# Every day at 8:00 AM
schedule.every().day.at("08:00").do(send_daily_report)
Two things worth flagging beyond the obvious “don’t hardcode credentials in source” (use environment variables or a secrets manager, always). First, this job has no retry logic at all: if the SMTP connection times out for any reason — a transient network blip, Gmail rate-limiting, a brief outage — the report for that day is simply gone, silently logged to stdout where nobody is watching. A daily report is exactly the kind of job where “fail loudly and retry” matters more than for a scrape job, because there’s no next hourly run to paper over the gap; the next opportunity is 24 hours away. Wrapping the send in a small retry-with-backoff (three attempts, a few seconds apart) costs a handful of lines and closes most of that gap.
Second, and this ties back to the multi-replica problem above: an email-sending job is one of the worst jobs to accidentally duplicate, because unlike a backup or a scrape, the failure is directly visible to your users and erodes trust in the product (“why did I get three copies of this?”). If this job lives inside a horizontally-scaled service, it is a strong candidate for extraction into its own single-instance scheduler or a celery beat task specifically because the blast radius of getting it wrong is a customer-facing one.
Running in the background
Windows Task Scheduler
# Register a Python script with Windows Task Scheduler
# 1. Open Task Scheduler
# 2. Create Basic Task
# 3. Program: python.exe
# 4. Arguments: C:\path\to\script.py
Linux cron
# Edit crontab
crontab -e
# Daily at 9:00 AM
0 9 * * * /usr/bin/python3 /path/to/script.py
# Every hour
0 * * * * /usr/bin/python3 /path/to/script.py
# Every Monday at 10:00 AM
0 10 * * 1 /usr/bin/python3 /path/to/script.py
Handing scheduling off to the OS avoids the in-process problems above almost entirely — cron and Task Scheduler are external to your application, so a crashed or redeployed app doesn’t erase the schedule, and (on a single machine) there’s no risk of “three replicas each thinking they own the job,” because cron itself only runs on one host. The trade-off moves elsewhere: plain cron does not retry or track missed runs either. If the machine is powered off or the cron daemon isn’t running at 0 9 * * *, that invocation simply never happens and cron has no memory of it — there is no catch-up. (anacron, common on Debian-based desktop distros, exists specifically to patch this gap for machines that aren’t always on, but it is not the default on most server distributions and shouldn’t be assumed present.) On a server that’s supposed to be up 24/7, this is rarely an issue in practice; on a scheduled batch machine that gets powered down overnight, it is exactly the same “the job silently didn’t run” problem discussed with schedule earlier, just moved one layer down the stack.
It’s also worth remembering that crontab -e edits belong to whichever user runs the command, and a cron job’s environment is deliberately minimal — it does not inherit your shell’s PATH, virtualenv activation, or environment variables. A script that imports fine when you run it manually from your shell can fail inside cron with ModuleNotFoundError simply because cron invoked a different python3 than the one you tested with. Always use an absolute path to the interpreter inside the virtualenv (e.g. /home/app/venv/bin/python3) rather than relying on PATH resolution inside a cron job.
Handling failures, timeouts, and observability in production
Once a schedule leaves your laptop, three concerns become non-optional: you need to know a job ran (or didn’t), a hung job cannot be allowed to block every job after it, and errors need to surface somewhere a human will actually see them.
import logging
logging.basicConfig(
filename='scheduler.log',
level=logging.INFO,
format='%(asctime)s - %(message)s'
)
def job():
logging.info("Job start")
# work here
logging.info("Job done")
def safe_job():
try:
job()
except Exception as e:
logging.error(f"Error: {e}")
import signal
def timeout_handler(signum, frame):
raise TimeoutError("Job timed out")
signal.signal(signal.SIGALRM, timeout_handler)
signal.alarm(300) # 5 minute cap
The logging start/end pair is more useful than it looks: without it, “did last night’s job actually run” is a question you can only answer by checking whether its side effects happened (does the backup file exist? did the email arrive?), which is a much slower feedback loop than grepping a log for yesterday’s Job start / Job done pair. Wrapping the call in safe_job() matters specifically because both schedule and a while True loop are single-threaded by default — an unhandled exception inside one job’s callback can kill the entire scheduling loop, silently taking every other future job down with it. That is a much worse outcome than the one failing job alone, and it is easy to miss in testing because your test runs rarely raise exceptions on purpose.
The SIGALRM-based timeout is a reasonable default on Unix, but it comes with a real limitation worth knowing before you rely on it: signal.alarm() only works on the main thread and only on Unix-like systems, so it silently does nothing useful on Windows and nothing at all if your job runs in a background thread. For cross-platform code, or code running inside a thread pool, concurrent.futures.ProcessPoolExecutor with .result(timeout=300) (which raises on timeout and can actually kill the worker process, unlike a thread-based timeout) is the more portable equivalent — worth the extra setup if you’ve been burned by a job that hangs on a network call with no timeout of its own.
Related Articles
- Python file automation | Organize, rename, and back up files
- Python series index
- Python environment setup
- Python basic syntax
- Python Web Scraping | BeautifulSoup and Selenium Explained
- Node.js Async Programming: Callbacks, Promises, and Async/Await
Frequently Asked Questions (FAQ)
Q. Why does my scheduled job run twice after I start the app with multiple Gunicorn workers?
A. Each worker is a separate process that imports your app, so each one starts its own scheduler, and every scheduler fires the job: with four workers you get four runs. Run the scheduler in exactly one dedicated process, use a shared job store with only one instance calling .start(), or move the job to cron, a Kubernetes CronJob or Celery beat, which trigger it once outside the web workers.