A Practical List of Annoying Click Tasks—and How to Automate Them with Python

Table of Contents
Automation
Automation
pathlib.Path
pathlib.Path
cron
cron
httpx
httpx
selectolax
selectolax
Automate Your Annoying Click Tasks with Python
If you've clicked the same sequence of buttons twice, it's a candidate for automation. Python excels at eliminating the boring parts of your day — from organizing files to scraping data, processing PDFs, and polling APIs. The key is staying disciplined: build tiny scripts that do one job well, not grand frameworks you'll never finish.
The Golden Rule
Don't build massive systems. Build tiny scripts to do one small job well.
How to Automate Safely
1. Pick a Hated Task
Identify something you do manually every day: renaming files, scraping a page, resizing images, processing PDFs, or organizing downloads. Write it down in one sentence — "Every Monday I download the CSV from the dashboard and upload it to the staging server." That sentence is your script.
2. Use the Standard Library First
Don't download heavy packages yet. pathlib (files), shutil (copying), csv (data), and datetime can handle 90% of your chores.
from pathlib import Path
for p in Path("downloads").glob("*.csv"):
print(f"Would rename {p.name}") # Always test with print first!
3. Treat Data with Respect
Never overwrite original data. Load the CSV, clean it (use pandas if it's complex), and write to a new output file. Add a timestamp to the output filename so you can audit runs: report_2026-09-22.csv.
4. Schedule and Log It
Use cron (Linux/Mac) or Task Scheduler (Windows) to run it automatically. Need help crafting cron schedules without syntax errors? Use our free visual Cron Expression Generator & Schedule Builder. Add logging so you know it worked when you weren't watching.
5. Add Alerting for Failures
A script that silently fails is worse than no script at all. Send yourself an email or push notification on error. The smtplib + ssl combo works with Gmail App Passwords and requires zero dependencies:
import smtplib, ssl
def send_failure_alert(error: str) -> None:
with smtplib.SMTP_SSL("smtp.gmail.com", 465, context=ssl.create_default_context()) as s:
s.login("you@gmail.com", "your-app-password")
s.sendmail("you@gmail.com", "you@gmail.com",
f"Subject: Script failed\n\n{error}")
Common Traps to Avoid
Don't assume perfect inputs
Files will arrive with weird names. APIs will time out. Always add simple validation checks and handle errors gracefully instead of letting the script crash silently.
| Trap | Bad Approach | Better Approach |
|---|---|---|
| File names with spaces/special chars | String concatenation | pathlib.Path for safe handling |
| API timeouts | No timeout (hangs forever) | httpx.get(url, timeout=10) |
| Overwriting source data | Write back to same file | Write to new output file |
| Silent failures | Empty except: pass | logging.error with traceback |
| Scraping dynamic pages | Parsing JS-rendered HTML with regex | httpx + selectolax for static, Playwright for dynamic |
Example 1: Downloads Organizer
from pathlib import Path
import shutil
import logging
logging.basicConfig(level=logging.INFO, format="%(asctime)s - %(message)s")
DOWNLOADS = Path.home() / "Downloads"
CATEGORIES = {
"Images": [".jpg", ".jpeg", ".png", ".gif", ".webp"],
"Documents": [".pdf", ".doc", ".docx", ".txt"],
"Archives": [".zip", ".tar", ".gz", ".rar"],
"Code": [".py", ".js", ".ts", ".json", ".yaml", ".yml"],
}
def main():
for file in DOWNLOADS.iterdir():
if not file.is_file():
continue
ext = file.suffix.lower()
for category, extensions in CATEGORIES.items():
if ext in extensions:
dest_dir = DOWNLOADS / category
dest_dir.mkdir(exist_ok=True)
dest = dest_dir / file.name
# Handle duplicates
counter = 1
while dest.exists():
stem = file.stem
dest = dest_dir / f"{stem}_{counter}{file.suffix}"
counter += 1
shutil.move(str(file), str(dest))
logging.info(f"Moved {file.name} → {category}/")
break
if __name__ == "__main__":
main()
Example 2: Web Scraping with httpx + selectolax
Use httpx (async-capable, HTTP/2 support) and selectolax (5–10× faster than BeautifulSoup) for scraping static pages:
import httpx
from selectolax.parser import HTMLParser
from pathlib import Path
import json
from datetime import datetime
HEADERS = {"User-Agent": "Mozilla/5.0 (compatible; PyScraper/1.0)"}
def scrape_hn_top() -> list[dict]:
resp = httpx.get("https://news.ycombinator.com", headers=HEADERS, timeout=10)
resp.raise_for_status()
tree = HTMLParser(resp.text)
items = []
for row in tree.css("tr.athing"):
title_node = row.css_first("span.titleline > a")
if not title_node:
continue
items.append({
"title": title_node.text(),
"url": title_node.attributes.get("href", ""),
})
return items
def main():
items = scrape_hn_top()
output = Path(f"hn_{datetime.now().strftime('%Y-%m-%d')}.json")
output.write_text(json.dumps(items, indent=2))
print(f"Saved {len(items)} items → {output}")
if __name__ == "__main__":
main()
Install with: pip install httpx selectolax
For pages that require JavaScript rendering, use playwright instead:
from playwright.async_api import async_playwright
import asyncio
async def scrape_dynamic(url: str) -> str:
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page()
await page.goto(url, wait_until="networkidle")
content = await page.content()
await browser.close()
return content
Example 3: PDF Text Extraction and Processing
Extracting text from PDFs is a common enterprise automation task — invoice processing, contract analysis, log extraction.
import pdfplumber
from pathlib import Path
import csv
def extract_tables_from_pdf(pdf_path: Path) -> list[list[str]]:
"""Extract all tables from a PDF into a flat list of rows."""
all_rows: list[list[str]] = []
with pdfplumber.open(pdf_path) as pdf:
for page in pdf.pages:
for table in page.extract_tables():
all_rows.extend(table)
return all_rows
def process_invoices(input_dir: Path, output_csv: Path) -> None:
rows: list[list[str]] = []
for pdf in sorted(input_dir.glob("invoice_*.pdf")):
extracted = extract_tables_from_pdf(pdf)
rows.extend(extracted)
with output_csv.open("w", newline="") as f:
writer = csv.writer(f)
writer.writerows(rows)
print(f"Processed {len(list(input_dir.glob('*.pdf')))} PDFs → {output_csv}")
Install with: pip install pdfplumber
Example 4: API Polling with Exponential Backoff
Don't hammer an API on fixed intervals — use exponential backoff with jitter when polling for job completion:
import httpx
import time
import random
import logging
log = logging.getLogger(__name__)
def poll_until_done(job_id: str, api_key: str, max_retries: int = 10) -> dict:
"""Poll a job endpoint until status is 'done' or 'failed'."""
base_url = f"https://api.example.com/jobs/{job_id}"
headers = {"Authorization": f"Bearer {api_key}"}
for attempt in range(max_retries):
try:
resp = httpx.get(base_url, headers=headers, timeout=15)
resp.raise_for_status()
data = resp.json()
if data["status"] == "done":
return data
if data["status"] == "failed":
raise RuntimeError(f"Job {job_id} failed: {data.get('error')}")
# Still pending — back off with jitter
wait = (2 ** attempt) + random.uniform(0, 1)
log.info(f"Job {job_id} pending, retry {attempt + 1}/{max_retries}, wait {wait:.1f}s")
time.sleep(wait)
except httpx.HTTPStatusError as e:
if e.response.status_code == 429: # Rate limited
wait = int(e.response.headers.get("Retry-After", 60))
log.warning(f"Rate limited, waiting {wait}s")
time.sleep(wait)
else:
raise
raise TimeoutError(f"Job {job_id} did not complete in {max_retries} retries")
Scheduling with cron
Once your script is solid, schedule it. Edit crontab with crontab -e:
# Run downloads organizer every weekday at 6pm
0 18 * * 1-5 /usr/bin/python3 /home/user/scripts/organize_downloads.py >> /var/log/organizer.log 2>&1
# Scrape HN headlines every morning at 8am
0 8 * * * /usr/bin/python3 /home/user/scripts/scrape_headlines.py >> /var/log/scraper.log 2>&1
Always use absolute paths for both the Python binary and script path in cron — cron has a minimal PATH that doesn't include your virtualenv. Use which python3 (or which uv) to find the right binary.
For Python projects managed with uv, the cron line becomes:
0 8 * * * /home/user/.local/bin/uv run /home/user/scripts/scrape_headlines.py
Where to Start?
Try writing a script that monitors your "Downloads" folder and automatically moves images to a "Pictures" folder and PDFs to "Documents". Get it working, schedule it, and enjoy the saved time!
You Might Also Like
Free In-Browser Developer Tools
Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.
Related Articles

Python Libraries I'd Bet On for 2026
I burned too many weekends on the wrong Python libraries so you don't have to. Here's what I actually use to ship things that stay working.
Read more
Polars vs Pandas: Performance and Memory Benchmarks
Comprehensive benchmark of Polars vs Pandas on memory utilization, Apache Arrow integration, lazy query optimization, and multi-core CPU scaling.
Read more
Python's Underscores: Conventions, Not Access Modifiers
Python uses underscores as naming conventions, not enforced access modifiers—understanding the difference prevents confusing code and helps you write better APIs.
Read more