•8 min read

A Practical List of Annoying Click Tasks—and How to Automate Them with Python

A Practical List of Annoying Click Tasks—and How to Automate Them with Python
Python Automation

Automation

Click to reveal
Python Automation
Writing a script to perform a repetitive manual task automatically. Goal: eliminate boring work, not build complex systems.

Automation

Python Automation

pathlib.Path

Click to reveal
Python Automation
Modern file path handling: `Path('dir') / 'file.txt'`. Cross-platform, chainable, has read_text()/write_text() methods.

pathlib.Path

Python Automation

cron

Click to reveal
Python Automation
Linux/macOS job scheduler. Runs scripts automatically on a schedule. `crontab -e` to edit. Format: minute hour day month dow command.

cron

Python Automation

httpx

Click to reveal
Python Automation
Modern async-capable HTTP client for Python. Supports HTTP/2, connection pooling, and timeouts. Replacement for requests in 2026.

httpx

Python Automation

selectolax

Click to reveal
Python Automation
Blazing-fast HTML parser using the Modest CSS engine. 5-10x faster than BeautifulSoup for scraping structured pages.

selectolax

Automate Your Annoying Click Tasks with Python

If you've clicked the same sequence of buttons twice, it's a candidate for automation. Python excels at eliminating the boring parts of your day — from organizing files to scraping data, processing PDFs, and polling APIs. The key is staying disciplined: build tiny scripts that do one job well, not grand frameworks you'll never finish.

Audio Briefing
0:00 / 0:00
The Golden Rule

Don't build massive systems. Build tiny scripts to do one small job well.


How to Automate Safely

1. Pick a Hated Task

Identify something you do manually every day: renaming files, scraping a page, resizing images, processing PDFs, or organizing downloads. Write it down in one sentence — "Every Monday I download the CSV from the dashboard and upload it to the staging server." That sentence is your script.

2. Use the Standard Library First

Don't download heavy packages yet. pathlib (files), shutil (copying), csv (data), and datetime can handle 90% of your chores.

from pathlib import Path
for p in Path("downloads").glob("*.csv"):
    print(f"Would rename {p.name}")  # Always test with print first!

3. Treat Data with Respect

Never overwrite original data. Load the CSV, clean it (use pandas if it's complex), and write to a new output file. Add a timestamp to the output filename so you can audit runs: report_2026-09-22.csv.

4. Schedule and Log It

Use cron (Linux/Mac) or Task Scheduler (Windows) to run it automatically. Need help crafting cron schedules without syntax errors? Use our free visual Cron Expression Generator & Schedule Builder. Add logging so you know it worked when you weren't watching.

5. Add Alerting for Failures

A script that silently fails is worse than no script at all. Send yourself an email or push notification on error. The smtplib + ssl combo works with Gmail App Passwords and requires zero dependencies:

import smtplib, ssl

def send_failure_alert(error: str) -> None:
    with smtplib.SMTP_SSL("smtp.gmail.com", 465, context=ssl.create_default_context()) as s:
        s.login("you@gmail.com", "your-app-password")
        s.sendmail("you@gmail.com", "you@gmail.com",
                   f"Subject: Script failed\n\n{error}")

Advertisement

Common Traps to Avoid

Don't assume perfect inputs

Files will arrive with weird names. APIs will time out. Always add simple validation checks and handle errors gracefully instead of letting the script crash silently.

TrapBad ApproachBetter Approach
File names with spaces/special chars
String concatenation
pathlib.Path for safe handling
API timeouts
No timeout (hangs forever)
httpx.get(url, timeout=10)
Overwriting source data
Write back to same file
Write to new output file
Silent failures
Empty except: pass
logging.error with traceback
Scraping dynamic pages
Parsing JS-rendered HTML with regex
httpx + selectolax for static, Playwright for dynamic

Example 1: Downloads Organizer

from pathlib import Path
import shutil
import logging

logging.basicConfig(level=logging.INFO, format="%(asctime)s - %(message)s")

DOWNLOADS = Path.home() / "Downloads"
CATEGORIES = {
    "Images": [".jpg", ".jpeg", ".png", ".gif", ".webp"],
    "Documents": [".pdf", ".doc", ".docx", ".txt"],
    "Archives": [".zip", ".tar", ".gz", ".rar"],
    "Code": [".py", ".js", ".ts", ".json", ".yaml", ".yml"],
}

def main():
    for file in DOWNLOADS.iterdir():
        if not file.is_file():
            continue
        
        ext = file.suffix.lower()
        for category, extensions in CATEGORIES.items():
            if ext in extensions:
                dest_dir = DOWNLOADS / category
                dest_dir.mkdir(exist_ok=True)
                dest = dest_dir / file.name
                
                # Handle duplicates
                counter = 1
                while dest.exists():
                    stem = file.stem
                    dest = dest_dir / f"{stem}_{counter}{file.suffix}"
                    counter += 1
                
                shutil.move(str(file), str(dest))
                logging.info(f"Moved {file.name} → {category}/")
                break

if __name__ == "__main__":
    main()

Example 2: Web Scraping with httpx + selectolax

Use httpx (async-capable, HTTP/2 support) and selectolax (5–10× faster than BeautifulSoup) for scraping static pages:

import httpx
from selectolax.parser import HTMLParser
from pathlib import Path
import json
from datetime import datetime

HEADERS = {"User-Agent": "Mozilla/5.0 (compatible; PyScraper/1.0)"}

def scrape_hn_top() -> list[dict]:
    resp = httpx.get("https://news.ycombinator.com", headers=HEADERS, timeout=10)
    resp.raise_for_status()
    
    tree = HTMLParser(resp.text)
    items = []
    
    for row in tree.css("tr.athing"):
        title_node = row.css_first("span.titleline > a")
        if not title_node:
            continue
        items.append({
            "title": title_node.text(),
            "url": title_node.attributes.get("href", ""),
        })
    
    return items

def main():
    items = scrape_hn_top()
    output = Path(f"hn_{datetime.now().strftime('%Y-%m-%d')}.json")
    output.write_text(json.dumps(items, indent=2))
    print(f"Saved {len(items)} items → {output}")

if __name__ == "__main__":
    main()

Install with: pip install httpx selectolax

For pages that require JavaScript rendering, use playwright instead:

from playwright.async_api import async_playwright
import asyncio

async def scrape_dynamic(url: str) -> str:
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        page = await browser.new_page()
        await page.goto(url, wait_until="networkidle")
        content = await page.content()
        await browser.close()
    return content

Advertisement

Example 3: PDF Text Extraction and Processing

Extracting text from PDFs is a common enterprise automation task — invoice processing, contract analysis, log extraction.

import pdfplumber
from pathlib import Path
import csv

def extract_tables_from_pdf(pdf_path: Path) -> list[list[str]]:
    """Extract all tables from a PDF into a flat list of rows."""
    all_rows: list[list[str]] = []
    with pdfplumber.open(pdf_path) as pdf:
        for page in pdf.pages:
            for table in page.extract_tables():
                all_rows.extend(table)
    return all_rows

def process_invoices(input_dir: Path, output_csv: Path) -> None:
    rows: list[list[str]] = []
    for pdf in sorted(input_dir.glob("invoice_*.pdf")):
        extracted = extract_tables_from_pdf(pdf)
        rows.extend(extracted)

    with output_csv.open("w", newline="") as f:
        writer = csv.writer(f)
        writer.writerows(rows)

    print(f"Processed {len(list(input_dir.glob('*.pdf')))} PDFs → {output_csv}")

Install with: pip install pdfplumber


Example 4: API Polling with Exponential Backoff

Don't hammer an API on fixed intervals — use exponential backoff with jitter when polling for job completion:

import httpx
import time
import random
import logging

log = logging.getLogger(__name__)

def poll_until_done(job_id: str, api_key: str, max_retries: int = 10) -> dict:
    """Poll a job endpoint until status is 'done' or 'failed'."""
    base_url = f"https://api.example.com/jobs/{job_id}"
    headers = {"Authorization": f"Bearer {api_key}"}
    
    for attempt in range(max_retries):
        try:
            resp = httpx.get(base_url, headers=headers, timeout=15)
            resp.raise_for_status()
            data = resp.json()
            
            if data["status"] == "done":
                return data
            if data["status"] == "failed":
                raise RuntimeError(f"Job {job_id} failed: {data.get('error')}")
            
            # Still pending — back off with jitter
            wait = (2 ** attempt) + random.uniform(0, 1)
            log.info(f"Job {job_id} pending, retry {attempt + 1}/{max_retries}, wait {wait:.1f}s")
            time.sleep(wait)
            
        except httpx.HTTPStatusError as e:
            if e.response.status_code == 429:  # Rate limited
                wait = int(e.response.headers.get("Retry-After", 60))
                log.warning(f"Rate limited, waiting {wait}s")
                time.sleep(wait)
            else:
                raise
    
    raise TimeoutError(f"Job {job_id} did not complete in {max_retries} retries")

Scheduling with cron

Once your script is solid, schedule it. Edit crontab with crontab -e:

# Run downloads organizer every weekday at 6pm
0 18 * * 1-5 /usr/bin/python3 /home/user/scripts/organize_downloads.py >> /var/log/organizer.log 2>&1

# Scrape HN headlines every morning at 8am
0 8 * * * /usr/bin/python3 /home/user/scripts/scrape_headlines.py >> /var/log/scraper.log 2>&1

Always use absolute paths for both the Python binary and script path in cron — cron has a minimal PATH that doesn't include your virtualenv. Use which python3 (or which uv) to find the right binary.

For Python projects managed with uv, the cron line becomes:

0 8 * * * /home/user/.local/bin/uv run /home/user/scripts/scrape_headlines.py
Where to Start?

Try writing a script that monitors your "Downloads" folder and automatically moves images to a "Pictures" folder and PDFs to "Documents". Get it working, schedule it, and enjoy the saved time!

PythonAutomationProductivity

You Might Also Like

Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement